System
The system addresses communication challenges for individuals with severe motor disabilities by using gaze input and speech synthesis, allowing seamless and natural interaction with voice output and emergency notifications.
Patent Information
- Application Number
- JP2024120578
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional communication tools for people with severe motor disabilities require external sensors or dedicated devices for gaze input, limiting seamless user experience, and those with speech difficulties are restricted to text-based communication, making voice communication difficult.
A system that captures a user's gaze using a device's camera, analyzes gaze data to identify coordinates, generates text data, and inputs it into a speech synthesis algorithm to enable speech, while storing data in a database for later analysis and sending notifications based on specific keywords.
Enables seamless and comfortable communication for individuals with severe motor impairments without external sensors, supports natural and emotionally rich interaction, and ensures user safety through prompt notifications.
Smart Images

Figure 2026019169000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional communication tools for people with severe motor disabilities require external sensors or dedicated devices to enable gaze input, preventing a seamless user experience. Furthermore, for people with speech difficulties, communication is limited to text, making voice communication difficult. As a result, there has been a lack of technology to enable people with severe motor disabilities to communicate comfortably. [Means for solving the problem]
[0005] The present invention provides a system that captures a user's gaze using a device's camera and analyzes the captured gaze data to identify coordinates corresponding to the gaze. The system also includes means for generating text data using the identified coordinates and inputting the generated text data into a speech synthesis algorithm to generate speech. This enables people with severe motor impairments to communicate more seamlessly and comfortably without the need for external sensors or specialized equipment. The system also includes means for storing the text data and speech data in a database for easy later analysis and reference. The system also includes means for generating notifications based on specific keywords and sending them to pre-defined contacts, enabling prompt and reliable response in emergencies.
[0006] "Device camera" means a camera built into an electronic device that is used for gaze capture and recognition of the user's face and eye position.
[0007] "Gaze data" is digital data that indicates the direction and position of the user's gaze.
[0008] "Coordinates" are numerical information indicating a specific position on the screen that is identified based on the line-of-sight data.
[0009] "Text data" is digital data expressed as a collection of characters selected based on the position of the gaze.
[0010] A "speech synthesis algorithm" is a computational method and program that converts text data into a human voice.
[0011] A "database" is a computer system for storing and managing text and audio data.
[0012] "Notifications" are alert messages generated based on specific keywords or events and sent to pre-defined contacts.
[0013] "Contacts" means the phone numbers, email addresses, or other correspondence addresses that you have pre-configured to receive notifications. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The system of the present invention utilizes the device's camera to capture a user's gaze and analyzes the gaze data to determine the coordinates where the user is looking. It then generates text data based on the coordinates and inputs the generated text data into a speech synthesis algorithm to generate speech so that the user can speak. It also includes functionality to store the generated text and speech data in a database and generate notifications based on specific keywords and send them to pre-defined contacts.
[0036] Program processing explanation
[0037] Gaze recognition
[0038] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The captured gaze data is pre-processed to remove noise and generate clear data. This gaze data is sent to a server and analyzed using a machine learning algorithm. After analysis, the coordinates where the user is looking are identified and the coordinate data is sent to the device.
[0039] Examples:
[0040] If a user is using their gaze to operate a virtual keyboard in an application and is trying to input the word "thank you," the device will generate the input string "thank you" as the user moves their gaze to each letter.
[0041] Text input
[0042] The coordinates obtained as a result of analyzing the gaze data are mapped to specific characters or icons, and the corresponding characters are entered. The device displays the received character data on the screen and provides real-time feedback to the user, ensuring that the user enters exactly what they intended.
[0043] Examples:
[0044] As the user types "Thank you," the characters are displayed on the screen in real time, providing a clear view of the user's progress.
[0045] Voice generation
[0046] The entered text data is sent from the device to the server, where a speech synthesis algorithm generates speech. The generated speech is customized using pre-stored samples of the user's voice and sent to the device, where it is played back, allowing the user to speak.
[0047] Examples:
[0048] When the text data "Thank you" is generated, the server converts it into speech and uses a pre-stored sample of the user's voice to generate a speech that pronounces "Thank you," which the device can then play back and convey to people around it.
[0049] Data storage and notification function
[0050] The generated text and voice data is stored in a database on the server, and the server also sends notifications to pre-defined contacts when specific keywords or events occur, enabling rapid response in emergencies.
[0051] Examples:
[0052] If a user types "help," the server analyzes the text data and sends an email or SMS notification to contacts registered as an emergency message.
[0053] This system enables people with severe motor disabilities to communicate seamlessly and comfortably using gaze input and generative AI without the need for external sensors or specialized devices. The addition of a voice generation function also enables more natural and emotionally rich communication. Furthermore, data storage and notification functions ensure user safety and comfort.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] When the application starts, the device activates the device's camera, which recognizes the user's face and eye position and begins capturing gaze data.
[0057] Step 2:
[0058] The device captures gaze data at regular intervals (e.g., every 100 milliseconds) and performs preprocessing, which involves removing noise and normalizing the data.
[0059] Step 3:
[0060] The device sends preprocessed gaze data, including the user's face position and gaze direction, to the server.
[0061] Step 4:
[0062] The server inputs the received gaze data into a machine learning algorithm for analysis, which identifies the coordinates where the user is looking.
[0063] Step 5:
[0064] The server transmits the identified coordinate data to the terminal.
[0065] Step 6:
[0066] Based on the received coordinate data, the device identifies the corresponding characters or icons and displays on the screen that the user has entered the input they intended.
[0067] Step 7:
[0068] The user operates the virtual keyboard with their eyes, selecting characters one after another. The selected characters are concatenated in real time to generate text data.
[0069] Step 8:
[0070] The terminal transmits the generated text data to the server.
[0071] Step 9:
[0072] The server feeds the text data into a speech synthesis algorithm to generate an audio file, which is customized using pre-stored samples of the user's voice.
[0073] Step 10:
[0074] The server transmits the generated audio file to the terminal.
[0075] Step 11:
[0076] The terminal plays the received audio file, and the user's intention is spoken aloud.
[0077] Step 12:
[0078] The server stores all text and audio data in a database and manages it so that it can be referenced and analyzed at a later date.
[0079] Step 13:
[0080] Users or medical staff configure notification settings based on specific keywords or conditions, and when the conditions are met, the server automatically sends notifications to designated contacts.
[0081] Step 14:
[0082] In the event of an emergency, if a preset keyword (e.g., "help") is entered, the server will send an emergency message to pre-defined contacts, allowing for a prompt and appropriate response.
[0083] These processing steps realize a system that integrates gaze recognition, character input, voice generation, data management and notification functions.
[0084] Example 1
[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0086] This system solves the problem of people with severe motor disabilities finding it difficult to achieve seamless and comfortable communication using gaze input and generative AI without the need for external sensors or specialized devices. It also enables more natural and emotional communication by reliably inputting the user's intended content and enabling it to be spoken aloud, and it also solves the problem of a lack of means to send prompt notifications in emergencies.
[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0088] In this invention, the server includes a means for analyzing gaze data and identifying the coordinates where the user is looking, a means for storing text data and voice data in a database, and a means for generating a notification based on a specific keyword and sending it to a pre-set contact, thereby making it possible to input characters based on the user's gaze, convert the input text data into voice, and send a notification under specific conditions based on the stored data.
[0089] "Device camera" refers to a camera used to capture a user's gaze, including cameras built into electronic devices such as smartphones, tablets, and computers.
[0090] "Gaze Data" refers to data about a user's eye movements and pupil position captured by a device's camera.
[0091] "Preprocessing" refers to the process of removing noise from captured gaze data and converting it into a format suitable for analysis.
[0092] The term "server" refers to a computer system that receives gaze data, analyzes it, and returns the necessary data to the terminal.
[0093] "Analysis" refers to the process of using machine learning algorithms to identify the coordinates where the user is looking from the gaze data.
[0094] "Coordinate data" refers to data indicating the position at which the user's gaze is pointing, as determined by analysis.
[0095] "Terminal" refers to an electronic device that has an interface with the user, accepts eye-gaze input, and receives data from a server.
[0096] "Text data" refers to character string data generated based on the user's line of sight.
[0097] "Speech synthesis algorithm" refers to the calculation procedures or programs used to convert text data into speech.
[0098] "Database" refers to an information system for systematically storing generated text and audio data.
[0099] "Specific keywords" refer to pre-defined important words or phrases contained in the text data entered by the user.
[0100] "Notifications" refer to messages or alerts sent to pre-defined contacts when certain keywords are detected.
[0101] The system of the present invention utilizes a device's camera to capture a user's gaze and analyzes the captured gaze data to determine the coordinates where the user is looking. This coordinate data is then used to generate text data, which is then input into a speech synthesis algorithm to generate speech. The system also stores the generated text and speech data, and generates notifications based on specific keywords and sends them to pre-defined contacts.
[0102] Gaze recognition
[0103] First, when the application is launched, the device starts the device camera and captures the user's gaze data at regular intervals. For this purpose, the camera needs to capture the position of the user's eyes at a high frame rate. For example, it is appropriate for the camera to capture images at 30 frames per second.
[0104] The device then preprocesses the captured gaze data to remove noise and generate clearer data, including high-frequency noise removal filters and grayscale conversion.
[0105] Data analysis
[0106] The preprocessed gaze data is encrypted and securely transmitted over the network to a server, where it is analyzed using machine learning algorithms to identify the coordinates where the user is looking. This analysis can potentially involve the application of deep learning models.
[0107] Text input
[0108] The analyzed coordinate data is sent to the device in real time, and characters and icons are identified based on that data. For example, when the user directs their gaze to each key on a virtual keyboard, the corresponding character is identified. Then, as the user moves their gaze, an input string is generated, and the device displays the string on the screen in real time.
[0109] Voice generation
[0110] The generated text data is sent from the device to a server, where it is converted into speech using a speech synthesis algorithm. The server uses pre-stored samples of the user's voice to generate a natural-sounding voice. This speech data is sent to the device and played back according to the user's settings.
[0111] Data Retention and Notification
[0112] The generated text and voice data is stored in a database on the server. In addition, the server sends notifications to pre-defined contacts when specific keywords or events occur, enabling a rapid response in the event of a user emergency.
[0113] Examples:
[0114] If a user is using their gaze to operate a virtual keyboard within an application and is trying to input the word "thank you," they can identify each character by directing their gaze at each letter. As the user looks at each letter in "thank you," the device generates the input string for "thank you." The generated text data is converted into speech by a connected speech synthesis algorithm, and the device plays it back. For example, the device can play back the audio "thank you" and convey the message to people around them.
[0115] Example prompt sentence:
[0116] "Please explain in natural language a program that uses the device's camera to capture the user's gaze, analyzes the gaze data, and identifies the coordinates where the user is looking. Please also explain the process of generating text data based on the coordinate data, and inputting that data into a speech synthesis algorithm to generate speech."
[0117] This system provides a simple and intuitive means of input using eye gaze control for people with severe motor disabilities, and supports richer communication by outputting the input as voice. It also ensures the user's safety by providing a notification function in case of an emergency.
[0118] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0119] Step 1:
[0120] When the application is launched, the device starts the device's camera and captures the user's gaze data at regular intervals. As input, it receives image data from the device's camera. The camera captures the user's eye position at 30 frames per second. As output, it generates captured gaze data. This gaze data includes the user's eye position information for each image frame.
[0121] Step 2:
[0122] The device preprocesses the captured gaze data to generate clear data with noise removed. It receives the captured gaze data as input, performs high-frequency noise removal filtering and grayscale conversion, and generates clear preprocessed gaze data as output.
[0123] Step 3:
[0124] The device sends the preprocessed gaze data to the server. As input, it receives preprocessed clear gaze data. The data is encrypted and securely transmitted over the network. As output, it generates the gaze data sent to the server.
[0125] Step 4:
[0126] The server analyzes the received gaze data using a machine learning algorithm to identify the coordinates where the user is looking. Preprocessed gaze data is received as input. A deep learning model is used as the machine learning algorithm. Specifically, it extracts features from the gaze data, inputs them into the model, and predicts the coordinates. As output, it generates coordinate data where the user is looking.
[0127] Step 5:
[0128] The server sends the parsed coordinate data to the device. As input, it receives the coordinate data of the user's view. The data is sent back to the device in real time. As output, it generates the coordinate data sent to the device.
[0129] Step 6:
[0130] The device identifies the character or icon the user is looking at based on the received coordinate data. As input, it receives the coordinate data the user is looking at. The coordinates are mapped to each key on the virtual keyboard, and when the gaze is directed, the character is identified. Specifically, it compares the coordinates with the key position on the virtual keyboard and identifies the matching character. As output, it generates the identified character data.
[0131] Step 7:
[0132] The terminal displays the identified character data on the screen in real time. As input, it receives the identified character data. As a specific action, it displays the character data in an appropriate position on the screen. As output, it provides visual feedback to the user.
[0133] Step 8:
[0134] The terminal generates text data based on the identified character data. As input, it receives a sequence of character data. As a specific operation, it concatenates the identified characters to create one piece of text data. As output, it obtains the generated text data.
[0135] Step 9:
[0136] The device sends the generated text data to the server and generates speech using a speech synthesis algorithm. The generated text data is received as input. The server generates natural-sounding speech using pre-stored samples of the user's voice. Specific operations include extracting phonemes from the text and generating a speech waveform. The generated speech data is obtained as output.
[0137] Step 10:
[0138] The server sends the generated voice data to the device, which then plays it. The generated voice data is received as input. Specific operations include passing the voice data to a playback device and outputting the voice. The output is the user's voice being transmitted to the surroundings.
[0139] Step 11:
[0140] The server stores the generated text and audio data in a database. As input, it receives the generated text and audio data. As a specific operation, it performs a write operation to the database. As output, it obtains the stored data, which can be accessed later.
[0141] Step 12:
[0142] The server sends notifications to pre-defined contacts when certain keywords or events occur. As input, it monitors the generated text data. Specific operations include checking whether the keywords are included and, if so, creating a notification message. As output, a notification is generated and sent to the pre-defined contacts.
[0143] (Application example 1)
[0144] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0145] In autonomous vehicles, passengers are expected to be able to intuitively set their destinations and respond quickly in emergencies, but current systems are cumbersome to operate, making them inconvenient for people with physical disabilities and those unfamiliar with technology. Furthermore, there is a lack of more natural, human-like means of communication, and improvements are needed to enhance passenger convenience and safety.
[0146] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0147] In this invention, the server includes means for capturing a user's gaze using the device's camera, means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze, means for generating text data using the identified coordinates, means for inputting the generated text data into a speech synthesis algorithm to generate speech, means for providing an interface for operating the autonomous vehicle, and means for selecting a destination based on the identified coordinates. This allows passengers to easily and intuitively set their destination using their gaze and receive feedback via voice response. Furthermore, in the event of an emergency, notifications are quickly generated based on specific keywords and sent to pre-defined contacts, ensuring passenger safety.
[0148] A "device camera" is an optical device used to capture a user's gaze.
[0149] "Gaze data" is information indicating the direction in which the user's gaze is directed.
[0150] "Coordinates" are numerical data used to indicate a specific point, and are usually composed of X-axis and Y-axis values.
[0151] "Text data" is digital information expressed as a string of characters or sentences.
[0152] A "speech synthesis algorithm" is a program or calculation method for converting input text data into speech.
[0153] "Speech" refers to sound signals generated by speech synthesis algorithms that mimic human speech.
[0154] An "autonomous vehicle" is a vehicle that can drive autonomously without the need for human driving.
[0155] An "interface" is a means or tool for exchanging information between a user and a system.
[0156] "Destination selection" refers to the operation or process by which a user specifies the place they want to go.
[0157] A "notification" is a message or alert that informs a user or other interested party of specific information.
[0158] A system embodying this invention utilizes gaze recognition technology and speech synthesis technology to improve passenger convenience and safety in autonomous vehicles. The system includes a camera for capturing gaze data and hardware and software for analyzing the data.
[0159] The main components of the system are:
[0160] 1. Camera
[0161] The camera is installed inside the vehicle and captures the direction and focus of the passenger's gaze, allowing the user's intended actions to be visually detected.
[0162] 2. Gaze Recognition Algorithm
[0163] The gaze data captured by the camera is analyzed using gaze recognition algorithms, which use machine learning libraries such as TensorFlow.
[0164] 3. Data analysis and text generation
[0165] The coordinates obtained by the gaze recognition algorithm are mapped to specific characters or actions, and the coordinate data is sent to the device, which generates text corresponding to the specified action.
[0166] 4. Speech Synthesis Algorithm
[0167] The generated text data is converted to audio data using a speech synthesis algorithm (e.g., Google TTS API), which is then played back through the device's speaker to provide feedback to the passenger.
[0168] 5. Database and Notification System
[0169] The generated text and voice data is sent to a server and stored in a database. If a specific keyword (e.g., "help") is entered, a notification is sent to pre-defined contacts using the Twilio API.
[0170] Illustrative Usage Scenarios
[0171] Passengers can access the destination selection screen by directing their gaze towards the camera inside the autonomous vehicle. For example, if they want to go to "Shibuya Station," they focus their gaze on the "Shibuya Station" icon. This gaze action generates the text data "Shibuya Station," and a speech synthesis algorithm generates a voice saying, "Your destination is Shibuya Station, right?" The voice is played over the in-car speaker, and passengers are asked to confirm. At the same time, this information is stored on a server, and emergency contacts are notified if necessary.
[0172] Prompt Sentence Examples
[0173] Choose your destination by sight. Choose from the list below:
[0174] 1. Shibuya Station
[0175] 2. Tokyo Tower
[0176] 3. Nearby restaurants
[0177] 4. Library
[0178] Focus your gaze on your chosen destination and stop scrolling.
[0179] When implemented, these elements and algorithms will work seamlessly together to design a system that allows passengers to operate the system intuitively. Passengers will be able to operate the system with their eyes and receive feedback through voice, significantly improving convenience and safety.
[0180] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0181] Step 1:
[0182] Capture gaze data
[0183] The user looks at the camera in the autonomous vehicle. The camera captures gaze data at regular intervals, recording the direction and focus of the gaze. The captured image data is then input into the gaze recognition algorithm.
[0184] Input: Image data captured by the camera
[0185] Output: Gaze data
[0186] Step 2:
[0187] Preprocessing and transmission of gaze data
[0188] The device preprocesses the captured gaze data to remove noise, and then sends the preprocessed data to the server.
[0189] Input: Captured gaze data
[0190] Output: Preprocessed gaze data
[0191] Step 3:
[0192] Analysis of gaze data
[0193] The server receives the preprocessed gaze data and uses a gaze recognition algorithm (such as TensorFlow) to identify the coordinates where the gaze is pointing. This coordinate data is then sent from the server to the device.
[0194] Input: Preprocessed gaze data
[0195] Output: Identified coordinate data
[0196] Step 4:
[0197] Generating text data
[0198] The device uses the received coordinate data to generate text data: gaze is mapped to specific icons and characters, and a corresponding string of characters is generated.
[0199] Input: Identified coordinate data
[0200] Output: Text data
[0201] Step 5:
[0202] Generate audio data
[0203] The device sends the generated text data to the server, which then generates voice data using a speech synthesis algorithm (Google TTS API), which is then sent to the device.
[0204] Input: Text data
[0205] Output: Audio data
[0206] Step 6:
[0207] Playing audio
[0208] The terminal plays the received audio data through the speaker, and passengers receive audio feedback.
[0209] Input: Audio data
[0210] Output: Play audio
[0211] Step 7:
[0212] Data storage
[0213] The server stores the generated text data and voice data in a database.
[0214] Input: Text data, audio data
[0215] Output: Saved data
[0216] Step 8:
[0217] Sending emergency notifications
[0218] If a specific keyword (e.g., help) is entered, the server uses the Twilio API to send a notification to pre-defined contacts.
[0219] Input: Specific keyword
[0220] Output: Send emergency notification
[0221] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0222] The system of the present invention captures a user's gaze using a device's camera, analyzes the gaze data to identify the coordinates where the user is looking, and then combines it with an emotion engine that recognizes the user's emotions. It then generates text data based on this coordinate data and adjusts the tone and intonation of the generated voice based on the results of the emotion engine's estimation and analysis of the user's emotions. Finally, the generated text data is input into a speech synthesis algorithm to generate speech, allowing the user to speak. The system also includes functionality to store the generated text data and voice data in a database and generate notifications based on specific keywords and send them to pre-defined contacts.
[0223] Program processing explanation
[0224] Gaze recognition
[0225] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The captured gaze data is pre-processed to remove noise and generate clear data. This gaze data is sent to a server and analyzed using a machine learning algorithm. After analysis, the coordinates where the user is looking are identified and the coordinate data is sent to the device.
[0226] Examples:
[0227] If a user is using their gaze to operate a virtual keyboard in an application and is trying to input the word "thank you," the device will generate the input string "thank you" as the user moves their gaze to each letter.
[0228] Text data generation and emotion recognition
[0229] The server maps the identified coordinates to corresponding character information to generate text data in real time. At the same time, the device uses gaze data and other related data (e.g., facial expressions, eye movements) to estimate the user's emotions in real time. This estimation result is input into the emotion engine for detailed emotion analysis.
[0230] Examples:
[0231] If the user smiles while inputting "thank you" with their gaze, the emotion engine will analyze the gaze data and facial expression data to determine that the user is feeling happy.
[0232] Voice generation
[0233] The input text data and emotion recognition results are sent from the device to a server, where a speech synthesis algorithm generates a voice. The generated voice is customized using pre-stored samples of the user's voice, and the tone and intonation of the voice are adjusted according to the estimated emotion. For example, if the user is expressing joy, the voice will be generated in a brighter tone.
[0234] Examples:
[0235] If the user types "thank you" and the emotion engine identifies the user's emotion as "joy," the server will generate a voice that pronounces "thank you" in a bright and emotional voice, which the device can then play back and convey to those around it.
[0236] Data storage and notification function
[0237] The generated text and voice data is stored in a database on the server, and the server also sends notifications to pre-defined contacts when specific keywords or events occur, enabling rapid response in emergencies.
[0238] Examples:
[0239] If a user types "help," the server analyzes the text data and sends an email or SMS notification to contacts registered as an emergency message.
[0240] This system enables people with severe motor disabilities to communicate seamlessly and comfortably using gaze input and emotion recognition without the need for external sensors or specialized devices. The addition of a voice generation function also enables more natural and emotionally rich communication. Furthermore, data storage and notification functions ensure user safety and comfort.
[0241] The processing flow will be explained below.
[0242] Step 1:
[0243] When the application starts, the device activates the device's camera, which recognizes the user's face and eye position and begins capturing gaze data.
[0244] Step 2:
[0245] The device captures gaze data at regular intervals (e.g., every 100 milliseconds) and preprocesses it, which involves noise removal and data normalization.
[0246] Step 3:
[0247] The device transmits the captured gaze data, including the user's face position and gaze direction, to the server.
[0248] Step 4:
[0249] The server inputs the received gaze data into a machine learning algorithm for analysis, which identifies the coordinates where the user is looking.
[0250] Step 5:
[0251] The server transmits the identified coordinate data to the terminal.
[0252] Step 6:
[0253] Based on the received coordinate data, the device identifies the corresponding characters or icons and displays on the screen that the user has entered the input they intended.
[0254] Step 7:
[0255] The user operates the virtual keyboard with their eyes, selecting characters one after another. The selected characters are concatenated in real time to generate text data.
[0256] Step 8:
[0257] The terminal transmits the generated text data to the server.
[0258] Step 9:
[0259] The device collects gaze data and related information (e.g., facial expressions, eye movements) and sends it to the emotion engine.
[0260] Step 10:
[0261] The server uses an emotion engine to recognize emotions and sends the results to the device. The emotion engine analyzes the user's gaze data and facial expression data to estimate the user's emotions.
[0262] Step 11:
[0263] The server uses the text data and emotion recognition results to input the speech synthesis algorithm and generate an audio file, adjusting the tone and intonation of the voice depending on the user's emotion.
[0264] Step 12:
[0265] The server transmits the generated audio file to the terminal.
[0266] Step 13:
[0267] The terminal plays the received audio file, and the user's intention is spoken aloud.
[0268] Step 14:
[0269] The server stores all text and audio data in a database and manages it so that it can be referenced and analyzed at a later date.
[0270] Step 15:
[0271] Users or medical staff configure notification settings based on specific keywords or conditions, and when the conditions are met, the server automatically sends notifications to designated contacts.
[0272] Step 16:
[0273] In the event of an emergency, if a preset keyword (e.g., "help") is entered, the server will send an emergency message to pre-defined contacts, allowing for a prompt and appropriate response.
[0274] These processing steps realize a system that integrates gaze recognition, character input, emotion recognition, voice generation, and data management and notification functions.
[0275] Example 2
[0276] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0277] Conventional communication systems require external sensors or specialized equipment for people with severe motor disabilities, which poses problems such as high costs and the hassle of installation. Furthermore, they only provide simple input methods, making it difficult for users to express their emotions naturally. Furthermore, they are unable to respond quickly in emergencies, limiting the means to ensure user safety.
[0278] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's gaze using the device's camera, means for preprocessing the captured gaze data and removing noise, means for analyzing the preprocessed gaze data to identify coordinates corresponding to the gaze, means for generating text data using the identified coordinates, means for recognizing the user's emotion using the gaze data and related data, means for adjusting the tone and intonation of voice based on the generated text data and the emotion recognition results, and means for inputting the generated text data into a speech synthesis algorithm to generate speech. This enables even people with severe motor impairments to communicate naturally and emotionally without the need for external sensors or dedicated devices. Furthermore, the generated text data and voice data can be stored in a database, and notifications can be generated based on specific keywords and sent to pre-defined contacts, ensuring a rapid response in emergencies.
[0279] "Device camera" refers to a device with a camera function used to capture the user's line of sight.
[0280] "Gaze data" refers to data regarding the position and movement of the user's gaze captured by a camera.
[0281] "Preprocessing" refers to the process of applying noise removal and other processing to captured gaze data to generate clear data suitable for analysis.
[0282] "Noise reduction" refers to the process of removing external environmental influences and unnecessary information from captured raw data.
[0283] "Coordinate data" is data that indicates where on the screen the user's line of sight is located, derived by analyzing the line of sight data.
[0284] "Text data" refers to character information generated using gaze data.
[0285] "Emotion recognition" refers to a technology that estimates and classifies a user's emotional state by analyzing gaze data and related data.
[0286] An "emotion engine" refers to an algorithm or system that analyzes a user's gaze data, facial expression data, etc. to infer emotions.
[0287] "Speech synthesis algorithm" refers to a computational method or algorithm for converting text data into speech data.
[0288] "Tone and intonation" refers to the pitch and intonation of sounds in audio data.
[0289] "Database" refers to a digital warehouse for efficiently storing and managing generated text and audio data.
[0290] "Notifications" refer to alerts or messages generated based on specific keywords and sending information to pre-defined contacts.
[0291] The system of the present invention utilizes the device's camera to capture the user's gaze and analyzes the gaze data to determine the user's gaze coordinates. It also incorporates an emotion engine to recognize the user's emotions and adjust the tone and intonation of the voice based on the results. Finally, the generated text data is input into a speech synthesis algorithm to generate speech for the user to speak.
[0292] Gaze recognition
[0293] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The gaze data is mainly used to record the user's eye position and focus object. The captured gaze data is pre-processed in real time to remove noise. This pre-processing is necessary to minimize the influence of the external environment. The pre-processed gaze data is sent to a server and analyzed by a machine learning algorithm. This analysis identifies the coordinates where the user is looking, and the coordinate data is sent back to the device.
[0294] Text data generation and emotion recognition
[0295] The server maps the identified coordinates to corresponding character information and generates text data in real time. At the same time, the device uses gaze data and other related data (e.g., facial expressions and eye movements) to infer the user's emotions in real time. The emotion recognition results are input into the emotion engine for detailed emotion analysis.
[0296] Voice generation
[0297] The text data and emotion recognition results are sent from the device to a server, where a speech synthesis algorithm generates a voice. The generated voice is customized using pre-stored samples of the user's voice. The tone and intonation of the voice are also adjusted according to the estimated emotion. For example, if the user is expressing joy, the voice will be generated with a brighter tone.
[0298] Data storage and notification function
[0299] The generated text and voice data is stored in a database on the server. When a specific keyword or event occurs, the server sends a notification to pre-defined contacts. This notification function allows for a quick response in the event of an emergency.
[0300] Examples:
[0301] If a user is using their eyes to operate a virtual keyboard within an application and is trying to type "thank you," the device identifies each character by directing their gaze to each letter. As the user looks at each letter in "thank you," the device generates the string "thank you." If the user is smiling, the emotion engine recognizes the user's emotion as "joy" and generates a voice message saying "thank you" in a bright tone. If the user types "help," the server analyzes the text data as an emergency message and sends a notification to pre-defined contacts.
[0302] Example prompt:
[0303] "Please explain a new communication system that uses gaze input and emotion recognition, including specific usage scenarios and technical details."
[0304] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0305] Step 1:
[0306] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The camera records the user's eye position and focus and outputs the data as gaze data. Specifically, when the user looks at the display, the camera automatically tracks the movement of the pupils. This data is stored in the device's memory as gaze data.
[0307] Step 2:
[0308] The device performs preprocessing on the captured gaze data. Because gaze data contains noise, a noise reduction filter is applied to generate clear data. The input is raw gaze data, and the output is gaze data with noise removed. Specifically, the device filters out inaccurate data caused by camera resolution and external ambient light.
[0309] Step 3:
[0310] The device sends preprocessed gaze data to a server, which uses a machine learning algorithm to calculate the coordinate data on the screen corresponding to the gaze focus. The input is the preprocessed gaze data, and the output is coordinate data. Specifically, the device analyzes the gaze direction and position information to determine which part of the screen the user is looking at.
[0311] Step 4:
[0312] The server maps the coordinate data to specific character information and generates text data in real time. The input is coordinate data and the output is text data. In concrete terms, when a user tries to input "thank you" on the virtual keyboard with their gaze, the server generates text data by combining characters corresponding to those coordinates one after another.
[0313] Step 5:
[0314] The device uses gaze data and related data (e.g., facial expressions and eye movements) to estimate the user's emotions in real time. This data is input into an emotion engine for detailed emotion analysis. The input is gaze data and facial expression data, and the output is emotion recognition results. Specific operations include recognizing smiling and angry facial expressions and classifying the emotions as "joy" or "anger."
[0315] Step 6:
[0316] The device sends the generated text data and emotion recognition results to the server. The server generates speech using a speech synthesis algorithm. The input is the text data and emotion recognition results, and the output is speech data. In concrete terms, if the text "Thank you" is sent and the emotion recognition result is "joy," a voice saying "Thank you" in a bright tone is generated.
[0317] Step 7:
[0318] The server stores the generated text and voice data in a database. When a specific keyword or event occurs, a notification is sent to pre-defined contacts. The input is the keyword or event data, and the output is the notification message. Specifically, when a user types "help," the server analyzes the data and sends an emergency notification to the contacts.
[0319] Through each processing step, we have created a system that enables even people with severe motor disabilities to communicate in a natural and emotional way.
[0320] (Application example 2)
[0321] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0322] Conventional factory robot operation requires a dedicated remote controller and a complex user interface, resulting in problems such as reduced work efficiency and increased operator burden. Furthermore, the inability to recognize the emotional state of the operator during operation can lead to a lack of safety and efficiency in work. Furthermore, intuitive operation methods based on gaze and emotions have yet to be introduced, creating a demand for a more natural operating environment.
[0323] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's gaze using the device's camera, means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze, means for recognizing the user's emotion, means for generating text data based on the identified coordinates and the recognized emotion, means for inputting the generated text data into a speech synthesis algorithm and generating speech with a tone and intonation corresponding to the emotion, and means for processing and storing the gaze data and emotion data. This enables operators to operate factory robots in an intuitive and natural way using gaze and emotion, thereby improving operation efficiency and safety.
[0324] A "device camera" is a photographic device used to capture a user's gaze data.
[0325] "Gaze data" refers to information about the user's gaze direction and viewpoint position obtained through the device's camera.
[0326] "Coordinates" are position information in two-dimensional or three-dimensional space that is specified based on line-of-sight data.
[0327] "Emotion" refers to the psychological state that the user shows during the gaze operation.
[0328] "Text data" is character information generated based on gaze data and emotion data.
[0329] A "speech synthesis algorithm" is a program or process that converts input text data into speech.
[0330] "Emotional tone and intonation" refers to the timbre and intonation of a voice that is adjusted to suit the user's emotional state.
[0331] The "means for processing and storing gaze data and emotion data" is a function for analyzing and recording captured gaze data and recognized emotion data.
[0332] This invention is a system that captures a user's gaze data using a device's camera and generates text data and voice based on the gaze data and the user's emotion data. This system is realized mainly using the following hardware and software:
[0333] 1. Hardware:
[0334] Device camera: The camera used to capture user gaze data. Examples include a typical webcam or a camera built into a head-mounted display (HMD).
[0335] Head-mounted display (HMD): A device worn by the user that has a camera for eye tracking. An example of a suitable device is the Microsoft HoloLens 2.
[0336] 2. Software:
[0337] OpenCV: A library used to capture and preprocess gaze data.
[0338] dlib: A library for facial landmark detection.
[0339] pyttsx3: A speech synthesis algorithm that converts text data into speech.
[0340] emotion_recognition: A software module for recognizing user emotions.
[0341] gaze_tracking: A library for gaze tracking.
[0342] System configuration and processing flow:
[0343] The server captures the user's gaze data from the device's camera and analyzes it using OpenCV and gaze_tracking. By processing this gaze data, the coordinates of the gaze direction are identified. It also recognizes the user's emotion data in real time using dlib and emotion_recognition.
[0344] Based on the identified coordinates and the recognized emotion data, text data is generated. The generated text data is then fed into a speech synthesis algorithm using pyttsx3 to generate a customized voice with a tone and intonation that corresponds to the user's emotion. The gaze data and emotion data are then appropriately processed and stored for future reference and analysis.
[0345] Examples:
[0346] For example, imagine a user wearing an HMD in a factory and directing their gaze toward a specific machine part. At this time, gaze data is captured and it is determined that the gaze is directed toward a specific coordinate (the location of the machine part). At the same time, if the user's facial expression is recognized as "satisfied," the speech synthesis algorithm will cheerfully report, "The part has been placed in the correct position."
[0347] Example prompt for a generative AI model:
[0348] It recognizes the user's gaze and checks their position if they are looking at a specific work area. It estimates the user's emotions from their facial expressions and, if they are "satisfied," notifies them by voice that "the work is completed."
[0349] In this way, by using the system of the present invention, it becomes possible to operate a robot in a factory intuitively and efficiently.
[0350] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0351] Step 1:
[0352] The user wears the device's camera (e.g., a head-mounted display) and the system is activated. The device's camera begins capturing the user's gaze. This input data is captured as image frames.
[0353] Step 2:
[0354] The gaze data captured from the camera is pre-processed using OpenCV. This pre-processing step involves denoising the image and extracting the desired parts. The output is clean gaze data.
[0355] Step 3:
[0356] The preprocessed gaze data is analyzed using the gaze_tracking library. A specific algorithm is used to determine the gaze direction and the coordinates of the viewpoint. The input of this step is the preprocessed gaze data, and the output is the coordinate information of the gaze pointing.
[0357] Step 4:
[0358] In parallel, we use the dlib library to detect facial landmarks and recognize user emotions through the emotion_recognition module. The input is the captured raw image data, and the output is the user's emotional state (e.g., happy, anger, sadness).
[0359] Step 5:
[0360] Based on the identified coordinates and the recognized emotion data, text data is generated in real time. The server analyzes these input data and generates text that reflects the object the user is looking at and its meaning. The output is the generated text data.
[0361] Step 6:
[0362] The generated text data is converted into speech data using the speech synthesis algorithm pyttsx3. The tone and intonation of the speech are adjusted depending on the emotional state. The input is the generated text data and emotion recognition results, and the output is customized speech data.
[0363] Step 7:
[0364] The generated voice data is fed back to the user through the device, for example, a message such as "The part has been placed correctly" is played in a bright tone. In this step, the user receives the voice feedback and can take further action based on it.
[0365] Step 8:
[0366] The gaze data and emotion data are processed and stored appropriately on the server. When a specific key event occurs, a notification is sent to the configured contacts. The input of this step is gaze data and emotion data, and the output is saving to a database and generating a notification.
[0367] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0368] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0369] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0370] [Second embodiment]
[0371] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0372] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0373] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0374] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0375] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0376] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0377] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0378] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0379] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0380] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0381] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0382] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0383] The system of the present invention utilizes the device's camera to capture a user's gaze and analyzes the gaze data to determine the coordinates where the user is looking. It then generates text data based on the coordinates and inputs the generated text data into a speech synthesis algorithm to generate speech so that the user can speak. It also includes functionality to store the generated text and speech data in a database and generate notifications based on specific keywords and send them to pre-defined contacts.
[0384] Program processing explanation
[0385] Gaze recognition
[0386] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The captured gaze data is pre-processed to remove noise and generate clear data. This gaze data is sent to a server and analyzed using a machine learning algorithm. After analysis, the coordinates where the user is looking are identified and the coordinate data is sent to the device.
[0387] Examples:
[0388] If a user is using their gaze to operate a virtual keyboard in an application and is trying to input the word "thank you," the device will generate the input string "thank you" as the user moves their gaze to each letter.
[0389] Text input
[0390] The coordinates obtained as a result of analyzing the gaze data are mapped to specific characters or icons, and the corresponding characters are entered. The device displays the received character data on the screen and provides real-time feedback to the user, ensuring that the user enters exactly what they intended.
[0391] Examples:
[0392] As the user types "Thank you," the characters are displayed on the screen in real time, providing a clear view of the user's progress.
[0393] Voice generation
[0394] The entered text data is sent from the device to the server, where a speech synthesis algorithm generates speech. The generated speech is customized using pre-stored samples of the user's voice and sent to the device, where it is played back, allowing the user to speak.
[0395] Examples:
[0396] When the text data "Thank you" is generated, the server converts it into speech and uses a pre-stored sample of the user's voice to generate a speech that pronounces "Thank you," which the device can then play back and convey to people around it.
[0397] Data storage and notification function
[0398] The generated text and voice data is stored in a database on the server, and the server also sends notifications to pre-defined contacts when specific keywords or events occur, enabling rapid response in emergencies.
[0399] Examples:
[0400] If a user types "help," the server analyzes the text data and sends an email or SMS notification to contacts registered as an emergency message.
[0401] This system enables people with severe motor disabilities to communicate seamlessly and comfortably using gaze input and generative AI without the need for external sensors or specialized devices. The addition of a voice generation function also enables more natural and emotionally rich communication. Furthermore, data storage and notification functions ensure user safety and comfort.
[0402] The processing flow will be explained below.
[0403] Step 1:
[0404] When the application starts, the device activates the device's camera, which recognizes the user's face and eye position and begins capturing gaze data.
[0405] Step 2:
[0406] The device captures gaze data at regular intervals (e.g., every 100 milliseconds) and performs preprocessing, which involves removing noise and normalizing the data.
[0407] Step 3:
[0408] The device sends preprocessed gaze data, including the user's face position and gaze direction, to the server.
[0409] Step 4:
[0410] The server inputs the received gaze data into a machine learning algorithm for analysis, which identifies the coordinates where the user is looking.
[0411] Step 5:
[0412] The server transmits the identified coordinate data to the terminal.
[0413] Step 6:
[0414] Based on the received coordinate data, the device identifies the corresponding characters or icons and displays on the screen that the user has entered the input they intended.
[0415] Step 7:
[0416] The user operates the virtual keyboard with their eyes, selecting characters one after another. The selected characters are concatenated in real time to generate text data.
[0417] Step 8:
[0418] The terminal transmits the generated text data to the server.
[0419] Step 9:
[0420] The server feeds the text data into a speech synthesis algorithm to generate an audio file, which is customized using pre-stored samples of the user's voice.
[0421] Step 10:
[0422] The server transmits the generated audio file to the terminal.
[0423] Step 11:
[0424] The terminal plays the received audio file, and the user's intention is spoken aloud.
[0425] Step 12:
[0426] The server stores all text and audio data in a database and manages it so that it can be referenced and analyzed at a later date.
[0427] Step 13:
[0428] Users or medical staff configure notification settings based on specific keywords or conditions, and when the conditions are met, the server automatically sends notifications to designated contacts.
[0429] Step 14:
[0430] In the event of an emergency, if a preset keyword (e.g., "help") is entered, the server will send an emergency message to pre-defined contacts, allowing for a prompt and appropriate response.
[0431] These processing steps realize a system that integrates gaze recognition, character input, voice generation, data management and notification functions.
[0432] Example 1
[0433] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0434] This system solves the problem of people with severe motor disabilities finding it difficult to achieve seamless and comfortable communication using gaze input and generative AI without the need for external sensors or specialized devices. It also enables more natural and emotional communication by reliably inputting the user's intended content and enabling it to be spoken aloud, and it also solves the problem of a lack of means to send prompt notifications in emergencies.
[0435] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0436] In this invention, the server includes a means for analyzing gaze data and identifying the coordinates where the user is looking, a means for storing text data and voice data in a database, and a means for generating a notification based on a specific keyword and sending it to a pre-set contact, thereby making it possible to input characters based on the user's gaze, convert the input text data into voice, and send a notification under specific conditions based on the stored data.
[0437] "Device camera" refers to a camera used to capture a user's gaze, including cameras built into electronic devices such as smartphones, tablets, and computers.
[0438] "Gaze Data" refers to data about a user's eye movements and pupil position captured by a device's camera.
[0439] "Preprocessing" refers to the process of removing noise from captured gaze data and converting it into a format suitable for analysis.
[0440] The term "server" refers to a computer system that receives gaze data, analyzes it, and returns the necessary data to the terminal.
[0441] "Analysis" refers to the process of using machine learning algorithms to identify the coordinates where the user is looking from the gaze data.
[0442] "Coordinate data" refers to data indicating the position at which the user's gaze is pointing, as determined by analysis.
[0443] "Terminal" refers to an electronic device that has an interface with the user, accepts eye-gaze input, and receives data from a server.
[0444] "Text data" refers to character string data generated based on the user's line of sight.
[0445] "Speech synthesis algorithm" refers to the calculation procedures or programs used to convert text data into speech.
[0446] "Database" refers to an information system for systematically storing generated text and audio data.
[0447] "Specific keywords" refer to pre-defined important words or phrases contained in the text data entered by the user.
[0448] "Notifications" refer to messages or alerts sent to pre-defined contacts when certain keywords are detected.
[0449] The system of the present invention utilizes a device's camera to capture a user's gaze and analyzes the captured gaze data to determine the coordinates where the user is looking. This coordinate data is then used to generate text data, which is then input into a speech synthesis algorithm to generate speech. The system also stores the generated text and speech data, and generates notifications based on specific keywords and sends them to pre-defined contacts.
[0450] Gaze recognition
[0451] First, when the application is launched, the device starts the device camera and captures the user's gaze data at regular intervals. For this purpose, the camera needs to capture the position of the user's eyes at a high frame rate. For example, it is appropriate for the camera to capture images at 30 frames per second.
[0452] The device then preprocesses the captured gaze data to remove noise and generate clearer data, including high-frequency noise removal filters and grayscale conversion.
[0453] Data analysis
[0454] The preprocessed gaze data is encrypted and securely transmitted over the network to a server, where it is analyzed using machine learning algorithms to identify the coordinates where the user is looking. This analysis can potentially involve the application of deep learning models.
[0455] Text input
[0456] The analyzed coordinate data is sent to the device in real time, and characters and icons are identified based on that data. For example, when the user directs their gaze to each key on a virtual keyboard, the corresponding character is identified. Then, as the user moves their gaze, an input string is generated, and the device displays the string on the screen in real time.
[0457] Voice generation
[0458] The generated text data is sent from the device to a server, where it is converted into speech using a speech synthesis algorithm. The server uses pre-stored samples of the user's voice to generate a natural-sounding voice. This speech data is sent to the device and played back according to the user's settings.
[0459] Data Retention and Notification
[0460] The generated text and voice data is stored in a database on the server. In addition, the server sends notifications to pre-defined contacts when specific keywords or events occur, enabling a rapid response in the event of a user emergency.
[0461] Examples:
[0462] If a user is using their gaze to operate a virtual keyboard within an application and is trying to input the word "thank you," they can identify each character by directing their gaze at each letter. As the user looks at each letter in "thank you," the device generates the input string for "thank you." The generated text data is converted into speech by a connected speech synthesis algorithm, and the device plays it back. For example, the device can play back the audio "thank you" and convey the message to people around them.
[0463] Example prompt sentence:
[0464] "Please explain in natural language a program that uses the device's camera to capture the user's gaze, analyzes the gaze data, and identifies the coordinates where the user is looking. Please also explain the process of generating text data based on the coordinate data, and inputting that data into a speech synthesis algorithm to generate speech."
[0465] This system provides a simple and intuitive means of input using eye gaze control for people with severe motor disabilities, and supports richer communication by outputting the input as voice. It also ensures the user's safety by providing a notification function in case of an emergency.
[0466] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0467] Step 1:
[0468] When the application is launched, the device starts the device's camera and captures the user's gaze data at regular intervals. As input, it receives image data from the device's camera. The camera captures the user's eye position at 30 frames per second. As output, it generates captured gaze data. This gaze data includes the user's eye position information for each image frame.
[0469] Step 2:
[0470] The device preprocesses the captured gaze data to generate clear data with noise removed. It receives the captured gaze data as input, performs high-frequency noise removal filtering and grayscale conversion, and generates clear preprocessed gaze data as output.
[0471] Step 3:
[0472] The device sends the preprocessed gaze data to the server. As input, it receives preprocessed clear gaze data. The data is encrypted and securely transmitted over the network. As output, it generates the gaze data sent to the server.
[0473] Step 4:
[0474] The server analyzes the received gaze data using a machine learning algorithm to identify the coordinates where the user is looking. Preprocessed gaze data is received as input. A deep learning model is used as the machine learning algorithm. Specifically, it extracts features from the gaze data, inputs them into the model, and predicts the coordinates. As output, it generates coordinate data where the user is looking.
[0475] Step 5:
[0476] The server sends the parsed coordinate data to the device. As input, it receives the coordinate data of the user's view. The data is sent back to the device in real time. As output, it generates the coordinate data sent to the device.
[0477] Step 6:
[0478] The device identifies the character or icon the user is looking at based on the received coordinate data. As input, it receives the coordinate data the user is looking at. The coordinates are mapped to each key on the virtual keyboard, and when the gaze is directed, the character is identified. Specifically, it compares the coordinates with the key position on the virtual keyboard and identifies the matching character. As output, it generates the identified character data.
[0479] Step 7:
[0480] The terminal displays the identified character data on the screen in real time. As input, it receives the identified character data. As a specific action, it displays the character data in an appropriate position on the screen. As output, it provides visual feedback to the user.
[0481] Step 8:
[0482] The terminal generates text data based on the identified character data. As input, it receives a sequence of character data. As a specific operation, it concatenates the identified characters to create one piece of text data. As output, it obtains the generated text data.
[0483] Step 9:
[0484] The device sends the generated text data to the server and generates speech using a speech synthesis algorithm. The generated text data is received as input. The server generates natural-sounding speech using pre-stored samples of the user's voice. Specific operations include extracting phonemes from the text and generating a speech waveform. The generated speech data is obtained as output.
[0485] Step 10:
[0486] The server sends the generated voice data to the device, which then plays it. The generated voice data is received as input. Specific operations include passing the voice data to a playback device and outputting the voice. The output is the user's voice being transmitted to the surroundings.
[0487] Step 11:
[0488] The server stores the generated text and audio data in a database. As input, it receives the generated text and audio data. As a specific operation, it performs a write operation to the database. As output, it obtains the stored data, which can be accessed later.
[0489] Step 12:
[0490] The server sends notifications to pre-defined contacts when certain keywords or events occur. As input, it monitors the generated text data. Specific operations include checking whether the keywords are included and, if so, creating a notification message. As output, a notification is generated and sent to the pre-defined contacts.
[0491] (Application example 1)
[0492] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0493] In autonomous vehicles, passengers are expected to be able to intuitively set their destinations and respond quickly in emergencies, but current systems are cumbersome to operate, making them inconvenient for people with physical disabilities and those unfamiliar with technology. Furthermore, there is a lack of more natural, human-like means of communication, and improvements are needed to enhance passenger convenience and safety.
[0494] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0495] In this invention, the server includes means for capturing a user's gaze using the device's camera, means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze, means for generating text data using the identified coordinates, means for inputting the generated text data into a speech synthesis algorithm to generate speech, means for providing an interface for operating the autonomous vehicle, and means for selecting a destination based on the identified coordinates. This allows passengers to easily and intuitively set their destination using their gaze and receive feedback via voice response. Furthermore, in the event of an emergency, notifications are quickly generated based on specific keywords and sent to pre-defined contacts, ensuring passenger safety.
[0496] A "device camera" is an optical device used to capture a user's gaze.
[0497] "Gaze data" is information indicating the direction in which the user's gaze is directed.
[0498] "Coordinates" are numerical data used to indicate a specific point, and are usually composed of X-axis and Y-axis values.
[0499] "Text data" is digital information expressed as a string of characters or sentences.
[0500] A "speech synthesis algorithm" is a program or calculation method for converting input text data into speech.
[0501] "Speech" refers to sound signals generated by speech synthesis algorithms that mimic human speech.
[0502] An "autonomous vehicle" is a vehicle that can drive autonomously without the need for human driving.
[0503] An "interface" is a means or tool for exchanging information between a user and a system.
[0504] "Destination selection" refers to the operation or process by which a user specifies the place they want to go.
[0505] A "notification" is a message or alert that informs a user or other interested party of specific information.
[0506] A system embodying this invention utilizes gaze recognition technology and speech synthesis technology to improve passenger convenience and safety in autonomous vehicles. The system includes a camera for capturing gaze data and hardware and software for analyzing the data.
[0507] The main components of the system are:
[0508] 1. Camera
[0509] The camera is installed inside the vehicle and captures the direction and focus of the passenger's gaze, allowing the user's intended actions to be visually detected.
[0510] 2. Gaze Recognition Algorithm
[0511] The gaze data captured by the camera is analyzed using gaze recognition algorithms, which use machine learning libraries such as TensorFlow.
[0512] 3. Data analysis and text generation
[0513] The coordinates obtained by the gaze recognition algorithm are mapped to specific characters or actions, and the coordinate data is sent to the device, which generates text corresponding to the specified action.
[0514] 4. Speech Synthesis Algorithm
[0515] The generated text data is converted to audio data using a speech synthesis algorithm (e.g., Google TTS API), which is then played back through the device's speaker to provide feedback to the passenger.
[0516] 5. Database and Notification System
[0517] The generated text and voice data is sent to a server and stored in a database. If a specific keyword (e.g., "help") is entered, a notification is sent to pre-defined contacts using the Twilio API.
[0518] Illustrative Usage Scenarios
[0519] Passengers can access the destination selection screen by directing their gaze towards the camera inside the autonomous vehicle. For example, if they want to go to "Shibuya Station," they focus their gaze on the "Shibuya Station" icon. This gaze action generates the text data "Shibuya Station," and a speech synthesis algorithm generates a voice saying, "Your destination is Shibuya Station, right?" The voice is played over the in-car speaker, and passengers are asked to confirm. At the same time, this information is stored on a server, and emergency contacts are notified if necessary.
[0520] Prompt Sentence Examples
[0521] Choose your destination by sight. Choose from the list below:
[0522] 1. Shibuya Station
[0523] 2. Tokyo Tower
[0524] 3. Nearby restaurants
[0525] 4. Library
[0526] Focus your gaze on your chosen destination and stop scrolling.
[0527] When implemented, these elements and algorithms will work seamlessly together to design a system that allows passengers to operate the system intuitively. Passengers will be able to operate the system with their eyes and receive feedback through voice, significantly improving convenience and safety.
[0528] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0529] Step 1:
[0530] Capture gaze data
[0531] The user looks at the camera in the autonomous vehicle. The camera captures gaze data at regular intervals, recording the direction and focus of the gaze. The captured image data is then input into the gaze recognition algorithm.
[0532] Input: Image data captured by the camera
[0533] Output: Gaze data
[0534] Step 2:
[0535] Preprocessing and transmission of gaze data
[0536] The device preprocesses the captured gaze data to remove noise, and then sends the preprocessed data to the server.
[0537] Input: Captured gaze data
[0538] Output: Preprocessed gaze data
[0539] Step 3:
[0540] Analysis of gaze data
[0541] The server receives the preprocessed gaze data and uses a gaze recognition algorithm (such as TensorFlow) to identify the coordinates where the gaze is pointing. This coordinate data is then sent from the server to the device.
[0542] Input: Preprocessed gaze data
[0543] Output: Identified coordinate data
[0544] Step 4:
[0545] Generating text data
[0546] The device uses the received coordinate data to generate text data: gaze is mapped to specific icons and characters, and a corresponding string of characters is generated.
[0547] Input: Identified coordinate data
[0548] Output: Text data
[0549] Step 5:
[0550] Generate audio data
[0551] The device sends the generated text data to the server, which then generates voice data using a speech synthesis algorithm (Google TTS API), which is then sent to the device.
[0552] Input: Text data
[0553] Output: Audio data
[0554] Step 6:
[0555] Playing audio
[0556] The terminal plays the received audio data through the speaker, and passengers receive audio feedback.
[0557] Input: Audio data
[0558] Output: Play audio
[0559] Step 7:
[0560] Data storage
[0561] The server stores the generated text data and voice data in a database.
[0562] Input: Text data, audio data
[0563] Output: Saved data
[0564] Step 8:
[0565] Sending emergency notifications
[0566] If a specific keyword (e.g., help) is entered, the server uses the Twilio API to send a notification to pre-defined contacts.
[0567] Input: Specific keyword
[0568] Output: Send emergency notification
[0569] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0570] The system of the present invention captures a user's gaze using a device's camera, analyzes the gaze data to identify the coordinates where the user is looking, and then combines it with an emotion engine that recognizes the user's emotions. It then generates text data based on this coordinate data and adjusts the tone and intonation of the generated voice based on the results of the emotion engine's estimation and analysis of the user's emotions. Finally, the generated text data is input into a speech synthesis algorithm to generate speech, allowing the user to speak. The system also includes functionality to store the generated text data and voice data in a database and generate notifications based on specific keywords and send them to pre-defined contacts.
[0571] Program processing explanation
[0572] Gaze recognition
[0573] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The captured gaze data is pre-processed to remove noise and generate clear data. This gaze data is sent to a server and analyzed using a machine learning algorithm. After analysis, the coordinates where the user is looking are identified and the coordinate data is sent to the device.
[0574] Examples:
[0575] If a user is using their gaze to operate a virtual keyboard in an application and is trying to input the word "thank you," the device will generate the input string "thank you" as the user moves their gaze to each letter.
[0576] Text data generation and emotion recognition
[0577] The server maps the identified coordinates to corresponding character information to generate text data in real time. At the same time, the device uses gaze data and other related data (e.g., facial expressions, eye movements) to estimate the user's emotions in real time. This estimation result is input into the emotion engine for detailed emotion analysis.
[0578] Examples:
[0579] If the user smiles while inputting "thank you" with their gaze, the emotion engine will analyze the gaze data and facial expression data to determine that the user is feeling happy.
[0580] Voice generation
[0581] The input text data and emotion recognition results are sent from the device to a server, where a speech synthesis algorithm generates a voice. The generated voice is customized using pre-stored samples of the user's voice, and the tone and intonation of the voice are adjusted according to the estimated emotion. For example, if the user is expressing joy, the voice will be generated in a brighter tone.
[0582] Examples:
[0583] If the user types "thank you" and the emotion engine identifies the user's emotion as "joy," the server will generate a voice that pronounces "thank you" in a bright and emotional voice, which the device can then play back and convey to those around it.
[0584] Data storage and notification function
[0585] The generated text and voice data is stored in a database on the server, and the server also sends notifications to pre-defined contacts when specific keywords or events occur, enabling rapid response in emergencies.
[0586] Examples:
[0587] If a user types "help," the server analyzes the text data and sends an email or SMS notification to contacts registered as an emergency message.
[0588] This system enables people with severe motor disabilities to communicate seamlessly and comfortably using gaze input and emotion recognition without the need for external sensors or specialized devices. The addition of a voice generation function also enables more natural and emotionally rich communication. Furthermore, data storage and notification functions ensure user safety and comfort.
[0589] The processing flow will be explained below.
[0590] Step 1:
[0591] When the application starts, the device activates the device's camera, which recognizes the user's face and eye position and begins capturing gaze data.
[0592] Step 2:
[0593] The device captures gaze data at regular intervals (e.g., every 100 milliseconds) and preprocesses it, which involves noise removal and data normalization.
[0594] Step 3:
[0595] The device transmits the captured gaze data, including the user's face position and gaze direction, to the server.
[0596] Step 4:
[0597] The server inputs the received gaze data into a machine learning algorithm for analysis, which identifies the coordinates where the user is looking.
[0598] Step 5:
[0599] The server transmits the identified coordinate data to the terminal.
[0600] Step 6:
[0601] Based on the received coordinate data, the device identifies the corresponding characters or icons and displays on the screen that the user has entered the input they intended.
[0602] Step 7:
[0603] The user operates the virtual keyboard with their eyes, selecting characters one after another. The selected characters are concatenated in real time to generate text data.
[0604] Step 8:
[0605] The terminal transmits the generated text data to the server.
[0606] Step 9:
[0607] The device collects gaze data and related information (e.g., facial expressions, eye movements) and sends it to the emotion engine.
[0608] Step 10:
[0609] The server uses an emotion engine to recognize emotions and sends the results to the device. The emotion engine analyzes the user's gaze data and facial expression data to estimate the user's emotions.
[0610] Step 11:
[0611] The server uses the text data and emotion recognition results to input the speech synthesis algorithm and generate an audio file, adjusting the tone and intonation of the voice depending on the user's emotion.
[0612] Step 12:
[0613] The server transmits the generated audio file to the terminal.
[0614] Step 13:
[0615] The terminal plays the received audio file, and the user's intention is spoken aloud.
[0616] Step 14:
[0617] The server stores all text and audio data in a database and manages it so that it can be referenced and analyzed at a later date.
[0618] Step 15:
[0619] Users or medical staff configure notification settings based on specific keywords or conditions, and when the conditions are met, the server automatically sends notifications to designated contacts.
[0620] Step 16:
[0621] In the event of an emergency, if a preset keyword (e.g., "help") is entered, the server will send an emergency message to pre-defined contacts, allowing for a prompt and appropriate response.
[0622] These processing steps realize a system that integrates gaze recognition, character input, emotion recognition, voice generation, and data management and notification functions.
[0623] Example 2
[0624] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0625] Conventional communication systems require external sensors or specialized equipment for people with severe motor disabilities, which poses problems such as high costs and the hassle of installation. Furthermore, they only provide simple input methods, making it difficult for users to express their emotions naturally. Furthermore, they are unable to respond quickly in emergencies, limiting the means to ensure user safety.
[0626] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's gaze using the device's camera, means for preprocessing the captured gaze data and removing noise, means for analyzing the preprocessed gaze data to identify coordinates corresponding to the gaze, means for generating text data using the identified coordinates, means for recognizing the user's emotion using the gaze data and related data, means for adjusting the tone and intonation of voice based on the generated text data and the emotion recognition results, and means for inputting the generated text data into a speech synthesis algorithm to generate speech. This enables even people with severe motor impairments to communicate naturally and emotionally without the need for external sensors or dedicated devices. Furthermore, the generated text data and voice data can be stored in a database, and notifications can be generated based on specific keywords and sent to pre-defined contacts, ensuring a rapid response in emergencies.
[0627] "Device camera" refers to a device with a camera function used to capture the user's line of sight.
[0628] "Gaze data" refers to data regarding the position and movement of the user's gaze captured by a camera.
[0629] "Preprocessing" refers to the process of applying noise removal and other processing to captured gaze data to generate clear data suitable for analysis.
[0630] "Noise reduction" refers to the process of removing external environmental influences and unnecessary information from captured raw data.
[0631] "Coordinate data" is data that indicates where on the screen the user's line of sight is located, derived by analyzing the line of sight data.
[0632] "Text data" refers to character information generated using gaze data.
[0633] "Emotion recognition" refers to a technology that estimates and classifies a user's emotional state by analyzing gaze data and related data.
[0634] An "emotion engine" refers to an algorithm or system that analyzes a user's gaze data, facial expression data, etc. to infer emotions.
[0635] "Speech synthesis algorithm" refers to a computational method or algorithm for converting text data into speech data.
[0636] "Tone and intonation" refers to the pitch and intonation of sounds in audio data.
[0637] "Database" refers to a digital warehouse for efficiently storing and managing generated text and audio data.
[0638] "Notifications" refer to alerts or messages generated based on specific keywords and sending information to pre-defined contacts.
[0639] The system of the present invention utilizes the device's camera to capture the user's gaze and analyzes the gaze data to determine the user's gaze coordinates. It also incorporates an emotion engine to recognize the user's emotions and adjust the tone and intonation of the voice based on the results. Finally, the generated text data is input into a speech synthesis algorithm to generate speech for the user to speak.
[0640] Gaze recognition
[0641] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The gaze data is mainly used to record the user's eye position and focus object. The captured gaze data is pre-processed in real time to remove noise. This pre-processing is necessary to minimize the influence of the external environment. The pre-processed gaze data is sent to a server and analyzed by a machine learning algorithm. This analysis identifies the coordinates where the user is looking, and the coordinate data is sent back to the device.
[0642] Text data generation and emotion recognition
[0643] The server maps the identified coordinates to corresponding character information and generates text data in real time. At the same time, the device uses gaze data and other related data (e.g., facial expressions and eye movements) to infer the user's emotions in real time. The emotion recognition results are input into the emotion engine for detailed emotion analysis.
[0644] Voice generation
[0645] The text data and emotion recognition results are sent from the device to a server, where a speech synthesis algorithm generates a voice. The generated voice is customized using pre-stored samples of the user's voice. The tone and intonation of the voice are also adjusted according to the estimated emotion. For example, if the user is expressing joy, the voice will be generated with a brighter tone.
[0646] Data storage and notification function
[0647] The generated text and voice data is stored in a database on the server. When a specific keyword or event occurs, the server sends a notification to pre-defined contacts. This notification function allows for a quick response in the event of an emergency.
[0648] Examples:
[0649] If a user is using their eyes to operate a virtual keyboard within an application and is trying to type "thank you," the device identifies each character by directing their gaze to each letter. As the user looks at each letter in "thank you," the device generates the string "thank you." If the user is smiling, the emotion engine recognizes the user's emotion as "joy" and generates a voice message saying "thank you" in a bright tone. If the user types "help," the server analyzes the text data as an emergency message and sends a notification to pre-defined contacts.
[0650] Example prompt:
[0651] "Please explain a new communication system that uses gaze input and emotion recognition, including specific usage scenarios and technical details."
[0652] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0653] Step 1:
[0654] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The camera records the user's eye position and focus and outputs the data as gaze data. Specifically, when the user looks at the display, the camera automatically tracks the movement of the pupils. This data is stored in the device's memory as gaze data.
[0655] Step 2:
[0656] The device performs preprocessing on the captured gaze data. Because gaze data contains noise, a noise reduction filter is applied to generate clear data. The input is raw gaze data, and the output is gaze data with noise removed. Specifically, the device filters out inaccurate data caused by camera resolution and external ambient light.
[0657] Step 3:
[0658] The device sends preprocessed gaze data to a server, which uses a machine learning algorithm to calculate the coordinate data on the screen corresponding to the gaze focus. The input is the preprocessed gaze data, and the output is coordinate data. Specifically, the device analyzes the gaze direction and position information to determine which part of the screen the user is looking at.
[0659] Step 4:
[0660] The server maps the coordinate data to specific character information and generates text data in real time. The input is coordinate data and the output is text data. In concrete terms, when a user tries to input "thank you" on the virtual keyboard with their gaze, the server generates text data by combining characters corresponding to those coordinates one after another.
[0661] Step 5:
[0662] The device uses gaze data and related data (e.g., facial expressions and eye movements) to estimate the user's emotions in real time. This data is input into an emotion engine for detailed emotion analysis. The input is gaze data and facial expression data, and the output is emotion recognition results. Specific operations include recognizing smiling and angry facial expressions and classifying the emotions as "joy" or "anger."
[0663] Step 6:
[0664] The device sends the generated text data and emotion recognition results to the server. The server generates speech using a speech synthesis algorithm. The input is the text data and emotion recognition results, and the output is speech data. In concrete terms, if the text "Thank you" is sent and the emotion recognition result is "joy," a voice saying "Thank you" in a bright tone is generated.
[0665] Step 7:
[0666] The server stores the generated text and voice data in a database. When a specific keyword or event occurs, a notification is sent to pre-defined contacts. The input is the keyword or event data, and the output is the notification message. Specifically, when a user types "help," the server analyzes the data and sends an emergency notification to the contacts.
[0667] Through each processing step, we have created a system that enables even people with severe motor disabilities to communicate in a natural and emotional way.
[0668] (Application example 2)
[0669] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0670] Conventional factory robot operation requires a dedicated remote controller and a complex user interface, resulting in problems such as reduced work efficiency and increased operator burden. Furthermore, the inability to recognize the emotional state of the operator during operation can lead to a lack of safety and efficiency in work. Furthermore, intuitive operation methods based on gaze and emotions have yet to be introduced, creating a demand for a more natural operating environment.
[0671] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's gaze using the device's camera, means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze, means for recognizing the user's emotion, means for generating text data based on the identified coordinates and the recognized emotion, means for inputting the generated text data into a speech synthesis algorithm and generating speech with a tone and intonation corresponding to the emotion, and means for processing and storing the gaze data and emotion data. This enables operators to operate factory robots in an intuitive and natural way using gaze and emotion, thereby improving operation efficiency and safety.
[0672] A "device camera" is a photographic device used to capture a user's gaze data.
[0673] "Gaze data" refers to information about the user's gaze direction and viewpoint position obtained through the device's camera.
[0674] "Coordinates" are position information in two-dimensional or three-dimensional space that is specified based on line-of-sight data.
[0675] "Emotion" refers to the psychological state that the user shows during the gaze operation.
[0676] "Text data" is character information generated based on gaze data and emotion data.
[0677] A "speech synthesis algorithm" is a program or process that converts input text data into speech.
[0678] "Emotional tone and intonation" refers to the timbre and intonation of a voice that is adjusted to suit the user's emotional state.
[0679] The "means for processing and storing gaze data and emotion data" is a function for analyzing and recording captured gaze data and recognized emotion data.
[0680] This invention is a system that captures a user's gaze data using a device's camera and generates text data and voice based on the gaze data and the user's emotion data. This system is realized mainly using the following hardware and software:
[0681] 1. Hardware:
[0682] Device camera: The camera used to capture user gaze data. Examples include a typical webcam or a camera built into a head-mounted display (HMD).
[0683] Head-mounted display (HMD): A device worn by the user that has a camera for eye tracking. An example of a suitable device is the Microsoft HoloLens 2.
[0684] 2. Software:
[0685] OpenCV: A library used to capture and preprocess gaze data.
[0686] dlib: A library for facial landmark detection.
[0687] pyttsx3: A speech synthesis algorithm that converts text data into speech.
[0688] emotion_recognition: A software module for recognizing user emotions.
[0689] gaze_tracking: A library for gaze tracking.
[0690] System configuration and processing flow:
[0691] The server captures the user's gaze data from the device's camera and analyzes it using OpenCV and gaze_tracking. By processing this gaze data, the coordinates of the gaze direction are identified. It also recognizes the user's emotion data in real time using dlib and emotion_recognition.
[0692] Based on the identified coordinates and the recognized emotion data, text data is generated. The generated text data is then fed into a speech synthesis algorithm using pyttsx3 to generate a customized voice with a tone and intonation that corresponds to the user's emotion. The gaze data and emotion data are then appropriately processed and stored for future reference and analysis.
[0693] Examples:
[0694] For example, imagine a user wearing an HMD in a factory and directing their gaze toward a specific machine part. At this time, gaze data is captured and it is determined that the gaze is directed toward a specific coordinate (the location of the machine part). At the same time, if the user's facial expression is recognized as "satisfied," the speech synthesis algorithm will cheerfully report, "The part has been placed in the correct position."
[0695] Example prompt for a generative AI model:
[0696] It recognizes the user's gaze and checks their position if they are looking at a specific work area. It estimates the user's emotions from their facial expressions and, if they are "satisfied," notifies them by voice that "the work is completed."
[0697] In this way, by using the system of the present invention, it becomes possible to operate a robot in a factory intuitively and efficiently.
[0698] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0699] Step 1:
[0700] The user wears the device's camera (e.g., a head-mounted display) and the system is activated. The device's camera begins capturing the user's gaze. This input data is captured as image frames.
[0701] Step 2:
[0702] The gaze data captured from the camera is pre-processed using OpenCV. This pre-processing step involves denoising the image and extracting the desired parts. The output is clean gaze data.
[0703] Step 3:
[0704] The preprocessed gaze data is analyzed using the gaze_tracking library. A specific algorithm is used to determine the gaze direction and the coordinates of the viewpoint. The input of this step is the preprocessed gaze data, and the output is the coordinate information of the gaze pointing.
[0705] Step 4:
[0706] In parallel, we use the dlib library to detect facial landmarks and recognize user emotions through the emotion_recognition module. The input is the captured raw image data, and the output is the user's emotional state (e.g., happy, anger, sadness).
[0707] Step 5:
[0708] Based on the identified coordinates and the recognized emotion data, text data is generated in real time. The server analyzes these input data and generates text that reflects the object the user is looking at and its meaning. The output is the generated text data.
[0709] Step 6:
[0710] The generated text data is converted into speech data using the speech synthesis algorithm pyttsx3. The tone and intonation of the speech are adjusted depending on the emotional state. The input is the generated text data and emotion recognition results, and the output is customized speech data.
[0711] Step 7:
[0712] The generated voice data is fed back to the user through the device, for example, a message such as "The part has been placed correctly" is played in a bright tone. In this step, the user receives the voice feedback and can take further action based on it.
[0713] Step 8:
[0714] The gaze data and emotion data are processed and stored appropriately on the server. When a specific key event occurs, a notification is sent to the configured contacts. The input of this step is gaze data and emotion data, and the output is saving to a database and generating a notification.
[0715] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0716] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0717] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0718] [Third embodiment]
[0719] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0720] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0721] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0722] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0723] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0724] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0725] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0726] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0727] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0728] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0729] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0730] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0731] The system of the present invention utilizes the device's camera to capture a user's gaze and analyzes the gaze data to determine the coordinates where the user is looking. It then generates text data based on the coordinates and inputs the generated text data into a speech synthesis algorithm to generate speech so that the user can speak. It also includes functionality to store the generated text and speech data in a database and generate notifications based on specific keywords and send them to pre-defined contacts.
[0732] Program processing explanation
[0733] Gaze recognition
[0734] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The captured gaze data is pre-processed to remove noise and generate clear data. This gaze data is sent to a server and analyzed using a machine learning algorithm. After analysis, the coordinates where the user is looking are identified and the coordinate data is sent to the device.
[0735] Examples:
[0736] If a user is using their gaze to operate a virtual keyboard in an application and is trying to input the word "thank you," the device will generate the input string "thank you" as the user moves their gaze to each letter.
[0737] Text input
[0738] The coordinates obtained as a result of analyzing the gaze data are mapped to specific characters or icons, and the corresponding characters are entered. The device displays the received character data on the screen and provides real-time feedback to the user, ensuring that the user enters exactly what they intended.
[0739] Examples:
[0740] As the user types "Thank you," the characters are displayed on the screen in real time, providing a clear view of the user's progress.
[0741] Voice generation
[0742] The entered text data is sent from the device to the server, where a speech synthesis algorithm generates speech. The generated speech is customized using pre-stored samples of the user's voice and sent to the device, where it is played back, allowing the user to speak.
[0743] Examples:
[0744] When the text data "Thank you" is generated, the server converts it into speech and uses a pre-stored sample of the user's voice to generate a speech that pronounces "Thank you," which the device can then play back and convey to people around it.
[0745] Data storage and notification function
[0746] The generated text and voice data is stored in a database on the server, and the server also sends notifications to pre-defined contacts when specific keywords or events occur, enabling rapid response in emergencies.
[0747] Examples:
[0748] If a user types "help," the server analyzes the text data and sends an email or SMS notification to contacts registered as an emergency message.
[0749] This system enables people with severe motor disabilities to communicate seamlessly and comfortably using gaze input and generative AI without the need for external sensors or specialized devices. The addition of a voice generation function also enables more natural and emotionally rich communication. Furthermore, data storage and notification functions ensure user safety and comfort.
[0750] The processing flow will be explained below.
[0751] Step 1:
[0752] When the application starts, the device activates the device's camera, which recognizes the user's face and eye position and begins capturing gaze data.
[0753] Step 2:
[0754] The device captures gaze data at regular intervals (e.g., every 100 milliseconds) and performs preprocessing, which involves removing noise and normalizing the data.
[0755] Step 3:
[0756] The device sends preprocessed gaze data, including the user's face position and gaze direction, to the server.
[0757] Step 4:
[0758] The server inputs the received gaze data into a machine learning algorithm for analysis, which identifies the coordinates where the user is looking.
[0759] Step 5:
[0760] The server transmits the identified coordinate data to the terminal.
[0761] Step 6:
[0762] Based on the received coordinate data, the device identifies the corresponding characters or icons and displays on the screen that the user has entered the input they intended.
[0763] Step 7:
[0764] The user operates the virtual keyboard with their eyes, selecting characters one after another. The selected characters are concatenated in real time to generate text data.
[0765] Step 8:
[0766] The terminal transmits the generated text data to the server.
[0767] Step 9:
[0768] The server feeds the text data into a speech synthesis algorithm to generate an audio file, which is customized using pre-stored samples of the user's voice.
[0769] Step 10:
[0770] The server transmits the generated audio file to the terminal.
[0771] Step 11:
[0772] The terminal plays the received audio file, and the user's intention is spoken aloud.
[0773] Step 12:
[0774] The server stores all text and audio data in a database and manages it so that it can be referenced and analyzed at a later date.
[0775] Step 13:
[0776] Users or medical staff configure notification settings based on specific keywords or conditions, and when the conditions are met, the server automatically sends notifications to designated contacts.
[0777] Step 14:
[0778] In the event of an emergency, if a preset keyword (e.g., "help") is entered, the server will send an emergency message to pre-defined contacts, allowing for a prompt and appropriate response.
[0779] These processing steps realize a system that integrates gaze recognition, character input, voice generation, data management and notification functions.
[0780] Example 1
[0781] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0782] This system solves the problem of people with severe motor disabilities finding it difficult to achieve seamless and comfortable communication using gaze input and generative AI without the need for external sensors or specialized devices. It also enables more natural and emotional communication by reliably inputting the user's intended content and enabling it to be spoken aloud, and it also solves the problem of a lack of means to send prompt notifications in emergencies.
[0783] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0784] In this invention, the server includes a means for analyzing gaze data and identifying the coordinates where the user is looking, a means for storing text data and voice data in a database, and a means for generating a notification based on a specific keyword and sending it to a pre-set contact, thereby making it possible to input characters based on the user's gaze, convert the input text data into voice, and send a notification under specific conditions based on the stored data.
[0785] "Device camera" refers to a camera used to capture a user's gaze, including cameras built into electronic devices such as smartphones, tablets, and computers.
[0786] "Gaze Data" refers to data about a user's eye movements and pupil position captured by a device's camera.
[0787] "Preprocessing" refers to the process of removing noise from captured gaze data and converting it into a format suitable for analysis.
[0788] The term "server" refers to a computer system that receives gaze data, analyzes it, and returns the necessary data to the terminal.
[0789] "Analysis" refers to the process of using machine learning algorithms to identify the coordinates where the user is looking from the gaze data.
[0790] "Coordinate data" refers to data indicating the position at which the user's gaze is pointing, as determined by analysis.
[0791] "Terminal" refers to an electronic device that has an interface with the user, accepts eye-gaze input, and receives data from a server.
[0792] "Text data" refers to character string data generated based on the user's line of sight.
[0793] "Speech synthesis algorithm" refers to the calculation procedures or programs used to convert text data into speech.
[0794] "Database" refers to an information system for systematically storing generated text and audio data.
[0795] "Specific keywords" refer to pre-defined important words or phrases contained in the text data entered by the user.
[0796] "Notifications" refer to messages or alerts sent to pre-defined contacts when certain keywords are detected.
[0797] The system of the present invention utilizes a device's camera to capture a user's gaze and analyzes the captured gaze data to determine the coordinates where the user is looking. This coordinate data is then used to generate text data, which is then input into a speech synthesis algorithm to generate speech. The system also stores the generated text and speech data, and generates notifications based on specific keywords and sends them to pre-defined contacts.
[0798] Gaze recognition
[0799] First, when the application is launched, the device starts the device camera and captures the user's gaze data at regular intervals. For this purpose, the camera needs to capture the position of the user's eyes at a high frame rate. For example, it is appropriate for the camera to capture images at 30 frames per second.
[0800] The device then preprocesses the captured gaze data to remove noise and generate clearer data, including high-frequency noise removal filters and grayscale conversion.
[0801] Data analysis
[0802] The preprocessed gaze data is encrypted and securely transmitted over the network to a server, where it is analyzed using machine learning algorithms to identify the coordinates where the user is looking. This analysis can potentially involve the application of deep learning models.
[0803] Text input
[0804] The analyzed coordinate data is sent to the device in real time, and characters and icons are identified based on that data. For example, when the user directs their gaze to each key on a virtual keyboard, the corresponding character is identified. Then, as the user moves their gaze, an input string is generated, and the device displays the string on the screen in real time.
[0805] Voice generation
[0806] The generated text data is sent from the device to a server, where it is converted into speech using a speech synthesis algorithm. The server uses pre-stored samples of the user's voice to generate a natural-sounding voice. This speech data is sent to the device and played back according to the user's settings.
[0807] Data Retention and Notification
[0808] The generated text and voice data is stored in a database on the server. In addition, the server sends notifications to pre-defined contacts when specific keywords or events occur, enabling a rapid response in the event of a user emergency.
[0809] Examples:
[0810] If a user is using their gaze to operate a virtual keyboard within an application and is trying to input the word "thank you," they can identify each character by directing their gaze at each letter. As the user looks at each letter in "thank you," the device generates the input string for "thank you." The generated text data is converted into speech by a connected speech synthesis algorithm, and the device plays it back. For example, the device can play back the audio "thank you" and convey the message to people around them.
[0811] Example prompt sentence:
[0812] "Please explain in natural language a program that uses the device's camera to capture the user's gaze, analyzes the gaze data, and identifies the coordinates where the user is looking. Please also explain the process of generating text data based on the coordinate data, and inputting that data into a speech synthesis algorithm to generate speech."
[0813] This system provides a simple and intuitive means of input using eye gaze control for people with severe motor disabilities, and supports richer communication by outputting the input as voice. It also ensures the user's safety by providing a notification function in case of an emergency.
[0814] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0815] Step 1:
[0816] When the application is launched, the device starts the device's camera and captures the user's gaze data at regular intervals. As input, it receives image data from the device's camera. The camera captures the user's eye position at 30 frames per second. As output, it generates captured gaze data. This gaze data includes the user's eye position information for each image frame.
[0817] Step 2:
[0818] The device preprocesses the captured gaze data to generate clear data with noise removed. It receives the captured gaze data as input, performs high-frequency noise removal filtering and grayscale conversion, and generates clear preprocessed gaze data as output.
[0819] Step 3:
[0820] The device sends the preprocessed gaze data to the server. As input, it receives preprocessed clear gaze data. The data is encrypted and securely transmitted over the network. As output, it generates the gaze data sent to the server.
[0821] Step 4:
[0822] The server analyzes the received gaze data using a machine learning algorithm to identify the coordinates where the user is looking. Preprocessed gaze data is received as input. A deep learning model is used as the machine learning algorithm. Specifically, it extracts features from the gaze data, inputs them into the model, and predicts the coordinates. As output, it generates coordinate data where the user is looking.
[0823] Step 5:
[0824] The server sends the parsed coordinate data to the device. As input, it receives the coordinate data of the user's view. The data is sent back to the device in real time. As output, it generates the coordinate data sent to the device.
[0825] Step 6:
[0826] The device identifies the character or icon the user is looking at based on the received coordinate data. As input, it receives the coordinate data the user is looking at. The coordinates are mapped to each key on the virtual keyboard, and when the gaze is directed, the character is identified. Specifically, it compares the coordinates with the key position on the virtual keyboard and identifies the matching character. As output, it generates the identified character data.
[0827] Step 7:
[0828] The terminal displays the identified character data on the screen in real time. As input, it receives the identified character data. As a specific action, it displays the character data in an appropriate position on the screen. As output, it provides visual feedback to the user.
[0829] Step 8:
[0830] The terminal generates text data based on the identified character data. As input, it receives a sequence of character data. As a specific operation, it concatenates the identified characters to create one piece of text data. As output, it obtains the generated text data.
[0831] Step 9:
[0832] The device sends the generated text data to the server and generates speech using a speech synthesis algorithm. The generated text data is received as input. The server generates natural-sounding speech using pre-stored samples of the user's voice. Specific operations include extracting phonemes from the text and generating a speech waveform. The generated speech data is obtained as output.
[0833] Step 10:
[0834] The server sends the generated voice data to the device, which then plays it. The generated voice data is received as input. Specific operations include passing the voice data to a playback device and outputting the voice. The output is the user's voice being transmitted to the surroundings.
[0835] Step 11:
[0836] The server stores the generated text and audio data in a database. As input, it receives the generated text and audio data. As a specific operation, it performs a write operation to the database. As output, it obtains the stored data, which can be accessed later.
[0837] Step 12:
[0838] The server sends notifications to pre-defined contacts when certain keywords or events occur. As input, it monitors the generated text data. Specific operations include checking whether the keywords are included and, if so, creating a notification message. As output, a notification is generated and sent to the pre-defined contacts.
[0839] (Application example 1)
[0840] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0841] In autonomous vehicles, passengers are expected to be able to intuitively set their destinations and respond quickly in emergencies, but current systems are cumbersome to operate, making them inconvenient for people with physical disabilities and those unfamiliar with technology. Furthermore, there is a lack of more natural, human-like means of communication, and improvements are needed to enhance passenger convenience and safety.
[0842] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0843] In this invention, the server includes means for capturing a user's gaze using the device's camera, means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze, means for generating text data using the identified coordinates, means for inputting the generated text data into a speech synthesis algorithm to generate speech, means for providing an interface for operating the autonomous vehicle, and means for selecting a destination based on the identified coordinates. This allows passengers to easily and intuitively set their destination using their gaze and receive feedback via voice response. Furthermore, in the event of an emergency, notifications are quickly generated based on specific keywords and sent to pre-defined contacts, ensuring passenger safety.
[0844] A "device camera" is an optical device used to capture a user's gaze.
[0845] "Gaze data" is information indicating the direction in which the user's gaze is directed.
[0846] "Coordinates" are numerical data used to indicate a specific point, and are usually composed of X-axis and Y-axis values.
[0847] "Text data" is digital information expressed as a string of characters or sentences.
[0848] A "speech synthesis algorithm" is a program or calculation method for converting input text data into speech.
[0849] "Speech" refers to sound signals generated by speech synthesis algorithms that mimic human speech.
[0850] An "autonomous vehicle" is a vehicle that can drive autonomously without the need for human driving.
[0851] An "interface" is a means or tool for exchanging information between a user and a system.
[0852] "Destination selection" refers to the operation or process by which a user specifies the place they want to go.
[0853] A "notification" is a message or alert that informs a user or other interested party of specific information.
[0854] A system embodying this invention utilizes gaze recognition technology and speech synthesis technology to improve passenger convenience and safety in autonomous vehicles. The system includes a camera for capturing gaze data and hardware and software for analyzing the data.
[0855] The main components of the system are:
[0856] 1. Camera
[0857] The camera is installed inside the vehicle and captures the direction and focus of the passenger's gaze, allowing the user's intended actions to be visually detected.
[0858] 2. Gaze Recognition Algorithm
[0859] The gaze data captured by the camera is analyzed using gaze recognition algorithms, which use machine learning libraries such as TensorFlow.
[0860] 3. Data analysis and text generation
[0861] The coordinates obtained by the gaze recognition algorithm are mapped to specific characters or actions, and the coordinate data is sent to the device, which generates text corresponding to the specified action.
[0862] 4. Speech Synthesis Algorithm
[0863] The generated text data is converted to audio data using a speech synthesis algorithm (e.g., Google TTS API), which is then played back through the device's speaker to provide feedback to the passenger.
[0864] 5. Database and Notification System
[0865] The generated text and voice data is sent to a server and stored in a database. If a specific keyword (e.g., "help") is entered, a notification is sent to pre-defined contacts using the Twilio API.
[0866] Illustrative Usage Scenarios
[0867] Passengers can access the destination selection screen by directing their gaze towards the camera inside the autonomous vehicle. For example, if they want to go to "Shibuya Station," they focus their gaze on the "Shibuya Station" icon. This gaze action generates the text data "Shibuya Station," and a speech synthesis algorithm generates a voice saying, "Your destination is Shibuya Station, right?" The voice is played over the in-car speaker, and passengers are asked to confirm. At the same time, this information is stored on a server, and emergency contacts are notified if necessary.
[0868] Prompt Sentence Examples
[0869] Choose your destination by sight. Choose from the list below:
[0870] 1. Shibuya Station
[0871] 2. Tokyo Tower
[0872] 3. Nearby restaurants
[0873] 4. Library
[0874] Focus your gaze on your chosen destination and stop scrolling.
[0875] When implemented, these elements and algorithms will work seamlessly together to design a system that allows passengers to operate the system intuitively. Passengers will be able to operate the system with their eyes and receive feedback through voice, significantly improving convenience and safety.
[0876] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0877] Step 1:
[0878] Capture gaze data
[0879] The user looks at the camera in the autonomous vehicle. The camera captures gaze data at regular intervals, recording the direction and focus of the gaze. The captured image data is then input into the gaze recognition algorithm.
[0880] Input: Image data captured by the camera
[0881] Output: Gaze data
[0882] Step 2:
[0883] Preprocessing and transmission of gaze data
[0884] The device preprocesses the captured gaze data to remove noise, and then sends the preprocessed data to the server.
[0885] Input: Captured gaze data
[0886] Output: Preprocessed gaze data
[0887] Step 3:
[0888] Analysis of gaze data
[0889] The server receives the preprocessed gaze data and uses a gaze recognition algorithm (such as TensorFlow) to identify the coordinates where the gaze is pointing. This coordinate data is then sent from the server to the device.
[0890] Input: Preprocessed gaze data
[0891] Output: Identified coordinate data
[0892] Step 4:
[0893] Generating text data
[0894] The device uses the received coordinate data to generate text data: gaze is mapped to specific icons and characters, and a corresponding string of characters is generated.
[0895] Input: Identified coordinate data
[0896] Output: Text data
[0897] Step 5:
[0898] Generate audio data
[0899] The device sends the generated text data to the server, which then generates voice data using a speech synthesis algorithm (Google TTS API), which is then sent to the device.
[0900] Input: Text data
[0901] Output: Audio data
[0902] Step 6:
[0903] Playing audio
[0904] The terminal plays the received audio data through the speaker, and passengers receive audio feedback.
[0905] Input: Audio data
[0906] Output: Play audio
[0907] Step 7:
[0908] Data storage
[0909] The server stores the generated text data and voice data in a database.
[0910] Input: Text data, audio data
[0911] Output: Saved data
[0912] Step 8:
[0913] Sending emergency notifications
[0914] If a specific keyword (e.g., help) is entered, the server uses the Twilio API to send a notification to pre-defined contacts.
[0915] Input: Specific keyword
[0916] Output: Send emergency notification
[0917] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0918] The system of the present invention captures a user's gaze using a device's camera, analyzes the gaze data to identify the coordinates where the user is looking, and then combines it with an emotion engine that recognizes the user's emotions. It then generates text data based on this coordinate data and adjusts the tone and intonation of the generated voice based on the results of the emotion engine's estimation and analysis of the user's emotions. Finally, the generated text data is input into a speech synthesis algorithm to generate speech, allowing the user to speak. The system also includes functionality to store the generated text data and voice data in a database and generate notifications based on specific keywords and send them to pre-defined contacts.
[0919] Program processing explanation
[0920] Gaze recognition
[0921] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The captured gaze data is pre-processed to remove noise and generate clear data. This gaze data is sent to a server and analyzed using a machine learning algorithm. After analysis, the coordinates where the user is looking are identified and the coordinate data is sent to the device.
[0922] Examples:
[0923] If a user is using their gaze to operate a virtual keyboard in an application and is trying to input the word "thank you," the device will generate the input string "thank you" as the user moves their gaze to each letter.
[0924] Text data generation and emotion recognition
[0925] The server maps the identified coordinates to corresponding character information to generate text data in real time. At the same time, the device uses gaze data and other related data (e.g., facial expressions, eye movements) to estimate the user's emotions in real time. This estimation result is input into the emotion engine for detailed emotion analysis.
[0926] Examples:
[0927] If the user smiles while inputting "thank you" with their gaze, the emotion engine will analyze the gaze data and facial expression data to determine that the user is feeling happy.
[0928] Voice generation
[0929] The input text data and emotion recognition results are sent from the device to a server, where a speech synthesis algorithm generates a voice. The generated voice is customized using pre-stored samples of the user's voice, and the tone and intonation of the voice are adjusted according to the estimated emotion. For example, if the user is expressing joy, the voice will be generated in a brighter tone.
[0930] Examples:
[0931] If the user types "thank you" and the emotion engine identifies the user's emotion as "joy," the server will generate a voice that pronounces "thank you" in a bright and emotional voice, which the device can then play back and convey to those around it.
[0932] Data storage and notification function
[0933] The generated text and voice data is stored in a database on the server, and the server also sends notifications to pre-defined contacts when specific keywords or events occur, enabling rapid response in emergencies.
[0934] Examples:
[0935] If a user types "help," the server analyzes the text data and sends an email or SMS notification to contacts registered as an emergency message.
[0936] This system enables people with severe motor disabilities to communicate seamlessly and comfortably using gaze input and emotion recognition without the need for external sensors or specialized devices. The addition of a voice generation function also enables more natural and emotionally rich communication. Furthermore, data storage and notification functions ensure user safety and comfort.
[0937] The processing flow will be explained below.
[0938] Step 1:
[0939] When the application starts, the device activates the device's camera, which recognizes the user's face and eye position and begins capturing gaze data.
[0940] Step 2:
[0941] The device captures gaze data at regular intervals (e.g., every 100 milliseconds) and preprocesses it, which involves noise removal and data normalization.
[0942] Step 3:
[0943] The device transmits the captured gaze data, including the user's face position and gaze direction, to the server.
[0944] Step 4:
[0945] The server inputs the received gaze data into a machine learning algorithm for analysis, which identifies the coordinates where the user is looking.
[0946] Step 5:
[0947] The server transmits the identified coordinate data to the terminal.
[0948] Step 6:
[0949] Based on the received coordinate data, the device identifies the corresponding characters or icons and displays on the screen that the user has entered the input they intended.
[0950] Step 7:
[0951] The user operates the virtual keyboard with their eyes, selecting characters one after another. The selected characters are concatenated in real time to generate text data.
[0952] Step 8:
[0953] The terminal transmits the generated text data to the server.
[0954] Step 9:
[0955] The device collects gaze data and related information (e.g., facial expressions, eye movements) and sends it to the emotion engine.
[0956] Step 10:
[0957] The server uses an emotion engine to recognize emotions and sends the results to the device. The emotion engine analyzes the user's gaze data and facial expression data to estimate the user's emotions.
[0958] Step 11:
[0959] The server uses the text data and emotion recognition results to input the speech synthesis algorithm and generate an audio file, adjusting the tone and intonation of the voice depending on the user's emotion.
[0960] Step 12:
[0961] The server transmits the generated audio file to the terminal.
[0962] Step 13:
[0963] The terminal plays the received audio file, and the user's intention is spoken aloud.
[0964] Step 14:
[0965] The server stores all text and audio data in a database and manages it so that it can be referenced and analyzed at a later date.
[0966] Step 15:
[0967] Users or medical staff configure notification settings based on specific keywords or conditions, and when the conditions are met, the server automatically sends notifications to designated contacts.
[0968] Step 16:
[0969] In the event of an emergency, if a preset keyword (e.g., "help") is entered, the server will send an emergency message to pre-defined contacts, allowing for a prompt and appropriate response.
[0970] These processing steps realize a system that integrates gaze recognition, character input, emotion recognition, voice generation, and data management and notification functions.
[0971] Example 2
[0972] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0973] Conventional communication systems require external sensors or specialized equipment for people with severe motor disabilities, which poses problems such as high costs and the hassle of installation. Furthermore, they only provide simple input methods, making it difficult for users to express their emotions naturally. Furthermore, they are unable to respond quickly in emergencies, limiting the means to ensure user safety.
[0974] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's gaze using the device's camera, means for preprocessing the captured gaze data and removing noise, means for analyzing the preprocessed gaze data to identify coordinates corresponding to the gaze, means for generating text data using the identified coordinates, means for recognizing the user's emotion using the gaze data and related data, means for adjusting the tone and intonation of voice based on the generated text data and the emotion recognition results, and means for inputting the generated text data into a speech synthesis algorithm to generate speech. This enables even people with severe motor impairments to communicate naturally and emotionally without the need for external sensors or dedicated devices. Furthermore, the generated text data and voice data can be stored in a database, and notifications can be generated based on specific keywords and sent to pre-defined contacts, ensuring a rapid response in emergencies.
[0975] "Device camera" refers to a device with a camera function used to capture the user's line of sight.
[0976] "Gaze data" refers to data regarding the position and movement of the user's gaze captured by a camera.
[0977] "Preprocessing" refers to the process of applying noise removal and other processing to captured gaze data to generate clear data suitable for analysis.
[0978] "Noise reduction" refers to the process of removing external environmental influences and unnecessary information from captured raw data.
[0979] "Coordinate data" is data that indicates where on the screen the user's line of sight is located, derived by analyzing the line of sight data.
[0980] "Text data" refers to character information generated using gaze data.
[0981] "Emotion recognition" refers to a technology that estimates and classifies a user's emotional state by analyzing gaze data and related data.
[0982] An "emotion engine" refers to an algorithm or system that analyzes a user's gaze data, facial expression data, etc. to infer emotions.
[0983] "Speech synthesis algorithm" refers to a computational method or algorithm for converting text data into speech data.
[0984] "Tone and intonation" refers to the pitch and intonation of sounds in audio data.
[0985] "Database" refers to a digital warehouse for efficiently storing and managing generated text and audio data.
[0986] "Notifications" refer to alerts or messages generated based on specific keywords and sending information to pre-defined contacts.
[0987] The system of the present invention utilizes the device's camera to capture the user's gaze and analyzes the gaze data to determine the user's gaze coordinates. It also incorporates an emotion engine to recognize the user's emotions and adjust the tone and intonation of the voice based on the results. Finally, the generated text data is input into a speech synthesis algorithm to generate speech for the user to speak.
[0988] Gaze recognition
[0989] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The gaze data is mainly used to record the user's eye position and focus object. The captured gaze data is pre-processed in real time to remove noise. This pre-processing is necessary to minimize the influence of the external environment. The pre-processed gaze data is sent to a server and analyzed by a machine learning algorithm. This analysis identifies the coordinates where the user is looking, and the coordinate data is sent back to the device.
[0990] Text data generation and emotion recognition
[0991] The server maps the identified coordinates to corresponding character information and generates text data in real time. At the same time, the device uses gaze data and other related data (e.g., facial expressions and eye movements) to infer the user's emotions in real time. The emotion recognition results are input into the emotion engine for detailed emotion analysis.
[0992] Voice generation
[0993] The text data and emotion recognition results are sent from the device to a server, where a speech synthesis algorithm generates a voice. The generated voice is customized using pre-stored samples of the user's voice. The tone and intonation of the voice are also adjusted according to the estimated emotion. For example, if the user is expressing joy, the voice will be generated with a brighter tone.
[0994] Data storage and notification function
[0995] The generated text and voice data is stored in a database on the server. When a specific keyword or event occurs, the server sends a notification to pre-defined contacts. This notification function allows for a quick response in the event of an emergency.
[0996] Examples:
[0997] If a user is using their eyes to operate a virtual keyboard within an application and is trying to type "thank you," the device identifies each character by directing their gaze to each letter. As the user looks at each letter in "thank you," the device generates the string "thank you." If the user is smiling, the emotion engine recognizes the user's emotion as "joy" and generates a voice message saying "thank you" in a bright tone. If the user types "help," the server analyzes the text data as an emergency message and sends a notification to pre-defined contacts.
[0998] Example prompt:
[0999] "Please explain a new communication system that uses gaze input and emotion recognition, including specific usage scenarios and technical details."
[1000] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1001] Step 1:
[1002] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The camera records the user's eye position and focus and outputs the data as gaze data. Specifically, when the user looks at the display, the camera automatically tracks the movement of the pupils. This data is stored in the device's memory as gaze data.
[1003] Step 2:
[1004] The device performs preprocessing on the captured gaze data. Because gaze data contains noise, a noise reduction filter is applied to generate clear data. The input is raw gaze data, and the output is gaze data with noise removed. Specifically, the device filters out inaccurate data caused by camera resolution and external ambient light.
[1005] Step 3:
[1006] The device sends preprocessed gaze data to a server, which uses a machine learning algorithm to calculate the coordinate data on the screen corresponding to the gaze focus. The input is the preprocessed gaze data, and the output is coordinate data. Specifically, the device analyzes the gaze direction and position information to determine which part of the screen the user is looking at.
[1007] Step 4:
[1008] The server maps the coordinate data to specific character information and generates text data in real time. The input is coordinate data and the output is text data. In concrete terms, when a user tries to input "thank you" on the virtual keyboard with their gaze, the server generates text data by combining characters corresponding to those coordinates one after another.
[1009] Step 5:
[1010] The device uses gaze data and related data (e.g., facial expressions and eye movements) to estimate the user's emotions in real time. This data is input into an emotion engine for detailed emotion analysis. The input is gaze data and facial expression data, and the output is emotion recognition results. Specific operations include recognizing smiling and angry facial expressions and classifying the emotions as "joy" or "anger."
[1011] Step 6:
[1012] The device sends the generated text data and emotion recognition results to the server. The server generates speech using a speech synthesis algorithm. The input is the text data and emotion recognition results, and the output is speech data. In concrete terms, if the text "Thank you" is sent and the emotion recognition result is "joy," a voice saying "Thank you" in a bright tone is generated.
[1013] Step 7:
[1014] The server stores the generated text and voice data in a database. When a specific keyword or event occurs, a notification is sent to pre-defined contacts. The input is the keyword or event data, and the output is the notification message. Specifically, when a user types "help," the server analyzes the data and sends an emergency notification to the contacts.
[1015] Through each processing step, we have created a system that enables even people with severe motor disabilities to communicate in a natural and emotional way.
[1016] (Application example 2)
[1017] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1018] Conventional factory robot operation requires a dedicated remote controller and a complex user interface, resulting in problems such as reduced work efficiency and increased operator burden. Furthermore, the inability to recognize the emotional state of the operator during operation can lead to a lack of safety and efficiency in work. Furthermore, intuitive operation methods based on gaze and emotions have yet to be introduced, creating a demand for a more natural operating environment.
[1019] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's gaze using the device's camera, means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze, means for recognizing the user's emotion, means for generating text data based on the identified coordinates and the recognized emotion, means for inputting the generated text data into a speech synthesis algorithm and generating speech with a tone and intonation corresponding to the emotion, and means for processing and storing the gaze data and emotion data. This enables operators to operate factory robots in an intuitive and natural way using gaze and emotion, thereby improving operation efficiency and safety.
[1020] A "device camera" is a photographic device used to capture a user's gaze data.
[1021] "Gaze data" refers to information about the user's gaze direction and viewpoint position obtained through the device's camera.
[1022] "Coordinates" are position information in two-dimensional or three-dimensional space that is specified based on line-of-sight data.
[1023] "Emotion" refers to the psychological state that the user shows during the gaze operation.
[1024] "Text data" is character information generated based on gaze data and emotion data.
[1025] A "speech synthesis algorithm" is a program or process that converts input text data into speech.
[1026] "Emotional tone and intonation" refers to the timbre and intonation of a voice that is adjusted to suit the user's emotional state.
[1027] The "means for processing and storing gaze data and emotion data" is a function for analyzing and recording captured gaze data and recognized emotion data.
[1028] This invention is a system that captures a user's gaze data using a device's camera and generates text data and voice based on the gaze data and the user's emotion data. This system is realized mainly using the following hardware and software:
[1029] 1. Hardware:
[1030] Device camera: The camera used to capture user gaze data. Examples include a typical webcam or a camera built into a head-mounted display (HMD).
[1031] Head-mounted display (HMD): A device worn by the user that has a camera for eye tracking. An example of a suitable device is the Microsoft HoloLens 2.
[1032] 2. Software:
[1033] OpenCV: A library used to capture and preprocess gaze data.
[1034] dlib: A library for facial landmark detection.
[1035] pyttsx3: A speech synthesis algorithm that converts text data into speech.
[1036] emotion_recognition: A software module for recognizing user emotions.
[1037] gaze_tracking: A library for gaze tracking.
[1038] System configuration and processing flow:
[1039] The server captures the user's gaze data from the device's camera and analyzes it using OpenCV and gaze_tracking. By processing this gaze data, the coordinates of the gaze direction are identified. It also recognizes the user's emotion data in real time using dlib and emotion_recognition.
[1040] Based on the identified coordinates and the recognized emotion data, text data is generated. The generated text data is then fed into a speech synthesis algorithm using pyttsx3 to generate a customized voice with a tone and intonation that corresponds to the user's emotion. The gaze data and emotion data are then appropriately processed and stored for future reference and analysis.
[1041] Examples:
[1042] For example, imagine a user wearing an HMD in a factory and directing their gaze toward a specific machine part. At this time, gaze data is captured and it is determined that the gaze is directed toward a specific coordinate (the location of the machine part). At the same time, if the user's facial expression is recognized as "satisfied," the speech synthesis algorithm will cheerfully report, "The part has been placed in the correct position."
[1043] Example prompt for a generative AI model:
[1044] It recognizes the user's gaze and checks their position if they are looking at a specific work area. It estimates the user's emotions from their facial expressions and, if they are "satisfied," notifies them by voice that "the work is completed."
[1045] In this way, by using the system of the present invention, it becomes possible to operate a robot in a factory intuitively and efficiently.
[1046] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1047] Step 1:
[1048] The user wears the device's camera (e.g., a head-mounted display) and the system is activated. The device's camera begins capturing the user's gaze. This input data is captured as image frames.
[1049] Step 2:
[1050] The gaze data captured from the camera is pre-processed using OpenCV. This pre-processing step involves denoising the image and extracting the desired parts. The output is clean gaze data.
[1051] Step 3:
[1052] The preprocessed gaze data is analyzed using the gaze_tracking library. A specific algorithm is used to determine the gaze direction and the coordinates of the viewpoint. The input of this step is the preprocessed gaze data, and the output is the coordinate information of the gaze pointing.
[1053] Step 4:
[1054] In parallel, we use the dlib library to detect facial landmarks and recognize user emotions through the emotion_recognition module. The input is the captured raw image data, and the output is the user's emotional state (e.g., happy, anger, sadness).
[1055] Step 5:
[1056] Based on the identified coordinates and the recognized emotion data, text data is generated in real time. The server analyzes these input data and generates text that reflects the object the user is looking at and its meaning. The output is the generated text data.
[1057] Step 6:
[1058] The generated text data is converted into speech data using the speech synthesis algorithm pyttsx3. The tone and intonation of the speech are adjusted depending on the emotional state. The input is the generated text data and emotion recognition results, and the output is customized speech data.
[1059] Step 7:
[1060] The generated voice data is fed back to the user through the device, for example, a message such as "The part has been placed correctly" is played in a bright tone. In this step, the user receives the voice feedback and can take further action based on it.
[1061] Step 8:
[1062] The gaze data and emotion data are processed and stored appropriately on the server. When a specific key event occurs, a notification is sent to the configured contacts. The input of this step is gaze data and emotion data, and the output is saving to a database and generating a notification.
[1063] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1064] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1065] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1066] [Fourth embodiment]
[1067] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1068] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1069] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1070] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1071] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1072] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1073] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1074] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1075] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1076] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1077] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1078] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1079] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1080] The system of the present invention utilizes the device's camera to capture a user's gaze and analyzes the gaze data to determine the coordinates where the user is looking. It then generates text data based on the coordinates and inputs the generated text data into a speech synthesis algorithm to generate speech so that the user can speak. It also includes functionality to store the generated text and speech data in a database and generate notifications based on specific keywords and send them to pre-defined contacts.
[1081] Program processing explanation
[1082] Gaze recognition
[1083] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The captured gaze data is pre-processed to remove noise and generate clear data. This gaze data is sent to a server and analyzed using a machine learning algorithm. After analysis, the coordinates where the user is looking are identified and the coordinate data is sent to the device.
[1084] Examples:
[1085] If a user is using their gaze to operate a virtual keyboard in an application and is trying to input the word "thank you," the device will generate the input string "thank you" as the user moves their gaze to each letter.
[1086] Text input
[1087] The coordinates obtained as a result of analyzing the gaze data are mapped to specific characters or icons, and the corresponding characters are entered. The device displays the received character data on the screen and provides real-time feedback to the user, ensuring that the user enters exactly what they intended.
[1088] Examples:
[1089] As the user types "Thank you," the characters are displayed on the screen in real time, providing a clear view of the user's progress.
[1090] Voice generation
[1091] The entered text data is sent from the device to the server, where a speech synthesis algorithm generates speech. The generated speech is customized using pre-stored samples of the user's voice and sent to the device, where it is played back, allowing the user to speak.
[1092] Examples:
[1093] When the text data "Thank you" is generated, the server converts it into speech and uses a pre-stored sample of the user's voice to generate a speech that pronounces "Thank you," which the device can then play back and convey to people around it.
[1094] Data storage and notification function
[1095] The generated text and voice data is stored in a database on the server, and the server also sends notifications to pre-defined contacts when specific keywords or events occur, enabling rapid response in emergencies.
[1096] Examples:
[1097] If a user types "help," the server analyzes the text data and sends an email or SMS notification to contacts registered as an emergency message.
[1098] This system enables people with severe motor disabilities to communicate seamlessly and comfortably using gaze input and generative AI without the need for external sensors or specialized devices. The addition of a voice generation function also enables more natural and emotionally rich communication. Furthermore, data storage and notification functions ensure user safety and comfort.
[1099] The processing flow will be explained below.
[1100] Step 1:
[1101] When the application starts, the device activates the device's camera, which recognizes the user's face and eye position and begins capturing gaze data.
[1102] Step 2:
[1103] The device captures gaze data at regular intervals (e.g., every 100 milliseconds) and performs preprocessing, which involves removing noise and normalizing the data.
[1104] Step 3:
[1105] The device sends preprocessed gaze data, including the user's face position and gaze direction, to the server.
[1106] Step 4:
[1107] The server inputs the received gaze data into a machine learning algorithm for analysis, which identifies the coordinates where the user is looking.
[1108] Step 5:
[1109] The server transmits the identified coordinate data to the terminal.
[1110] Step 6:
[1111] Based on the received coordinate data, the device identifies the corresponding characters or icons and displays on the screen that the user has entered the input they intended.
[1112] Step 7:
[1113] The user operates the virtual keyboard with their eyes, selecting characters one after another. The selected characters are concatenated in real time to generate text data.
[1114] Step 8:
[1115] The terminal transmits the generated text data to the server.
[1116] Step 9:
[1117] The server feeds the text data into a speech synthesis algorithm to generate an audio file, which is customized using pre-stored samples of the user's voice.
[1118] Step 10:
[1119] The server transmits the generated audio file to the terminal.
[1120] Step 11:
[1121] The terminal plays the received audio file, and the user's intention is spoken aloud.
[1122] Step 12:
[1123] The server stores all text and audio data in a database and manages it so that it can be referenced and analyzed at a later date.
[1124] Step 13:
[1125] Users or medical staff configure notification settings based on specific keywords or conditions, and when the conditions are met, the server automatically sends notifications to designated contacts.
[1126] Step 14:
[1127] In the event of an emergency, if a preset keyword (e.g., "help") is entered, the server will send an emergency message to pre-defined contacts, allowing for a prompt and appropriate response.
[1128] These processing steps realize a system that integrates gaze recognition, character input, voice generation, data management and notification functions.
[1129] Example 1
[1130] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1131] This system solves the problem of people with severe motor disabilities finding it difficult to achieve seamless and comfortable communication using gaze input and generative AI without the need for external sensors or specialized devices. It also enables more natural and emotional communication by reliably inputting the user's intended content and enabling it to be spoken aloud, and it also solves the problem of a lack of means to send prompt notifications in emergencies.
[1132] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1133] In this invention, the server includes a means for analyzing gaze data and identifying the coordinates where the user is looking, a means for storing text data and voice data in a database, and a means for generating a notification based on a specific keyword and sending it to a pre-set contact, thereby making it possible to input characters based on the user's gaze, convert the input text data into voice, and send a notification under specific conditions based on the stored data.
[1134] "Device camera" refers to a camera used to capture a user's gaze, including cameras built into electronic devices such as smartphones, tablets, and computers.
[1135] "Gaze Data" refers to data about a user's eye movements and pupil position captured by a device's camera.
[1136] "Preprocessing" refers to the process of removing noise from captured gaze data and converting it into a format suitable for analysis.
[1137] The term "server" refers to a computer system that receives gaze data, analyzes it, and returns the necessary data to the terminal.
[1138] "Analysis" refers to the process of using machine learning algorithms to identify the coordinates where the user is looking from the gaze data.
[1139] "Coordinate data" refers to data indicating the position at which the user's gaze is pointing, as determined by analysis.
[1140] "Terminal" refers to an electronic device that has an interface with the user, accepts eye-gaze input, and receives data from a server.
[1141] "Text data" refers to character string data generated based on the user's line of sight.
[1142] "Speech synthesis algorithm" refers to the calculation procedures or programs used to convert text data into speech.
[1143] "Database" refers to an information system for systematically storing generated text and audio data.
[1144] "Specific keywords" refer to pre-defined important words or phrases contained in the text data entered by the user.
[1145] "Notifications" refer to messages or alerts sent to pre-defined contacts when certain keywords are detected.
[1146] The system of the present invention utilizes a device's camera to capture a user's gaze and analyzes the captured gaze data to determine the coordinates where the user is looking. This coordinate data is then used to generate text data, which is then input into a speech synthesis algorithm to generate speech. The system also stores the generated text and speech data, and generates notifications based on specific keywords and sends them to pre-defined contacts.
[1147] Gaze recognition
[1148] First, when the application is launched, the device starts the device camera and captures the user's gaze data at regular intervals. For this purpose, the camera needs to capture the position of the user's eyes at a high frame rate. For example, it is appropriate for the camera to capture images at 30 frames per second.
[1149] The device then preprocesses the captured gaze data to remove noise and generate clearer data, including high-frequency noise removal filters and grayscale conversion.
[1150] Data analysis
[1151] The preprocessed gaze data is encrypted and securely transmitted over the network to a server, where it is analyzed using machine learning algorithms to identify the coordinates where the user is looking. This analysis can potentially involve the application of deep learning models.
[1152] Text input
[1153] The analyzed coordinate data is sent to the device in real time, and characters and icons are identified based on that data. For example, when the user directs their gaze to each key on a virtual keyboard, the corresponding character is identified. Then, as the user moves their gaze, an input string is generated, and the device displays the string on the screen in real time.
[1154] Voice generation
[1155] The generated text data is sent from the device to a server, where it is converted into speech using a speech synthesis algorithm. The server uses pre-stored samples of the user's voice to generate a natural-sounding voice. This speech data is sent to the device and played back according to the user's settings.
[1156] Data Retention and Notification
[1157] The generated text and voice data is stored in a database on the server. In addition, the server sends notifications to pre-defined contacts when specific keywords or events occur, enabling a rapid response in the event of a user emergency.
[1158] Examples:
[1159] If a user is using their gaze to operate a virtual keyboard within an application and is trying to input the word "thank you," they can identify each character by directing their gaze at each letter. As the user looks at each letter in "thank you," the device generates the input string for "thank you." The generated text data is converted into speech by a connected speech synthesis algorithm, and the device plays it back. For example, the device can play back the audio "thank you" and convey the message to people around them.
[1160] Example prompt sentence:
[1161] "Please explain in natural language a program that uses the device's camera to capture the user's gaze, analyzes the gaze data, and identifies the coordinates where the user is looking. Please also explain the process of generating text data based on the coordinate data, and inputting that data into a speech synthesis algorithm to generate speech."
[1162] This system provides a simple and intuitive means of input using eye gaze control for people with severe motor disabilities, and supports richer communication by outputting the input as voice. It also ensures the user's safety by providing a notification function in case of an emergency.
[1163] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1164] Step 1:
[1165] When the application is launched, the device starts the device's camera and captures the user's gaze data at regular intervals. As input, it receives image data from the device's camera. The camera captures the user's eye position at 30 frames per second. As output, it generates captured gaze data. This gaze data includes the user's eye position information for each image frame.
[1166] Step 2:
[1167] The device preprocesses the captured gaze data to generate clear data with noise removed. It receives the captured gaze data as input, performs high-frequency noise removal filtering and grayscale conversion, and generates clear preprocessed gaze data as output.
[1168] Step 3:
[1169] The device sends the preprocessed gaze data to the server. As input, it receives preprocessed clear gaze data. The data is encrypted and securely transmitted over the network. As output, it generates the gaze data sent to the server.
[1170] Step 4:
[1171] The server analyzes the received gaze data using a machine learning algorithm to identify the coordinates where the user is looking. Preprocessed gaze data is received as input. A deep learning model is used as the machine learning algorithm. Specifically, it extracts features from the gaze data, inputs them into the model, and predicts the coordinates. As output, it generates coordinate data where the user is looking.
[1172] Step 5:
[1173] The server sends the parsed coordinate data to the device. As input, it receives the coordinate data of the user's view. The data is sent back to the device in real time. As output, it generates the coordinate data sent to the device.
[1174] Step 6:
[1175] The device identifies the character or icon the user is looking at based on the received coordinate data. As input, it receives the coordinate data the user is looking at. The coordinates are mapped to each key on the virtual keyboard, and when the gaze is directed, the character is identified. Specifically, it compares the coordinates with the key position on the virtual keyboard and identifies the matching character. As output, it generates the identified character data.
[1176] Step 7:
[1177] The terminal displays the identified character data on the screen in real time. As input, it receives the identified character data. As a specific action, it displays the character data in an appropriate position on the screen. As output, it provides visual feedback to the user.
[1178] Step 8:
[1179] The terminal generates text data based on the identified character data. As input, it receives a sequence of character data. As a specific operation, it concatenates the identified characters to create one piece of text data. As output, it obtains the generated text data.
[1180] Step 9:
[1181] The device sends the generated text data to the server and generates speech using a speech synthesis algorithm. The generated text data is received as input. The server generates natural-sounding speech using pre-stored samples of the user's voice. Specific operations include extracting phonemes from the text and generating a speech waveform. The generated speech data is obtained as output.
[1182] Step 10:
[1183] The server sends the generated voice data to the device, which then plays it. The generated voice data is received as input. Specific operations include passing the voice data to a playback device and outputting the voice. The output is the user's voice being transmitted to the surroundings.
[1184] Step 11:
[1185] The server stores the generated text and audio data in a database. As input, it receives the generated text and audio data. As a specific operation, it performs a write operation to the database. As output, it obtains the stored data, which can be accessed later.
[1186] Step 12:
[1187] The server sends notifications to pre-defined contacts when certain keywords or events occur. As input, it monitors the generated text data. Specific operations include checking whether the keywords are included and, if so, creating a notification message. As output, a notification is generated and sent to the pre-defined contacts.
[1188] (Application example 1)
[1189] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1190] In autonomous vehicles, passengers are expected to be able to intuitively set their destinations and respond quickly in emergencies, but current systems are cumbersome to operate, making them inconvenient for people with physical disabilities and those unfamiliar with technology. Furthermore, there is a lack of more natural, human-like means of communication, and improvements are needed to enhance passenger convenience and safety.
[1191] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1192] In this invention, the server includes means for capturing a user's gaze using the device's camera, means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze, means for generating text data using the identified coordinates, means for inputting the generated text data into a speech synthesis algorithm to generate speech, means for providing an interface for operating the autonomous vehicle, and means for selecting a destination based on the identified coordinates. This allows passengers to easily and intuitively set their destination using their gaze and receive feedback via voice response. Furthermore, in the event of an emergency, notifications are quickly generated based on specific keywords and sent to pre-defined contacts, ensuring passenger safety.
[1193] A "device camera" is an optical device used to capture a user's gaze.
[1194] "Gaze data" is information indicating the direction in which the user's gaze is directed.
[1195] "Coordinates" are numerical data used to indicate a specific point, and are usually composed of X-axis and Y-axis values.
[1196] "Text data" is digital information expressed as a string of characters or sentences.
[1197] A "speech synthesis algorithm" is a program or calculation method for converting input text data into speech.
[1198] "Speech" refers to sound signals generated by speech synthesis algorithms that mimic human speech.
[1199] An "autonomous vehicle" is a vehicle that can drive autonomously without the need for human driving.
[1200] An "interface" is a means or tool for exchanging information between a user and a system.
[1201] "Destination selection" refers to the operation or process by which a user specifies the place they want to go.
[1202] A "notification" is a message or alert that informs a user or other interested party of specific information.
[1203] A system embodying this invention utilizes gaze recognition technology and speech synthesis technology to improve passenger convenience and safety in autonomous vehicles. The system includes a camera for capturing gaze data and hardware and software for analyzing the data.
[1204] The main components of the system are:
[1205] 1. Camera
[1206] The camera is installed inside the vehicle and captures the direction and focus of the passenger's gaze, allowing the user's intended actions to be visually detected.
[1207] 2. Gaze Recognition Algorithm
[1208] The gaze data captured by the camera is analyzed using gaze recognition algorithms, which use machine learning libraries such as TensorFlow.
[1209] 3. Data analysis and text generation
[1210] The coordinates obtained by the gaze recognition algorithm are mapped to specific characters or actions, and the coordinate data is sent to the device, which generates text corresponding to the specified action.
[1211] 4. Speech Synthesis Algorithm
[1212] The generated text data is converted to audio data using a speech synthesis algorithm (e.g., Google TTS API), which is then played back through the device's speaker to provide feedback to the passenger.
[1213] 5. Database and Notification System
[1214] The generated text and voice data is sent to a server and stored in a database. If a specific keyword (e.g., "help") is entered, a notification is sent to pre-defined contacts using the Twilio API.
[1215] Illustrative Usage Scenarios
[1216] Passengers can access the destination selection screen by directing their gaze towards the camera inside the autonomous vehicle. For example, if they want to go to "Shibuya Station," they focus their gaze on the "Shibuya Station" icon. This gaze action generates the text data "Shibuya Station," and a speech synthesis algorithm generates a voice saying, "Your destination is Shibuya Station, right?" The voice is played over the in-car speaker, and passengers are asked to confirm. At the same time, this information is stored on a server, and emergency contacts are notified if necessary.
[1217] Prompt Sentence Examples
[1218] Choose your destination by sight. Choose from the list below:
[1219] 1. Shibuya Station
[1220] 2. Tokyo Tower
[1221] 3. Nearby restaurants
[1222] 4. Library
[1223] Focus your gaze on your chosen destination and stop scrolling.
[1224] When implemented, these elements and algorithms will work seamlessly together to design a system that allows passengers to operate the system intuitively. Passengers will be able to operate the system with their eyes and receive feedback through voice, significantly improving convenience and safety.
[1225] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1226] Step 1:
[1227] Capture gaze data
[1228] The user looks at the camera in the autonomous vehicle. The camera captures gaze data at regular intervals, recording the direction and focus of the gaze. The captured image data is then input into the gaze recognition algorithm.
[1229] Input: Image data captured by the camera
[1230] Output: Gaze data
[1231] Step 2:
[1232] Preprocessing and transmission of gaze data
[1233] The device preprocesses the captured gaze data to remove noise, and then sends the preprocessed data to the server.
[1234] Input: Captured gaze data
[1235] Output: Preprocessed gaze data
[1236] Step 3:
[1237] Analysis of gaze data
[1238] The server receives the preprocessed gaze data and uses a gaze recognition algorithm (such as TensorFlow) to identify the coordinates where the gaze is pointing. This coordinate data is then sent from the server to the device.
[1239] Input: Preprocessed gaze data
[1240] Output: Identified coordinate data
[1241] Step 4:
[1242] Generating text data
[1243] The device uses the received coordinate data to generate text data: gaze is mapped to specific icons and characters, and a corresponding string of characters is generated.
[1244] Input: Identified coordinate data
[1245] Output: Text data
[1246] Step 5:
[1247] Generate audio data
[1248] The device sends the generated text data to the server, which then generates voice data using a speech synthesis algorithm (Google TTS API), which is then sent to the device.
[1249] Input: Text data
[1250] Output: Audio data
[1251] Step 6:
[1252] Playing audio
[1253] The terminal plays the received audio data through the speaker, and passengers receive audio feedback.
[1254] Input: Audio data
[1255] Output: Play audio
[1256] Step 7:
[1257] Data storage
[1258] The server stores the generated text data and voice data in a database.
[1259] Input: Text data, audio data
[1260] Output: Saved data
[1261] Step 8:
[1262] Sending emergency notifications
[1263] If a specific keyword (e.g., help) is entered, the server uses the Twilio API to send a notification to pre-defined contacts.
[1264] Input: Specific keyword
[1265] Output: Send emergency notification
[1266] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1267] The system of the present invention captures a user's gaze using a device's camera, analyzes the gaze data to identify the coordinates where the user is looking, and then combines it with an emotion engine that recognizes the user's emotions. It then generates text data based on this coordinate data and adjusts the tone and intonation of the generated voice based on the results of the emotion engine's estimation and analysis of the user's emotions. Finally, the generated text data is input into a speech synthesis algorithm to generate speech, allowing the user to speak. The system also includes functionality to store the generated text data and voice data in a database and generate notifications based on specific keywords and send them to pre-defined contacts.
[1268] Program processing explanation
[1269] Gaze recognition
[1270] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The captured gaze data is pre-processed to remove noise and generate clear data. This gaze data is sent to a server and analyzed using a machine learning algorithm. After analysis, the coordinates where the user is looking are identified and the coordinate data is sent to the device.
[1271] Examples:
[1272] If a user is using their gaze to operate a virtual keyboard in an application and is trying to input the word "thank you," the device will generate the input string "thank you" as the user moves their gaze to each letter.
[1273] Text data generation and emotion recognition
[1274] The server maps the identified coordinates to corresponding character information to generate text data in real time. At the same time, the device uses gaze data and other related data (e.g., facial expressions, eye movements) to estimate the user's emotions in real time. This estimation result is input into the emotion engine for detailed emotion analysis.
[1275] Examples:
[1276] If the user smiles while inputting "thank you" with their gaze, the emotion engine will analyze the gaze data and facial expression data to determine that the user is feeling happy.
[1277] Voice generation
[1278] The input text data and emotion recognition results are sent from the device to a server, where a speech synthesis algorithm generates a voice. The generated voice is customized using pre-stored samples of the user's voice, and the tone and intonation of the voice are adjusted according to the estimated emotion. For example, if the user is expressing joy, the voice will be generated in a brighter tone.
[1279] Examples:
[1280] If the user types "thank you" and the emotion engine identifies the user's emotion as "joy," the server will generate a voice that pronounces "thank you" in a bright and emotional voice, which the device can then play back and convey to those around it.
[1281] Data storage and notification function
[1282] The generated text and voice data is stored in a database on the server, and the server also sends notifications to pre-defined contacts when specific keywords or events occur, enabling rapid response in emergencies.
[1283] Examples:
[1284] If a user types "help," the server analyzes the text data and sends an email or SMS notification to contacts registered as an emergency message.
[1285] This system enables people with severe motor disabilities to communicate seamlessly and comfortably using gaze input and emotion recognition without the need for external sensors or specialized devices. The addition of a voice generation function also enables more natural and emotionally rich communication. Furthermore, data storage and notification functions ensure user safety and comfort.
[1286] The processing flow will be explained below.
[1287] Step 1:
[1288] When the application starts, the device activates the device's camera, which recognizes the user's face and eye position and begins capturing gaze data.
[1289] Step 2:
[1290] The device captures gaze data at regular intervals (e.g., every 100 milliseconds) and preprocesses it, which involves noise removal and data normalization.
[1291] Step 3:
[1292] The device transmits the captured gaze data, including the user's face position and gaze direction, to the server.
[1293] Step 4:
[1294] The server inputs the received gaze data into a machine learning algorithm for analysis, which identifies the coordinates where the user is looking.
[1295] Step 5:
[1296] The server transmits the identified coordinate data to the terminal.
[1297] Step 6:
[1298] Based on the received coordinate data, the device identifies the corresponding characters or icons and displays on the screen that the user has entered the input they intended.
[1299] Step 7:
[1300] The user operates the virtual keyboard with their eyes, selecting characters one after another. The selected characters are concatenated in real time to generate text data.
[1301] Step 8:
[1302] The terminal transmits the generated text data to the server.
[1303] Step 9:
[1304] The device collects gaze data and related information (e.g., facial expressions, eye movements) and sends it to the emotion engine.
[1305] Step 10:
[1306] The server uses an emotion engine to recognize emotions and sends the results to the device. The emotion engine analyzes the user's gaze data and facial expression data to estimate the user's emotions.
[1307] Step 11:
[1308] The server uses the text data and emotion recognition results to input the speech synthesis algorithm and generate an audio file, adjusting the tone and intonation of the voice depending on the user's emotion.
[1309] Step 12:
[1310] The server transmits the generated audio file to the terminal.
[1311] Step 13:
[1312] The terminal plays the received audio file, and the user's intention is spoken aloud.
[1313] Step 14:
[1314] The server stores all text and audio data in a database and manages it so that it can be referenced and analyzed at a later date.
[1315] Step 15:
[1316] Users or medical staff configure notification settings based on specific keywords or conditions, and when the conditions are met, the server automatically sends notifications to designated contacts.
[1317] Step 16:
[1318] In the event of an emergency, if a preset keyword (e.g., "help") is entered, the server will send an emergency message to pre-defined contacts, allowing for a prompt and appropriate response.
[1319] These processing steps realize a system that integrates gaze recognition, character input, emotion recognition, voice generation, and data management and notification functions.
[1320] Example 2
[1321] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1322] Conventional communication systems require external sensors or specialized equipment for people with severe motor disabilities, which poses problems such as high costs and the hassle of installation. Furthermore, they only provide simple input methods, making it difficult for users to express their emotions naturally. Furthermore, they are unable to respond quickly in emergencies, limiting the means to ensure user safety.
[1323] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's gaze using the device's camera, means for preprocessing the captured gaze data and removing noise, means for analyzing the preprocessed gaze data to identify coordinates corresponding to the gaze, means for generating text data using the identified coordinates, means for recognizing the user's emotion using the gaze data and related data, means for adjusting the tone and intonation of voice based on the generated text data and the emotion recognition results, and means for inputting the generated text data into a speech synthesis algorithm to generate speech. This enables even people with severe motor impairments to communicate naturally and emotionally without the need for external sensors or dedicated devices. Furthermore, the generated text data and voice data can be stored in a database, and notifications can be generated based on specific keywords and sent to pre-defined contacts, ensuring a rapid response in emergencies.
[1324] "Device camera" refers to a device with a camera function used to capture the user's line of sight.
[1325] "Gaze data" refers to data regarding the position and movement of the user's gaze captured by a camera.
[1326] "Preprocessing" refers to the process of applying noise removal and other processing to captured gaze data to generate clear data suitable for analysis.
[1327] "Noise reduction" refers to the process of removing external environmental influences and unnecessary information from captured raw data.
[1328] "Coordinate data" is data that indicates where on the screen the user's line of sight is located, derived by analyzing the line of sight data.
[1329] "Text data" refers to character information generated using gaze data.
[1330] "Emotion recognition" refers to a technology that estimates and classifies a user's emotional state by analyzing gaze data and related data.
[1331] An "emotion engine" refers to an algorithm or system that analyzes a user's gaze data, facial expression data, etc. to infer emotions.
[1332] "Speech synthesis algorithm" refers to a computational method or algorithm for converting text data into speech data.
[1333] "Tone and intonation" refers to the pitch and intonation of sounds in audio data.
[1334] "Database" refers to a digital warehouse for efficiently storing and managing generated text and audio data.
[1335] "Notifications" refer to alerts or messages generated based on specific keywords and sending information to pre-defined contacts.
[1336] The system of the present invention utilizes the device's camera to capture the user's gaze and analyzes the gaze data to determine the user's gaze coordinates. It also incorporates an emotion engine to recognize the user's emotions and adjust the tone and intonation of the voice based on the results. Finally, the generated text data is input into a speech synthesis algorithm to generate speech for the user to speak.
[1337] Gaze recognition
[1338] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The gaze data is mainly used to record the user's eye position and focus object. The captured gaze data is pre-processed in real time to remove noise. This pre-processing is necessary to minimize the influence of the external environment. The pre-processed gaze data is sent to a server and analyzed by a machine learning algorithm. This analysis identifies the coordinates where the user is looking, and the coordinate data is sent back to the device.
[1339] Text data generation and emotion recognition
[1340] The server maps the identified coordinates to corresponding character information and generates text data in real time. At the same time, the device uses gaze data and other related data (e.g., facial expressions and eye movements) to infer the user's emotions in real time. The emotion recognition results are input into the emotion engine for detailed emotion analysis.
[1341] Voice generation
[1342] The text data and emotion recognition results are sent from the device to a server, where a speech synthesis algorithm generates a voice. The generated voice is customized using pre-stored samples of the user's voice. The tone and intonation of the voice are also adjusted according to the estimated emotion. For example, if the user is expressing joy, the voice will be generated with a brighter tone.
[1343] Data storage and notification function
[1344] The generated text and voice data is stored in a database on the server. When a specific keyword or event occurs, the server sends a notification to pre-defined contacts. This notification function allows for a quick response in the event of an emergency.
[1345] Examples:
[1346] If a user is using their eyes to operate a virtual keyboard within an application and is trying to type "thank you," the device identifies each character by directing their gaze to each letter. As the user looks at each letter in "thank you," the device generates the string "thank you." If the user is smiling, the emotion engine recognizes the user's emotion as "joy" and generates a voice message saying "thank you" in a bright tone. If the user types "help," the server analyzes the text data as an emergency message and sends a notification to pre-defined contacts.
[1347] Example prompt:
[1348] "Please explain a new communication system that uses gaze input and emotion recognition, including specific usage scenarios and technical details."
[1349] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1350] Step 1:
[1351] When the application is launched, the device activates the device's camera and captures the user's gaze data at regular intervals. The camera records the user's eye position and focus and outputs the data as gaze data. Specifically, when the user looks at the display, the camera automatically tracks the movement of the pupils. This data is stored in the device's memory as gaze data.
[1352] Step 2:
[1353] The device performs preprocessing on the captured gaze data. Because gaze data contains noise, a noise reduction filter is applied to generate clear data. The input is raw gaze data, and the output is gaze data with noise removed. Specifically, the device filters out inaccurate data caused by camera resolution and external ambient light.
[1354] Step 3:
[1355] The device sends preprocessed gaze data to a server, which uses a machine learning algorithm to calculate the coordinate data on the screen corresponding to the gaze focus. The input is the preprocessed gaze data, and the output is coordinate data. Specifically, the device analyzes the gaze direction and position information to determine which part of the screen the user is looking at.
[1356] Step 4:
[1357] The server maps the coordinate data to specific character information and generates text data in real time. The input is coordinate data and the output is text data. In concrete terms, when a user tries to input "thank you" on the virtual keyboard with their gaze, the server generates text data by combining characters corresponding to those coordinates one after another.
[1358] Step 5:
[1359] The device uses gaze data and related data (e.g., facial expressions and eye movements) to estimate the user's emotions in real time. This data is input into an emotion engine for detailed emotion analysis. The input is gaze data and facial expression data, and the output is emotion recognition results. Specific operations include recognizing smiling and angry facial expressions and classifying the emotions as "joy" or "anger."
[1360] Step 6:
[1361] The device sends the generated text data and emotion recognition results to the server. The server generates speech using a speech synthesis algorithm. The input is the text data and emotion recognition results, and the output is speech data. In concrete terms, if the text "Thank you" is sent and the emotion recognition result is "joy," a voice saying "Thank you" in a bright tone is generated.
[1362] Step 7:
[1363] The server stores the generated text and voice data in a database. When a specific keyword or event occurs, a notification is sent to pre-defined contacts. The input is the keyword or event data, and the output is the notification message. Specifically, when a user types "help," the server analyzes the data and sends an emergency notification to the contacts.
[1364] Through each processing step, we have created a system that enables even people with severe motor disabilities to communicate in a natural and emotional way.
[1365] (Application example 2)
[1366] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1367] Conventional factory robot operation requires a dedicated remote controller and a complex user interface, resulting in problems such as reduced work efficiency and increased operator burden. Furthermore, the inability to recognize the emotional state of the operator during operation can lead to a lack of safety and efficiency in work. Furthermore, intuitive operation methods based on gaze and emotions have yet to be introduced, creating a demand for a more natural operating environment.
[1368] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the user's gaze using the device's camera, means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze, means for recognizing the user's emotion, means for generating text data based on the identified coordinates and the recognized emotion, means for inputting the generated text data into a speech synthesis algorithm and generating speech with a tone and intonation corresponding to the emotion, and means for processing and storing the gaze data and emotion data. This enables operators to operate factory robots in an intuitive and natural way using gaze and emotion, thereby improving operation efficiency and safety.
[1369] A "device camera" is a photographic device used to capture a user's gaze data.
[1370] "Gaze data" refers to information about the user's gaze direction and viewpoint position obtained through the device's camera.
[1371] "Coordinates" are position information in two-dimensional or three-dimensional space that is specified based on line-of-sight data.
[1372] "Emotion" refers to the psychological state that the user shows during the gaze operation.
[1373] "Text data" is character information generated based on gaze data and emotion data.
[1374] A "speech synthesis algorithm" is a program or process that converts input text data into speech.
[1375] "Emotional tone and intonation" refers to the timbre and intonation of a voice that is adjusted to suit the user's emotional state.
[1376] The "means for processing and storing gaze data and emotion data" is a function for analyzing and recording captured gaze data and recognized emotion data.
[1377] This invention is a system that captures a user's gaze data using a device's camera and generates text data and voice based on the gaze data and the user's emotion data. This system is realized mainly using the following hardware and software:
[1378] 1. Hardware:
[1379] Device camera: The camera used to capture user gaze data. Examples include a typical webcam or a camera built into a head-mounted display (HMD).
[1380] Head-mounted display (HMD): A device worn by the user that has a camera for eye tracking. An example of a suitable device is the Microsoft HoloLens 2.
[1381] 2. Software:
[1382] OpenCV: A library used to capture and preprocess gaze data.
[1383] dlib: A library for facial landmark detection.
[1384] pyttsx3: A speech synthesis algorithm that converts text data into speech.
[1385] emotion_recognition: A software module for recognizing user emotions.
[1386] gaze_tracking: A library for gaze tracking.
[1387] System configuration and processing flow:
[1388] The server captures the user's gaze data from the device's camera and analyzes it using OpenCV and gaze_tracking. By processing this gaze data, the coordinates of the gaze direction are identified. It also recognizes the user's emotion data in real time using dlib and emotion_recognition.
[1389] Based on the identified coordinates and the recognized emotion data, text data is generated. The generated text data is then fed into a speech synthesis algorithm using pyttsx3 to generate a customized voice with a tone and intonation that corresponds to the user's emotion. The gaze data and emotion data are then appropriately processed and stored for future reference and analysis.
[1390] Examples:
[1391] For example, imagine a user wearing an HMD in a factory and directing their gaze toward a specific machine part. At this time, gaze data is captured and it is determined that the gaze is directed toward a specific coordinate (the location of the machine part). At the same time, if the user's facial expression is recognized as "satisfied," the speech synthesis algorithm will cheerfully report, "The part has been placed in the correct position."
[1392] Example prompt for a generative AI model:
[1393] It recognizes the user's gaze and checks their position if they are looking at a specific work area. It estimates the user's emotions from their facial expressions and, if they are "satisfied," notifies them by voice that "the work is completed."
[1394] In this way, by using the system of the present invention, it becomes possible to operate a robot in a factory intuitively and efficiently.
[1395] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1396] Step 1:
[1397] The user wears the device's camera (e.g., a head-mounted display) and the system is activated. The device's camera begins capturing the user's gaze. This input data is captured as image frames.
[1398] Step 2:
[1399] The gaze data captured from the camera is pre-processed using OpenCV. This pre-processing step involves denoising the image and extracting the desired parts. The output is clean gaze data.
[1400] Step 3:
[1401] The preprocessed gaze data is analyzed using the gaze_tracking library. A specific algorithm is used to determine the gaze direction and the coordinates of the viewpoint. The input of this step is the preprocessed gaze data, and the output is the coordinate information of the gaze pointing.
[1402] Step 4:
[1403] In parallel, we use the dlib library to detect facial landmarks and recognize user emotions through the emotion_recognition module. The input is the captured raw image data, and the output is the user's emotional state (e.g., happy, anger, sadness).
[1404] Step 5:
[1405] Based on the identified coordinates and the recognized emotion data, text data is generated in real time. The server analyzes these input data and generates text that reflects the object the user is looking at and its meaning. The output is the generated text data.
[1406] Step 6:
[1407] The generated text data is converted into speech data using the speech synthesis algorithm pyttsx3. The tone and intonation of the speech are adjusted depending on the emotional state. The input is the generated text data and emotion recognition results, and the output is customized speech data.
[1408] Step 7:
[1409] The generated voice data is fed back to the user through the device, for example, a message such as "The part has been placed correctly" is played in a bright tone. In this step, the user receives the voice feedback and can take further action based on it.
[1410] Step 8:
[1411] The gaze data and emotion data are processed and stored appropriately on the server. When a specific key event occurs, a notification is sent to the configured contacts. The input of this step is gaze data and emotion data, and the output is saving to a database and generating a notification.
[1412] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1413] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1414] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1415] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1416] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1417] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1418] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1419] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1420] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1421] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1422] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1423] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1424] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1425] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1426] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1427] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1428] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1429] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1430] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1431] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1432] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1433] The following is further disclosed regarding the above embodiment.
[1434] (Claim 1)
[1435] means for capturing a user's gaze using a camera on the device;
[1436] means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze;
[1437] means for generating text data using the identified coordinates;
[1438] The system includes means for inputting the generated text data into a speech synthesis algorithm to generate speech.
[1439] (Claim 2)
[1440] 10. The system of claim 1,
[1441] A system including means for storing text data and audio data in a database.
[1442] (Claim 3)
[1443] 10. The system of claim 1,
[1444] A system that includes a means for generating notifications based on specific keywords and sending them to pre-defined contacts.
[1445] "Example 1"
[1446] (Claim 1)
[1447] means for capturing a user's gaze using a camera on the device;
[1448] means for preprocessing and denoising the captured gaze data;
[1449] means for transmitting the preprocessed gaze data to a server;
[1450] A means for the server to analyze the gaze data and identify the coordinates where the user is looking;
[1451] means for returning the identified coordinate data to the terminal;
[1452] means for generating text data using the coordinate data;
[1453] a means for inputting the generated text data into a speech synthesis algorithm to generate speech;
[1454] means for storing the text data and audio data in a database;
[1455] A system that includes a means for generating notifications based on specific keywords and sending them to pre-defined contacts.
[1456] (Claim 2)
[1457] 10. The system of claim 1, wherein the generated text data and audio data are stored in a database.
[1458] (Claim 3)
[1459] 10. The system of claim 1, wherein notifications are generated and sent to pre-defined contacts based on specific keywords.
[1460] "Application Example 1"
[1461] (Claim 1)
[1462] means for capturing a user's gaze using a camera on the device;
[1463] means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze;
[1464] means for generating text data using the identified coordinates;
[1465] a means for inputting the generated text data into a speech synthesis algorithm to generate speech;
[1466] a means for providing an interface for operating the automated driving vehicle;
[1467] The system includes a means for making a destination selection based on the identified coordinates.
[1468] (Claim 2)
[1469] 10. The system of claim 1, wherein the text data and the audio data are stored in a database.
[1470] (Claim 3)
[1471] 10. The system of claim 1, wherein notifications are generated based on specific keywords and sent to predefined contacts.
[1472] "Example 2: Combining Emotion Engines"
[1473] (Claim 1)
[1474] means for capturing a user's gaze using a camera on the device;
[1475] means for preprocessing and denoising the captured gaze data;
[1476] means for analyzing the preprocessed gaze data to identify coordinates corresponding to the gaze;
[1477] means for generating text data using the identified coordinates;
[1478] means for recognizing a user's emotion using gaze data and related data;
[1479] means for adjusting the tone and intonation of the voice based on the generated text data and the emotion recognition results;
[1480] The system includes means for inputting the generated text data into a speech synthesis algorithm to generate speech.
[1481] (Claim 2)
[1482] 10. The system of claim 1, further comprising means for storing the generated text data and audio data in a database.
[1483] (Claim 3)
[1484] 10. The system of claim 1, further comprising means for generating and sending notifications to predefined contacts based on specific keywords.
[1485] "Application example 2 when combining emotion engines"
[1486] (Claim 1)
[1487] means for capturing a user's gaze using a camera on the device;
[1488] means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze;
[1489] means for recognizing a user's emotion;
[1490] means for generating text data based on the identified coordinates and the recognized emotion;
[1491] A means for inputting the generated text data into a speech synthesis algorithm to generate speech with a tone and intonation corresponding to the emotion;
[1492] A system including means for processing and storing gaze data and emotion data.
[1493] (Claim 2)
[1494] 10. The system of claim 1, further comprising means for storing the text data and the audio data in a database.
[1495] (Claim 3)
[1496] 10. The system of claim 1, further comprising means for generating and sending notifications to pre-defined contacts based on specific keywords. [Explanation of symbols]
[1497] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for capturing a user's gaze using a camera on the device; means for analyzing the captured gaze data and identifying coordinates corresponding to the gaze; means for generating text data using the identified coordinates; The system includes means for inputting the generated text data into a speech synthesis algorithm to generate speech.
2. 10. The system of claim 1, A system including means for storing text data and audio data in a database.
3. 10. The system of claim 1, A system that includes a means for generating notifications based on specific keywords and sending them to pre-defined contacts.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A