system

The system addresses aphasic patients' communication challenges by analyzing audio and video data to generate natural responses, improving social interactions and reducing caregiver stress through secure, intuitive communication.

JP2026074939APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Aphasic patients face significant challenges in normal conversations and expressing intentions, leading to social isolation and communication difficulties for both patients and caregivers, necessitating a system that can interpret user intentions and emotions through audio and video data to facilitate natural communication.

Method used

A system that analyzes audio and video data to interpret user intentions and emotions, generating natural responses using speech synthesis and encryption for secure communication, incorporating facial and gesture recognition to enhance data security and prevent unauthorized access.

Benefits of technology

Enables aphasic patients to communicate smoothly with others by accurately interpreting non-verbal cues and generating appropriate responses, enhancing social connections and reducing stress for caregivers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074939000001_ABST
    Figure 2026074939000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for acquiring voice data and video data from a user, Means for converting the voice data into text data, Means for extracting face information from the video data and estimating an emotional state, Means for analyzing gesture information from the video data and interpreting the user's intention, Means for generating a response based on the text data and the user's intention, Means for converting the response into voice data, Means for presenting the voice data or text data to the user, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Aphasic patients have great difficulties in normal conversations and expressing intentions, and thus are severely restricted in their daily lives and social activities. For this reason, the increase in social isolation due to communication disorders and the lack of support in the welfare and medical fields have become problems. In addition, not only the patients themselves but also caregivers and family members have difficulties in communication and are under stress. Therefore, there is a need to develop a support system that allows users to express emotions and intentions in a natural way through voice and gestures and communicate smoothly with others.

Means for Solving the Problems

[0005] This invention provides a system that can appropriately interpret a user's intentions and emotions by analyzing audio and video data acquired from the user, and generate natural responses based on this interpretation. Specifically, it includes means for converting audio data into text data, means for extracting the user's facial information from video data and estimating their emotional state, and means for analyzing gesture information and interpreting the user's intentions. Furthermore, by providing means for generating a response based on the information obtained and presenting it to the user as audio or text data, the system enables aphasic patients to communicate smoothly with others. In addition, to enhance data security and prevent unauthorized use of audio and video data, the system includes means for encryption and access control.

[0006] "Users" refer to aphasic patients who attempt to communicate using this system, as well as those receiving support from them.

[0007] "Audio data" refers to information obtained by converting audio signals collected from users using microphones or similar devices into a digital format.

[0008] "Video data" refers to visual information, such as a user's face or gestures, collected using a camera or similar device and stored in digital format.

[0009] "Text data" refers to character data converted from audio data by a generation engine, which enables the expression of intentions and information in written form.

[0010] "Facial information" refers to information used to analyze emotional states and changes in facial expressions based on the user's facial features extracted from video data.

[0011] "Gesture information" refers to information used to understand the user's intentions and will by analyzing hand and body movements obtained from video data.

[0012] "Response" refers to a natural conversation or reaction created by the generative AI based on analyzed text data and gesture information.

[0013] A "speech synthesis engine" is a system element that uses technology to convert text data into speech signals and generate easily understandable speech.

[0014] "Encryption" refers to the technology used to enhance data security by converting digital data into a format that cannot be deciphered by third parties.

[0015] "Access control" refers to a management method that grants access rights to data only to specific users or devices, thereby preventing unauthorized access. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] Displays an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be described.

[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] The system for implementing the present invention constructs a personal assistant as a program to facilitate smooth communication with the user. This system analyzes a large amount of data obtained from the user's voice and video and generates appropriate responses, thereby enabling aphasic patients to engage in natural conversations with others.

[0038] System Configuration

[0039] The device is equipped with a microphone and camera, which capture the user's voice signals and video in real time. This allows for detailed capture of the user's speech, facial expressions, and gestures.

[0040] The server is equipped with a speech recognition module to analyze the acquired audio data and instantly convert it into text data. The speech recognition technology utilizes the latest AI, enabling accurate extraction of meaning even if the user's speech is unclear.

[0041] The video data is processed on the server using face recognition and gesture recognition algorithms. The face recognition module analyzes the user's facial expressions and features to estimate their emotions. The gesture recognition module understands the content of gestures and sign language from the user's hand movements and overall body movements.

[0042] The server integrates the analyzed audio-text data with gesture information obtained from the video, and uses generative AI to generate natural and contextual responses. The generated responses are output as text data.

[0043] The text data of the response is converted into a speech signal using a speech synthesis engine. This synthesized speech is played back at a speed and tone that is easy for the user to understand, enabling natural conversation.

[0044] Specific example

[0045] The user makes a mouth gesture indicating they want to drink tea.

[0046] The device captures voice and gestures and sends them to the server.

[0047] The server converts the speech into text and generates the response, "Would you like some tea?"

[0048] The server synthesizes the generated response into speech, and the terminal plays the audio for the user.

[0049] In this way, this system can provide appropriate responses based on various nonverbal information provided by the user, making it possible to maintain and strengthen social connections for aphasia patients and provide them with new enjoyment and a sense of security in their daily lives.

[0050] The following describes the processing flow.

[0051] Step 1:

[0052] To initiate communication, the user speaks to the device or makes gestures.

[0053] Step 2:

[0054] The device acquires user voice data via the microphone and video data via the camera. This data is transmitted to the server in real time.

[0055] Step 3:

[0056] The server inputs the received audio data into a speech recognition engine, which converts the audio into text data. At this stage, it uses the latest speech analysis technology to understand the context even if the speech is unclear.

[0057] Step 4:

[0058] The server extracts facial information from video data and estimates emotional states using a facial recognition algorithm. Simultaneously, it analyzes the user's hand and body movements using a gesture recognition algorithm to understand their intentions and requests.

[0059] Step 5:

[0060] The server integrates voice-to-text and gesture information and uses generative AI to generate appropriate and natural responses. This makes it possible to create conversations that match the user's intentions.

[0061] Step 6:

[0062] The generated response is sent from the server to the terminal and stored as text data. This text data is then converted into speech by a speech synthesis engine.

[0063] Step 7:

[0064] The device plays synthesized speech data to the user through its speaker and displays it as text on the screen when visual feedback is needed. This feedback allows the user to recognize the system's response and take the next step in communication.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] When users communicate smoothly and naturally with others using voice and gestures, their speech may be unclear or their intentions may not be properly conveyed. Furthermore, information security and privacy protection are also important issues. In this context, there is a need to develop systems that enable people with aphasia and other speech disorders to connect more smoothly with society and reduce the difficulties they face in daily life.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes means for acquiring audio and video information, means for converting acquired audio information into text information, means for analyzing facial information from video information and estimating emotional states, means for analyzing gestures and inferring the user's intentions, means for generating a response based on the text information and the user's intentions, means for converting the generated response into audio information, and means for providing audio or text information to the user. This makes it possible to correctly understand the user's ambiguous utterances and nonverbal information, realize natural dialogue, and protect audio and video information.

[0070] A "device for acquiring audio and video information" is a device used to record the user's voice and body movements in real time and to process that data.

[0071] A "device that converts acquired audio information into text information" is a processing device that analyzes audio data and converts its content into an appropriate text format.

[0072] A "device that analyzes facial information from video data to estimate emotional state" is an analytical device that extracts the user's facial features from acquired video data and infers emotions based on those facial expressions.

[0073] A "device that analyzes gestures and infers the user's intentions" is a device that interprets the movements of a user's hands and body from video data and identifies the intentions that those movements convey.

[0074] A "device that generates responses based on textual information and user intent" is a device that creates an appropriate response based on transcribed audio data and the user's nonverbal expressions of intent.

[0075] A "device that converts generated responses into audio information" is a device that synthesizes responses created in text into a format that can be played back as audio.

[0076] A "device that provides audio or textual information to a user" is a device used to present a generated response to a user through auditory or visual means.

[0077] A "device that performs encryption and access restriction" is a system that protects data to ensure the privacy of audio and video information and prevent unauthorized access.

[0078] This invention is a system for supporting interaction with users, and in particular, provides a mechanism for achieving natural communication that does not rely on voice or gestures. The core function of the system is to utilize voice and video data.

[0079] The device is responsible for acquiring user audio and video data in real time using a microphone and camera. This makes it possible to capture not only the user's speech but also non-verbal information such as facial expressions and gestures. The data acquired from the device is transmitted to a server via the network.

[0080] The server uses speech recognition software to convert audio data into text data. This speech recognition utilizes the latest AI technology, specifically generative AI models, which can accurately understand and transcribe even unclear speech. The server also runs facial recognition and gesture recognition algorithms to analyze video data and infer the user's emotional state and intentions.

[0081] Based on the analysis results, the server uses generative AI to generate appropriate responses for the user. This response generation process is configured with prompts designed to deeply understand the user's intent. By setting prompts such as, "If the user is asking for a drink, please make the best suggestion," the system's responses become accurate and natural.

[0082] The generated text-based response is converted back into audio data via speech synthesis software. This audio is then played back to the user through the device, resulting in a more engaging conversation compared to text-based feedback.

[0083] For example, a user might say "I want some tea" aloud or indicate their intention with a gesture. In this case, the device sends this data to the server, which generates and speaks an appropriate response, such as "Would you like some tea?", and plays it back to the user, thereby achieving direct and natural interaction.

[0084] This system provides a comprehensive solution for users to communicate smoothly using voice and gestures. The convenience offered by the invention will support a better everyday communication experience.

[0085] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0086] Step 1:

[0087] The device uses a microphone and camera to acquire audio and video data from the user in real time. This input data includes the user's speech and nonverbal gestures. The device converts this data into a digital format and prepares it for transmission over the network to the server.

[0088] Step 2:

[0089] The server receives audio data transmitted from the terminal. Using the received audio data as input, the speech recognition software on the server converts the audio into text data using a generative AI model. In doing so, it removes noise and analyzes the audio signal to obtain a meaningful string of characters.

[0090] Step 3:

[0091] The server processes video data as input. A facial recognition algorithm extracts facial features from the video data and estimates the user's emotional state based on their facial expressions. Simultaneously, a gesture recognition algorithm analyzes hand and body movements and identifies the intentions behind the gestures. Through these processes, the server interprets the user's nonverbal information.

[0092] Step 4:

[0093] The server integrates text data obtained from audio with emotion and gesture information from video. Using generative AI, it generates contextually appropriate responses based on prompts. Specifically, it outputs an appropriate response to prompts such as, "If the user is asking for a drink, please make the best suggestion." The response generated in this step is in text format.

[0094] Step 5:

[0095] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. During this process, the speed and tone are adjusted to enhance the naturalness of the synthesized speech. The synthesized speech data is then output to the user in a playable format.

[0096] Step 6:

[0097] The terminal receives audio data transmitted from the server and plays it back to the user. Using the speaker, it provides the user with a synthesized speech response, completing two-way communication. The played audio is adjusted to a volume and quality that is easy for the user to understand.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] Modern industrial activities require smooth communication with all users. In particular, providing appropriate communication even to users who have difficulty communicating is a social demand and an important challenge. This invention aims to facilitate dialogue with such diverse users and provide a more natural experience in customer service and support in real-world settings.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes means for acquiring data from a user, means for converting the data into text information, and means for utilizing the proposed response in dialogue in industrial activities. This facilitates smooth dialogue with all users and enables appropriate and natural dialogue even in industrial settings.

[0103] "Means of acquiring data from users" refers to devices or processes that collect users' voice and video in real time.

[0104] "Means for converting the data into text information" refers to a technology or process for converting an audio signal into text format using speech recognition technology.

[0105] "Means for extracting identification information and estimating a state" refers to technologies that recognize a user's face and emotions from video data and determine their psychological or physical state.

[0106] "Means for analyzing motion information and interpreting user intentions" refers to technologies that analyze user gestures and hand movements and understand the user's intentions based on them.

[0107] "Means for generating responses" refers to a process or algorithm for automatically generating appropriate responses for the user based on acquired data.

[0108] "Means for utilizing proposed responses in dialogue in industrial activities" refers to technologies that apply generated responses to communication in multiple industrial scenarios in real society.

[0109] This invention is a system for facilitating smooth communication with users who have difficulty communicating in industrial settings. This system includes a process for collecting and analyzing audio and video data in real time. The main components of this system and their operation are described below.

[0110] The server receives audio and video data collected from users and converts it into text using speech recognition technology. By utilizing a highly accurate AI model for speech recognition, accurate text transcription is possible even when the user's voice is unclear. For example, Google® Cloud Speech-to-Text can be used for speech recognition, and OpenAI® GPT-based models can be used for AI.

[0111] Next, the server uses facial recognition technology to estimate emotions and intentions from the video data. Libraries such as OpenCV are utilized for facial recognition and gesture analysis, analyzing the user's facial expressions and movements. Based on these results, the movements are interpreted to understand the user's underlying intentions.

[0112] Next, the server generates an appropriate response for the user through a generative AI model. This response is then adapted for natural communication specific to the industry. The generated response is then presented to the user as natural speech using a speech synthesis engine such as Amazon Polly.

[0113] As a concrete example, consider a scenario in a store where a customer asks, "What does this product taste like?" The customer's actions and voice are collected by sensors in smart glasses and sent to a server. The server analyzes this data and generates a response such as, "This tea has a matcha flavor and is very mild," which is then made available for staff to hear directly.

[0114] An example of a prompt is given as follows: "The customer is pointing to a product and saying, 'What does this product taste like?' Based on this information, the generating AI should respond with a product description." This approach allows for accurate understanding of the user's intent and the provision of information in a natural manner.

[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0116] Step 1:

[0117] The user makes requests using voice and gestures. The device's microphone and camera capture this information in real time. Specifically, the camera captures the user's voice and actions such as pointing.

[0118] Step 2:

[0119] The terminal transmits the acquired audio and video data to the server. During this process, the data is compressed as needed for efficient transmission. The input consists of audio signals and video frames, while the output is the transmitted digital data.

[0120] Step 3:

[0121] The server uses a speech recognition module to analyze the audio data and convert it into text. This process analyzes the phonemes within the audio and converts them into text. For example, the audio "What does this product taste like?" is converted into text.

[0122] Step 4:

[0123] The server analyzes the video data and extracts the user's facial features and gesture information. It processes the video frames and recognizes the user's facial expressions and movements. Using a facial recognition algorithm, it can recognize actions such as pointing.

[0124] Step 5:

[0125] The server uses a generative AI model to generate an appropriate response based on the speech recognition and gesture recognition results. Using the "generative AI model and prompt sentence," for example, a response such as "This tea has a matcha flavor and is very mild." is generated.

[0126] Step 6:

[0127] The generated response is converted into audio data and sent to the terminal. A speech synthesis engine is used to convert the text information into a natural-sounding audio signal. This audio signal is then returned to the terminal.

[0128] Step 7:

[0129] The device uses the played audio to present responses to the user. This allows the user to receive information in a natural conversational format. As a result, the user finds it easier to understand, and the conversation proceeds smoothly.

[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0131] This invention supports smooth communication for aphasia patients by using a system that incorporates an emotion engine that recognizes the user's emotions. This system generates responses that correspond to the user's emotions during dialogue, thereby achieving natural and realistic communication.

[0132] System Configuration

[0133] The device is equipped with a microphone and camera to collect the user's voice signals and video data. This allows it to acquire non-verbal information such as the user's facial expressions and body movements.

[0134] The server processes the acquired audio data using a speech recognition engine and converts it into text data. Advanced AI technology is used for speech recognition to extract appropriate context even in situations where vocalization is difficult.

[0135] The video data is processed on the server using face recognition and gesture recognition algorithms. The face recognition module analyzes the user's facial expressions and works with the emotion engine to determine their emotional state. The gesture recognition module recognizes the user's intentions from their hand and body movements.

[0136] The emotion engine integrates facial and gesture information to more accurately estimate the user's emotions. This estimation result is then reflected in the generated responses, providing responses that align with the user's emotions.

[0137] The server uses generative AI to create natural-sounding responses based on text data, gesture information, and emotion recognition results. These responses are then synthesized into audio data using speech synthesis technology and transmitted to the user.

[0138] Specific example

[0139] The user smiles and says, "I'm happy today."

[0140] The device captures audio and video data and sends it to the server.

[0141] The server converts the audio into text, and the emotion engine determines that a smile on the face indicates "joy."

[0142] The server generates and speaks a joyful response saying, "Today is a wonderful day."

[0143] The device presents the generated audio to the user, facilitating dialogue that allows for the sharing of emotions.

[0144] This system enables dialogue that is sensitive to the user's emotions, allowing aphasia patients to deepen their social connections. By responding quickly to changes in emotions, it is expected to provide more harmonious communication and improve the user's quality of life.

[0145] The following describes the processing flow.

[0146] Step 1:

[0147] The user initiates communication by speaking or making gestures towards the device.

[0148] Step 2:

[0149] The device uses a microphone and camera to collect audio and video data in real time. This data is sent to a server for processing.

[0150] Step 3:

[0151] The server activates its speech recognition engine and converts the incoming audio data into text data. During this process, it also analyzes the intonation and tone of speech to aim for more accurate text output.

[0152] Step 4:

[0153] The server analyzes the video data and uses a facial recognition algorithm to analyze the user's facial expressions. This estimates the user's emotional state, and the emotion engine determines that emotion.

[0154] Step 5:

[0155] Simultaneously, the server performs gesture recognition to analyze the user's hand and body movements, interpreting the user's intentions and requests from their actions. Based on this information, it prepares a more contextually appropriate response.

[0156] Step 6:

[0157] The emotion engine integrates facial and gesture information to determine the user's emotion and sends the result to the server. This becomes a crucial element in the content of the generated response.

[0158] Step 7:

[0159] The server uses generative AI to generate natural, emotion-optimized responses by integrating text, gesture information, and recognized emotions from the speech.

[0160] Step 8:

[0161] The generated text response is converted into speech data through a speech synthesis engine.

[0162] Step 9:

[0163] The device plays the generated voice response to the user and, if necessary, displays a text response on the screen to visually supplement it. This enables smooth communication with the user.

[0164] (Example 2)

[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0166] The goal is to provide a system that enables users with aphasia or other conditions that make verbal communication difficult to engage in natural, emotionally charged conversations. Furthermore, it requires the system to more accurately recognize the user's complex emotions and intentions and generate appropriate responses based on that understanding. Ensuring the security of audio and video data is also a key challenge.

[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0168] In this invention, the server includes means for acquiring audio and video data from the user, means for converting the audio data into text data, and means for integrating facial information and motion information to estimate the user's emotions. This makes it possible to accurately grasp the user's emotions and provide natural dialogue that corresponds to them.

[0169] A "user" refers to an individual who interacts with a system, provides audio and video data, and communicates with it.

[0170] "Audio data" refers to digital signals that include sound information such as the user's voice and tone.

[0171] "Video data" refers to digital video that includes visual information such as the user's facial expressions and movements.

[0172] "Text data" refers to data that represents the result of converting audio data into text information.

[0173] "Facial features" refers to information extracted from video data to identify the user's facial features and expressions.

[0174] "Motion information" refers to the results of analyzing the user's hand movements and body movements.

[0175] "Emotional state" refers to the state of mind determined from the user's facial expressions and actions.

[0176] "Intention" refers to the purpose or will interpreted based on the user's actions.

[0177] A "generative AI model" refers to an algorithm that uses artificial intelligence to generate natural language responses.

[0178] A "prompt sentence" refers to a sentence that, when input into a generative AI model, prompts the model to generate an appropriate response.

[0179] "Means of converting to audio data" refers to technologies or devices that have the function of converting text data into audio information.

[0180] "Protection" refers to security measures that prevent the unauthorized use or alteration of audio and video data.

[0181] This invention provides a system that supports communication, particularly for patients with aphasia, by recognizing the user's emotions and generating a corresponding response. The following describes embodiments for carrying out the invention.

[0182] First, the device is equipped with a microphone and camera to collect the user's voice, facial expressions, and gestures. This allows for the acquisition of both audio and video data from the user. The audio data includes the user's speech, and its intonation and tone can also be considered. The video data includes non-verbal information such as the user's facial expressions and hand movements.

[0183] Next, the server converts the received audio data into text data using speech recognition software (for example, a general-purpose speech recognition engine). This conversion utilizes advanced speech processing technology to accurately convert even unclear speech into text.

[0184] Furthermore, the server uses facial recognition technology (e.g., a basic facial recognition algorithm) to analyze facial features in the video data and determine the user's emotional state. It also uses gesture recognition technology to analyze the user's movements and determine their intentions.

[0185] Based on this information, the server's emotion engine integrates facial and motion data to estimate the user's emotions. Using this estimation, the server inputs prompt sentences into a generative AI model (for example, a general natural language generation algorithm) to generate a natural response.

[0186] As a concrete example, consider a case where a user smiles and says, "I'm happy today." In this case, the device sends audio and video data to the server. After converting the audio to text, the server uses an emotion engine to determine that the emotion is "joy." An example of a prompt is: "The user's emotion is joy. Please generate an appropriate message for a happy day." Based on this prompt, the generative AI model generates the response, "Today is a wonderful day."

[0187] Subsequently, the generated text responses are converted into audio data by a speech synthesis engine (e.g., basic speech synthesis technology) and provided to the user through the terminal. This process enables users to engage in natural, emotionally charged conversations, thereby improving the quality of communication.

[0188] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0189] Step 1:

[0190] The user speaks aloud to the system and simultaneously makes facial expressions and gestures in front of the camera. This generates audio and video data.

[0191] Step 2:

[0192] The device captures audio data from the user using a microphone and simultaneously captures video data using a camera. The input for this step is the user's raw audio and visual information, and the output is digital audio and video data.

[0193] Step 3:

[0194] The terminal sends the acquired audio and video data to the server. At this stage, the terminal performs the data transfer and prepares to wait for processing on the server side.

[0195] Step 4:

[0196] The server inputs the received audio data into the speech recognition engine and converts it into text data. The input is digital audio data, and the output is text data. Specifically, it performs frequency analysis of the audio and replaces the audio patterns with characters.

[0197] Step 5:

[0198] The server uses a facial recognition module based on video data to analyze the user's facial expressions. The input is video data, and the output is the emotion estimation result based on the facial expressions. The process involves extracting facial feature points and matching them against known emotion patterns.

[0199] Step 6:

[0200] The server analyzes user movement information using a gesture recognition module and determines the user's intention. Video data is input, and the user's intentions based on their movements are output. Specifically, the trajectory of hand movements is tracked, and corresponding movement labels are assigned.

[0201] Step 7:

[0202] The server inputs the results of face recognition and gesture recognition into an emotion engine that estimates the overall emotional state. The input consists of individual emotion estimates and intention information from actions, and the output is an integrated emotion estimate. As an operation, statistical estimation is performed that takes into account different emotional elements.

[0203] Step 8:

[0204] The server inputs a prompt sentence into a generative AI model and generates an appropriate response. The input consists of text data and sentiment estimation results, and the generated response text is output. Specifically, it is a process in which the AI ​​understands the context based on the prompt sentence and generates a new response.

[0205] Step 9:

[0206] The server inputs the response text into a speech synthesis engine and converts it into speech data. The generated response text is then output as natural-sounding speech data. This process involves synthesizing speech waveforms to produce fluent speech.

[0207] Step 10:

[0208] The device presents the generated audio data to the user. The user receives the generated response through the audio emitted from the device. In this step, conversation becomes possible through the audio output device.

[0209] (Application Example 2)

[0210] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0211] Modern brick-and-mortar stores demand quick and appropriate communication with customers. However, it is difficult for employees to instantly understand the diverse emotions of customers and provide appropriate service. Therefore, improving customer satisfaction is a challenge.

[0212] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0213] In this invention, the server includes a device for acquiring audio and video data from a user, a device for converting the audio data into text data, a device for extracting person information and estimating emotional state, a device for analyzing behavioral information and interpreting the customer's intentions, and a device for displaying the emotion recognition results and recommended responses on a wearable visual display device. This enables store employees to grasp the customer's emotions in real time and provide customer service that is appropriate to those emotions.

[0214] "Audio data" refers to sound information obtained from the user, including conversations and ambient sounds.

[0215] "Video data" refers to visual information acquired using cameras or other devices, such as recordings of the user's face and movements.

[0216] "Personal information" refers to information extracted from video data, such as facial features and expressions, used to identify an individual and estimate their condition.

[0217] "Emotional state" refers to the result of estimating the user's emotions based on personal information, and includes emotions such as joy and anger.

[0218] "Motion information" refers to information obtained by analyzing the user's gestures and body movements, used to interpret their intentions and will.

[0219] A "visual presentation device" is a device that can be worn by an individual and is used to visually present video information to the user.

[0220] "Emotion recognition results" refer to the estimated emotions obtained by analyzing emotional states, behavioral information, and other factors.

[0221] A "recommended response" is a suggested response generated based on emotion recognition results, designed to facilitate smoother interactions with the user.

[0222] The system that realizes this invention is configured as follows to support interaction with customers. First, a visual presentation device that can be worn by the user acquires audio and video. This device is equipped with a microphone and a camera and collects audio and video data in real time.

[0223] On the server, AI-based speech recognition software converts the collected audio data into text data. Video data is analyzed to extract person information, and facial recognition technology is used to estimate emotional states. In this process, emotions are analyzed based on the user's facial expressions and movements, and movement information is extracted simultaneously. The emotion recognition engine integrates the facial and movement information to generate emotion recognition results, more accurately estimating the user's emotions.

[0224] Based on these results, the generative AI model automatically generates appropriate recommended responses. These responses are displayed on a visual presentation device and shown to the user along with the emotion recognition results. This device allows the user to smoothly engage in conversations with customers.

[0225] For example, if a customer is smiling while looking at a particular product in a store, the server analyzes the image and recognizes it as "joy." The visual display device then presents the staff with a recommendation, such as, "We highly recommend that product."

[0226] An example of a prompt message is, "The customer is speaking with a smile. Please generate a positive comment."

[0227] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0228] Step 1:

[0229] The terminal acquires audio and video data. The user interacts with the customer through a visual presentation device, during which the microphone and camera capture audio and video information. At this stage, audio signals and visual images are inputs, and they are output as acquired data.

[0230] Step 2:

[0231] The server receives the audio data transmitted from the terminal and converts it into text data using speech recognition software. AI technology analyzes the audio waveform and converts it into words and sentences. This converted text data is the output.

[0232] Step 3:

[0233] The server analyzes person information based on the received video data. Using a face recognition engine, it extracts facial features from the video and estimates the emotional state based on this. The input is video data, and the output is an emotional state such as "joy" or "surprise."

[0234] Step 4:

[0235] The server analyzes video data using gesture recognition software to extract movement information and interpret the user's intentions. It extracts features such as hand and body movements and infers intentions based on them. The input for this step is video data, and the output is the interpreted intention information.

[0236] Step 5:

[0237] The server integrates emotional state and interpreted intention information to generate an emotion recognition result, and then uses a generative AI model to generate a recommended response. Prompts are used to allow the AI ​​to create a natural, emotion-based response. The input is the emotion recognition result and intention information, and the output is the generated recommended response.

[0238] Step 6:

[0239] The server generates a recommended response, which is sent to the terminal and displayed on the visual display device. The user then uses this to facilitate smooth interaction with the customer. The input is the recommended response, and the output is the content displayed on the visual display device.

[0240] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0241] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0242] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0243] [Second Embodiment]

[0244] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0245] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0246] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0247] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0248] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0249] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0250] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0251] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0252] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0253] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0254] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0255] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0256] The system for implementing the present invention constructs a personal assistant as a program to facilitate smooth communication with the user. This system analyzes a large amount of data obtained from the user's voice and video and generates appropriate responses, thereby enabling aphasic patients to engage in natural conversations with others.

[0257] System Configuration

[0258] The device is equipped with a microphone and camera, which capture the user's voice signals and video in real time. This allows for detailed capture of the user's speech, facial expressions, and gestures.

[0259] The server is equipped with a speech recognition module to analyze the acquired audio data and instantly convert it into text data. The speech recognition technology utilizes the latest AI, enabling accurate extraction of meaning even if the user's speech is unclear.

[0260] The video data is processed on the server using face recognition and gesture recognition algorithms. The face recognition module analyzes the user's facial expressions and features to estimate their emotions. The gesture recognition module understands the content of gestures and sign language from the user's hand movements and overall body movements.

[0261] The server integrates the analyzed audio-text data with gesture information obtained from the video, and uses generative AI to generate natural and contextual responses. The generated responses are output as text data.

[0262] The text data of the response is converted into a speech signal using a speech synthesis engine. This synthesized speech is played back at a speed and tone that is easy for the user to understand, enabling natural conversation.

[0263] Specific example

[0264] The user makes a mouth gesture indicating they want to drink tea.

[0265] The device captures voice and gestures and sends them to the server.

[0266] The server converts the speech into text and generates the response, "Would you like some tea?"

[0267] The server synthesizes the generated response into speech, and the terminal plays the audio for the user.

[0268] In this way, this system can provide appropriate responses based on various nonverbal information provided by the user, making it possible to maintain and strengthen social connections for aphasia patients and provide them with new enjoyment and a sense of security in their daily lives.

[0269] The following describes the processing flow.

[0270] Step 1:

[0271] To initiate communication, the user speaks to the device or makes gestures.

[0272] Step 2:

[0273] The device acquires user voice data via the microphone and video data via the camera. This data is transmitted to the server in real time.

[0274] Step 3:

[0275] The server inputs the received audio data into a speech recognition engine, which converts the audio into text data. At this stage, it uses the latest speech analysis technology to understand the context even if the speech is unclear.

[0276] Step 4:

[0277] The server extracts facial information from video data and estimates emotional states using a facial recognition algorithm. Simultaneously, it analyzes the user's hand and body movements using a gesture recognition algorithm to understand their intentions and requests.

[0278] Step 5:

[0279] The server integrates voice-to-text and gesture information and uses generative AI to generate appropriate and natural responses. This makes it possible to create conversations that match the user's intentions.

[0280] Step 6:

[0281] The generated response is sent from the server to the terminal and is also retained as text data. This text data is converted into speech by a speech synthesis engine.

[0282] Step 7:

[0283] The terminal plays the speech-synthesized data for the user through a speaker and displays it as text on the screen if visual feedback is required. With this feedback, the user can recognize the response from the system and take the next communication step.

[0284] (Example 1)

[0285] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0286] When a user communicates smoothly and naturally with others using voice or gestures, the speech may be unclear or the intention may not be conveyed appropriately. Also, information security and privacy protection are important issues. In such situations, there is a demand for the development of a system that enables aphasic patients and others to connect more smoothly with society and reduce the obstacles in daily life.

[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0288] In this invention, the server includes means for acquiring audio and video information, means for converting acquired audio information into text information, means for analyzing facial information from video information and estimating emotional states, means for analyzing gestures and inferring the user's intentions, means for generating a response based on the text information and the user's intentions, means for converting the generated response into audio information, and means for providing audio or text information to the user. This makes it possible to correctly understand the user's ambiguous utterances and nonverbal information, realize natural dialogue, and protect audio and video information.

[0289] A "device for acquiring audio and video information" is a device used to record the user's voice and body movements in real time and to process that data.

[0290] A "device that converts acquired audio information into text information" is a processing device that analyzes audio data and converts its content into an appropriate text format.

[0291] A "device that analyzes facial information from video data to estimate emotional state" is an analytical device that extracts the user's facial features from acquired video data and infers emotions based on those facial expressions.

[0292] A "device that analyzes gestures and infers the user's intentions" is a device that interprets the movements of a user's hands and body from video data and identifies the intentions that those movements convey.

[0293] A "device that generates responses based on textual information and user intent" is a device that creates an appropriate response based on transcribed audio data and the user's nonverbal expressions of intent.

[0294] A "device that converts generated responses into audio information" is a device that synthesizes responses created in text into a format that can be played back as audio.

[0295] A "device that provides audio or textual information to a user" is a device used to present a generated response to a user through auditory or visual means.

[0296] A "device that performs encryption and access restriction" is a system that protects data to ensure the privacy of audio and video information and prevent unauthorized access.

[0297] This invention is a system for supporting interaction with users, and in particular, provides a mechanism for achieving natural communication that does not rely on voice or gestures. The core function of the system is to utilize voice and video data.

[0298] The device is responsible for acquiring user audio and video data in real time using a microphone and camera. This makes it possible to capture not only the user's speech but also non-verbal information such as facial expressions and gestures. The data acquired from the device is transmitted to a server via the network.

[0299] The server uses speech recognition software to convert audio data into text data. This speech recognition utilizes the latest AI technology, specifically generative AI models, which can accurately understand and transcribe even unclear speech. The server also runs facial recognition and gesture recognition algorithms to analyze video data and infer the user's emotional state and intentions.

[0300] Based on the analysis results, the server uses generative AI to generate appropriate responses for the user. This response generation process is configured with prompts designed to deeply understand the user's intent. By setting prompts such as, "If the user is asking for a drink, please make the best suggestion," the system's responses become accurate and natural.

[0301] The generated text-form response is reconverted into audio data via text-to-speech software. This audio is played back to the user through the terminal, enabling a more familiar interaction compared to text-based feedback.

[0302] As a specific example, the user may verbally state "I want to drink tea" or indicate that intention through a gesture. In such a case, the terminal transfers this data to the server, and the server generates an appropriate response such as "Do you want to drink tea?" and converts it to speech to play it back to the user, thus realizing a direct and natural interaction.

[0303] This system provides a comprehensive solution for the user to smoothly communicate using voice and gestures. The convenience brought by the invention makes it possible to support a better daily communication experience.

[0304] The flow of the specific process in Example 1 will be described using FIG. 11.

[0305] Step 1:

[0306] The terminal uses a microphone and a camera to acquire audio data and video data from the user in real time. This input data includes the user's speech content and non-verbal gestures. The terminal converts these data into digital format and prepares to transmit them through the network in order to send them to the server.

[0307] Step 2:

[0308] The server receives the audio data transmitted from the terminal. Using the received audio data as input, the speech recognition software on the server converts the audio into text data using a generative AI model. At that time, noise removal and analysis of the audio signal are performed to obtain a meaningful string.

[0309] Step 3:

[0310] The server processes video data as input. A facial recognition algorithm extracts facial features from the video data and estimates the user's emotional state based on their facial expressions. Simultaneously, a gesture recognition algorithm analyzes hand and body movements and identifies the intentions behind the gestures. Through these processes, the server interprets the user's nonverbal information.

[0311] Step 4:

[0312] The server integrates text data obtained from audio with emotion and gesture information from video. Using generative AI, it generates contextually appropriate responses based on prompts. Specifically, it outputs an appropriate response to prompts such as, "If the user is asking for a drink, please make the best suggestion." The response generated in this step is in text format.

[0313] Step 5:

[0314] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. During this process, the speed and tone are adjusted to enhance the naturalness of the synthesized speech. The synthesized speech data is then output to the user in a playable format.

[0315] Step 6:

[0316] The terminal receives audio data transmitted from the server and plays it back to the user. Using the speaker, it provides the user with a synthesized speech response, completing two-way communication. The played audio is adjusted to a volume and quality that is easy for the user to understand.

[0317] (Application Example 1)

[0318] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0319] Modern industrial activities require smooth communication with all users. In particular, providing appropriate communication even to users who have difficulty communicating is a social demand and an important challenge. This invention aims to facilitate dialogue with such diverse users and provide a more natural experience in customer service and support in real-world settings.

[0320] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0321] In this invention, the server includes means for acquiring data from a user, means for converting the data into text information, and means for utilizing the proposed response in dialogue in industrial activities. This facilitates smooth dialogue with all users and enables appropriate and natural dialogue even in industrial settings.

[0322] "Means of acquiring data from users" refers to devices or processes that collect users' voice and video in real time.

[0323] "Means for converting the data into text information" refers to a technology or process for converting an audio signal into text format using speech recognition technology.

[0324] "Means for extracting identification information and estimating a state" refers to technologies that recognize a user's face and emotions from video data and determine their psychological or physical state.

[0325] "Means for analyzing motion information and interpreting user intentions" refers to technologies that analyze user gestures and hand movements and understand the user's intentions based on them.

[0326] "Means for generating responses" refers to a process or algorithm for automatically generating appropriate responses for the user based on acquired data.

[0327] "Means for utilizing proposed responses in dialogue in industrial activities" refers to technologies that apply generated responses to communication in multiple industrial scenarios in real society.

[0328] This invention is a system for facilitating smooth communication with users who have difficulty communicating in industrial settings. This system includes a process for collecting and analyzing audio and video data in real time. The main components of this system and their operation are described below.

[0329] The server receives audio and video data collected from users and converts it into text using speech recognition technology. By utilizing a highly accurate AI model for speech recognition, accurate text transcription is possible even when the user's speech is unclear. For example, Google Cloud Speech-to-Text can be used for speech recognition, and OpenAI GPT-based models can be used for AI.

[0330] Next, the server uses facial recognition technology to estimate emotions and intentions from the video data. Libraries such as OpenCV are utilized for facial recognition and gesture analysis, analyzing the user's facial expressions and movements. Based on these results, the movements are interpreted to understand the user's underlying intentions.

[0331] Next, the server generates an appropriate response for the user through a generative AI model. This response is then adapted for natural communication specific to the industry. The generated response is then presented to the user as natural speech using a speech synthesis engine such as Amazon Polly.

[0332] As a concrete example, consider a scenario in a store where a customer asks, "What does this product taste like?" The customer's actions and voice are collected by sensors in smart glasses and sent to a server. The server analyzes this data and generates a response such as, "This tea has a matcha flavor and is very mild," which is then made available for staff to hear directly.

[0333] An example of a prompt is given as follows: "The customer is pointing to a product and saying, 'What does this product taste like?' Based on this information, the generating AI should respond with a product description." This approach allows for accurate understanding of the user's intent and the provision of information in a natural manner.

[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0335] Step 1:

[0336] The user makes requests using voice and gestures. The device's microphone and camera capture this information in real time. Specifically, the camera captures the user's voice and actions such as pointing.

[0337] Step 2:

[0338] The terminal transmits the acquired audio and video data to the server. During this process, the data is compressed as needed for efficient transmission. The input consists of audio signals and video frames, while the output is the transmitted digital data.

[0339] Step 3:

[0340] The server uses a speech recognition module to analyze the audio data and convert it into text. This process analyzes the phonemes within the audio and converts them into text. For example, the audio "What does this product taste like?" is converted into text.

[0341] Step 4:

[0342] The server analyzes the video data and extracts the user's facial features and gesture information. It processes the video frames and recognizes the user's facial expressions and movements. Using a facial recognition algorithm, it can recognize actions such as pointing.

[0343] Step 5:

[0344] The server uses a generative AI model to generate an appropriate response based on the speech recognition and gesture recognition results. Using the "generative AI model and prompt sentence," for example, a response such as "This tea has a matcha flavor and is very mild." is generated.

[0345] Step 6:

[0346] The generated response is converted into audio data and sent to the terminal. A speech synthesis engine is used to convert the text information into a natural-sounding audio signal. This audio signal is then returned to the terminal.

[0347] Step 7:

[0348] The device uses the played audio to present responses to the user. This allows the user to receive information in a natural conversational format. As a result, the user finds it easier to understand, and the conversation proceeds smoothly.

[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0350] This invention supports smooth communication for aphasia patients by using a system that incorporates an emotion engine that recognizes the user's emotions. This system generates responses that correspond to the user's emotions during dialogue, thereby achieving natural and realistic communication.

[0351] System Configuration

[0352] The device is equipped with a microphone and camera to collect the user's voice signals and video data. This allows it to acquire non-verbal information such as the user's facial expressions and body movements.

[0353] The server processes the acquired audio data using a speech recognition engine and converts it into text data. Advanced AI technology is used for speech recognition to extract appropriate context even in situations where vocalization is difficult.

[0354] The video data is processed on the server using face recognition and gesture recognition algorithms. The face recognition module analyzes the user's facial expressions and works with the emotion engine to determine their emotional state. The gesture recognition module recognizes the user's intentions from their hand and body movements.

[0355] The emotion engine integrates facial and gesture information to more accurately estimate the user's emotions. This estimation result is then reflected in the generated responses, providing responses that align with the user's emotions.

[0356] The server uses generative AI to create natural-sounding responses based on text data, gesture information, and emotion recognition results. These responses are then synthesized into audio data using speech synthesis technology and transmitted to the user.

[0357] Specific example

[0358] The user smiles and says, "I'm happy today."

[0359] The device captures audio and video data and sends it to the server.

[0360] The server converts the audio into text, and the emotion engine determines that a smile on the face indicates "joy."

[0361] The server generates and speaks a joyful response saying, "Today is a wonderful day."

[0362] The device presents the generated audio to the user, facilitating dialogue that allows for the sharing of emotions.

[0363] This system enables dialogue that is sensitive to the user's emotions, allowing aphasia patients to deepen their social connections. By responding quickly to changes in emotions, it is expected to provide more harmonious communication and improve the user's quality of life.

[0364] The following describes the processing flow.

[0365] Step 1:

[0366] The user initiates communication by speaking or making gestures towards the device.

[0367] Step 2:

[0368] The device uses a microphone and camera to collect audio and video data in real time. This data is sent to a server for processing.

[0369] Step 3:

[0370] The server activates its speech recognition engine and converts the incoming audio data into text data. During this process, it also analyzes the intonation and tone of speech to aim for more accurate text output.

[0371] Step 4:

[0372] The server analyzes the video data and uses a facial recognition algorithm to analyze the user's facial expressions. This estimates the user's emotional state, and the emotion engine determines that emotion.

[0373] Step 5:

[0374] Simultaneously, the server performs gesture recognition to analyze the user's hand and body movements, interpreting the user's intentions and requests from their actions. Based on this information, it prepares a more contextually appropriate response.

[0375] Step 6:

[0376] The emotion engine integrates facial and gesture information to determine the user's emotion and sends the result to the server. This becomes a crucial element in the content of the generated response.

[0377] Step 7:

[0378] The server uses generative AI to generate natural, emotion-optimized responses by integrating text, gesture information, and recognized emotions from the speech.

[0379] Step 8:

[0380] The generated text response is converted into speech data through a speech synthesis engine.

[0381] Step 9:

[0382] The device plays the generated voice response to the user and, if necessary, displays a text response on the screen to visually supplement it. This enables smooth communication with the user.

[0383] (Example 2)

[0384] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0385] The goal is to provide a system that enables users with aphasia or other conditions that make verbal communication difficult to engage in natural, emotionally charged conversations. Furthermore, it requires the system to more accurately recognize the user's complex emotions and intentions and generate appropriate responses based on that understanding. Ensuring the security of audio and video data is also a key challenge.

[0386] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0387] In this invention, the server includes means for acquiring audio and video data from the user, means for converting the audio data into text data, and means for integrating facial information and motion information to estimate the user's emotions. This makes it possible to accurately grasp the user's emotions and provide natural dialogue that corresponds to them.

[0388] A "user" refers to an individual who interacts with a system, provides audio and video data, and communicates with it.

[0389] "Audio data" refers to digital signals that include sound information such as the user's voice and tone.

[0390] "Video data" refers to digital video that includes visual information such as the user's facial expressions and movements.

[0391] "Text data" refers to data that represents the result of converting audio data into text information.

[0392] "Facial features" refers to information extracted from video data to identify the user's facial features and expressions.

[0393] "Motion information" refers to the results of analyzing the user's hand movements and body movements.

[0394] "Emotional state" refers to the state of mind determined from the user's facial expressions and actions.

[0395] "Intention" refers to the purpose or will interpreted based on the user's actions.

[0396] A "generative AI model" refers to an algorithm that uses artificial intelligence to generate natural language responses.

[0397] A "prompt sentence" refers to a sentence that, when input into a generative AI model, prompts the model to generate an appropriate response.

[0398] "Means of converting to audio data" refers to technologies or devices that have the function of converting text data into audio information.

[0399] "Protection" refers to security measures that prevent the unauthorized use or alteration of audio and video data.

[0400] This invention provides a system that supports communication, particularly for patients with aphasia, by recognizing the user's emotions and generating a corresponding response. The following describes embodiments for carrying out the invention.

[0401] First, the device is equipped with a microphone and camera to collect the user's voice, facial expressions, and gestures. This allows for the acquisition of both audio and video data from the user. The audio data includes the user's speech, and its intonation and tone can also be considered. The video data includes non-verbal information such as the user's facial expressions and hand movements.

[0402] Next, the server converts the received audio data into text data using speech recognition software (for example, a general-purpose speech recognition engine). This conversion utilizes advanced speech processing technology to accurately convert even unclear speech into text.

[0403] Furthermore, the server uses facial recognition technology (e.g., a basic facial recognition algorithm) to analyze facial features in the video data and determine the user's emotional state. It also uses gesture recognition technology to analyze the user's movements and determine their intentions.

[0404] Based on this information, the server's emotion engine integrates facial and motion data to estimate the user's emotions. Using this estimation, the server inputs prompt sentences into a generative AI model (for example, a general natural language generation algorithm) to generate a natural response.

[0405] As a concrete example, consider a case where a user smiles and says, "I'm happy today." In this case, the device sends audio and video data to the server. After converting the audio to text, the server uses an emotion engine to determine that the emotion is "joy." An example of a prompt is: "The user's emotion is joy. Please generate an appropriate message for a happy day." Based on this prompt, the generative AI model generates the response, "Today is a wonderful day."

[0406] Subsequently, the generated text responses are converted into audio data by a speech synthesis engine (e.g., basic speech synthesis technology) and provided to the user through the terminal. This process enables users to engage in natural, emotionally charged conversations, thereby improving the quality of communication.

[0407] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0408] Step 1:

[0409] The user speaks aloud to the system and simultaneously makes facial expressions and gestures in front of the camera. This generates audio and video data.

[0410] Step 2:

[0411] The device captures audio data from the user using a microphone and simultaneously captures video data using a camera. The input for this step is the user's raw audio and visual information, and the output is digital audio and video data.

[0412] Step 3:

[0413] The terminal sends the acquired audio and video data to the server. At this stage, the terminal performs the data transfer and prepares to wait for processing on the server side.

[0414] Step 4:

[0415] The server inputs the received audio data into the speech recognition engine and converts it into text data. The input is digital audio data, and the output is text data. Specifically, it performs frequency analysis of the audio and replaces the audio patterns with characters.

[0416] Step 5:

[0417] The server uses a facial recognition module based on video data to analyze the user's facial expressions. The input is video data, and the output is the emotion estimation result based on the facial expressions. The process involves extracting facial feature points and matching them against known emotion patterns.

[0418] Step 6:

[0419] The server analyzes user movement information using a gesture recognition module and determines the user's intention. Video data is input, and the user's intentions based on their movements are output. Specifically, the trajectory of hand movements is tracked, and corresponding movement labels are assigned.

[0420] Step 7:

[0421] The server inputs the results of face recognition and gesture recognition into an emotion engine that estimates the overall emotional state. The input consists of individual emotion estimates and intention information from actions, and the output is an integrated emotion estimate. As an operation, statistical estimation is performed that takes into account different emotional elements.

[0422] Step 8:

[0423] The server inputs a prompt sentence into a generative AI model and generates an appropriate response. The input consists of text data and sentiment estimation results, and the generated response text is output. Specifically, it is a process in which the AI ​​understands the context based on the prompt sentence and generates a new response.

[0424] Step 9:

[0425] The server inputs the response text into a speech synthesis engine and converts it into speech data. The generated response text is then output as natural-sounding speech data. This process involves synthesizing speech waveforms to produce fluent speech.

[0426] Step 10:

[0427] The device presents the generated audio data to the user. The user receives the generated response through the audio emitted from the device. In this step, conversation becomes possible through the audio output device.

[0428] (Application Example 2)

[0429] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0430] Modern brick-and-mortar stores demand quick and appropriate communication with customers. However, it is difficult for employees to instantly understand the diverse emotions of customers and provide appropriate service. Therefore, improving customer satisfaction is a challenge.

[0431] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0432] In this invention, the server includes a device for acquiring audio and video data from a user, a device for converting the audio data into text data, a device for extracting person information and estimating emotional state, a device for analyzing behavioral information and interpreting the customer's intentions, and a device for displaying the emotion recognition results and recommended responses on a wearable visual display device. This enables store employees to grasp the customer's emotions in real time and provide customer service that is appropriate to those emotions.

[0433] "Audio data" refers to sound information obtained from the user, including conversations and ambient sounds.

[0434] "Video data" refers to visual information acquired using cameras or other devices, such as recordings of the user's face and movements.

[0435] "Personal information" refers to information extracted from video data, such as facial features and expressions, used to identify an individual and estimate their condition.

[0436] "Emotional state" refers to the result of estimating the user's emotions based on personal information, and includes emotions such as joy and anger.

[0437] "Motion information" refers to information obtained by analyzing the user's gestures and body movements, used to interpret their intentions and will.

[0438] A "visual presentation device" is a device that can be worn by an individual and is used to visually present video information to the user.

[0439] "Emotion recognition results" refer to the estimated emotions obtained by analyzing emotional states, behavioral information, and other factors.

[0440] A "recommended response" is a suggested response generated based on emotion recognition results, designed to facilitate smoother interactions with the user.

[0441] The system that realizes this invention is configured as follows to support interaction with customers. First, a visual presentation device that can be worn by the user acquires audio and video. This device is equipped with a microphone and a camera and collects audio and video data in real time.

[0442] On the server, AI-based speech recognition software converts the collected audio data into text data. Video data is analyzed to extract person information, and facial recognition technology is used to estimate emotional states. In this process, emotions are analyzed based on the user's facial expressions and movements, and movement information is extracted simultaneously. The emotion recognition engine integrates the facial and movement information to generate emotion recognition results, more accurately estimating the user's emotions.

[0443] Based on these results, the generative AI model automatically generates appropriate recommended responses. These responses are displayed on a visual presentation device and shown to the user along with the emotion recognition results. This device allows the user to smoothly engage in conversations with customers.

[0444] For example, if a customer is smiling while looking at a particular product in a store, the server analyzes the image and recognizes it as "joy." The visual display device then presents the staff with a recommendation, such as, "We highly recommend that product."

[0445] An example of a prompt message is, "The customer is speaking with a smile. Please generate a positive comment."

[0446] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0447] Step 1:

[0448] The terminal acquires audio and video data. The user interacts with the customer through a visual presentation device, during which the microphone and camera capture audio and video information. At this stage, audio signals and visual images are inputs, and they are output as acquired data.

[0449] Step 2:

[0450] The server receives the audio data transmitted from the terminal and converts it into text data using speech recognition software. AI technology analyzes the audio waveform and converts it into words and sentences. This converted text data is the output.

[0451] Step 3:

[0452] The server analyzes person information based on the received video data. Using a face recognition engine, it extracts facial features from the video and estimates the emotional state based on this. The input is video data, and the output is an emotional state such as "joy" or "surprise."

[0453] Step 4:

[0454] The server analyzes video data using gesture recognition software to extract movement information and interpret the user's intentions. It extracts features such as hand and body movements and infers intentions based on them. The input for this step is video data, and the output is the interpreted intention information.

[0455] Step 5:

[0456] The server integrates emotional state and interpreted intention information to generate an emotion recognition result, and then uses a generative AI model to generate a recommended response. Prompts are used to allow the AI ​​to create a natural, emotion-based response. The input is the emotion recognition result and intention information, and the output is the generated recommended response.

[0457] Step 6:

[0458] The server generates a recommended response, which is sent to the terminal and displayed on the visual display device. The user then uses this to facilitate smooth interaction with the customer. The input is the recommended response, and the output is the content displayed on the visual display device.

[0459] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0460] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0461] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0462] [Third Embodiment]

[0463] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0464] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0465] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0466] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0467] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0468] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0469] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0470] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0471] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0472] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0473] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0474] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0475] The system for implementing the present invention constructs a personal assistant as a program to facilitate smooth communication with the user. This system analyzes a large amount of data obtained from the user's voice and video and generates appropriate responses, thereby enabling aphasic patients to engage in natural conversations with others.

[0476] System Configuration

[0477] The device is equipped with a microphone and camera, which capture the user's voice signals and video in real time. This allows for detailed capture of the user's speech, facial expressions, and gestures.

[0478] The server is equipped with a speech recognition module to analyze the acquired audio data and instantly convert it into text data. The speech recognition technology utilizes the latest AI, enabling accurate extraction of meaning even if the user's speech is unclear.

[0479] The video data is processed on the server using face recognition and gesture recognition algorithms. The face recognition module analyzes the user's facial expressions and features to estimate their emotions. The gesture recognition module understands the content of gestures and sign language from the user's hand movements and overall body movements.

[0480] The server integrates the analyzed audio-text data with gesture information obtained from the video, and uses generative AI to generate natural and contextual responses. The generated responses are output as text data.

[0481] The text data of the response is converted into a speech signal using a speech synthesis engine. This synthesized speech is played back at a speed and tone that is easy for the user to understand, enabling natural conversation.

[0482] Specific example

[0483] The user makes a mouth gesture indicating they want to drink tea.

[0484] The device captures voice and gestures and sends them to the server.

[0485] The server converts the speech into text and generates the response, "Would you like some tea?"

[0486] The server synthesizes the generated response into speech, and the terminal plays the audio for the user.

[0487] In this way, this system can provide appropriate responses based on various nonverbal information provided by the user, making it possible to maintain and strengthen social connections for aphasia patients and provide them with new enjoyment and a sense of security in their daily lives.

[0488] The following describes the processing flow.

[0489] Step 1:

[0490] To initiate communication, the user speaks to the device or makes gestures.

[0491] Step 2:

[0492] The device acquires user voice data via the microphone and video data via the camera. This data is transmitted to the server in real time.

[0493] Step 3:

[0494] The server inputs the received audio data into a speech recognition engine, which converts the audio into text data. At this stage, it uses the latest speech analysis technology to understand the context even if the speech is unclear.

[0495] Step 4:

[0496] The server extracts facial information from video data and estimates emotional states using a facial recognition algorithm. Simultaneously, it analyzes the user's hand and body movements using a gesture recognition algorithm to understand their intentions and requests.

[0497] Step 5:

[0498] The server integrates voice-to-text and gesture information and uses generative AI to generate appropriate and natural responses. This makes it possible to create conversations that match the user's intentions.

[0499] Step 6:

[0500] The generated response is sent from the server to the terminal and stored as text data. This text data is then converted into speech by a speech synthesis engine.

[0501] Step 7:

[0502] The device plays synthesized speech data to the user through its speaker and displays it as text on the screen when visual feedback is needed. This feedback allows the user to recognize the system's response and take the next step in communication.

[0503] (Example 1)

[0504] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0505] When users communicate smoothly and naturally with others using voice and gestures, their speech may be unclear or their intentions may not be properly conveyed. Furthermore, information security and privacy protection are also important issues. In this context, there is a need to develop systems that enable people with aphasia and other speech disorders to connect more smoothly with society and reduce the difficulties they face in daily life.

[0506] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0507] In this invention, the server includes means for acquiring audio and video information, means for converting acquired audio information into text information, means for analyzing facial information from video information and estimating emotional states, means for analyzing gestures and inferring the user's intentions, means for generating a response based on the text information and the user's intentions, means for converting the generated response into audio information, and means for providing audio or text information to the user. This makes it possible to correctly understand the user's ambiguous utterances and nonverbal information, realize natural dialogue, and protect audio and video information.

[0508] A "device for acquiring audio and video information" is a device used to record the user's voice and body movements in real time and to process that data.

[0509] A "device that converts acquired audio information into text information" is a processing device that analyzes audio data and converts its content into an appropriate text format.

[0510] A "device that analyzes facial information from video data to estimate emotional state" is an analytical device that extracts the user's facial features from acquired video data and infers emotions based on those facial expressions.

[0511] A "device that analyzes gestures and infers the user's intentions" is a device that interprets the movements of a user's hands and body from video data and identifies the intentions that those movements convey.

[0512] A "device that generates responses based on textual information and user intent" is a device that creates an appropriate response based on transcribed audio data and the user's nonverbal expressions of intent.

[0513] A "device that converts generated responses into audio information" is a device that synthesizes responses created in text into a format that can be played back as audio.

[0514] A "device that provides audio or textual information to a user" is a device used to present a generated response to a user through auditory or visual means.

[0515] A "device that performs encryption and access restriction" is a system that protects data to ensure the privacy of audio and video information and prevent unauthorized access.

[0516] This invention is a system for supporting interaction with users, and in particular, provides a mechanism for achieving natural communication that does not rely on voice or gestures. The core function of the system is to utilize voice and video data.

[0517] The device is responsible for acquiring user audio and video data in real time using a microphone and camera. This makes it possible to capture not only the user's speech but also non-verbal information such as facial expressions and gestures. The data acquired from the device is transmitted to a server via the network.

[0518] The server uses speech recognition software to convert audio data into text data. This speech recognition utilizes the latest AI technology, specifically generative AI models, which can accurately understand and transcribe even unclear speech. The server also runs facial recognition and gesture recognition algorithms to analyze video data and infer the user's emotional state and intentions.

[0519] Based on the analysis results, the server uses generative AI to generate appropriate responses for the user. This response generation process is configured with prompts designed to deeply understand the user's intent. By setting prompts such as, "If the user is asking for a drink, please make the best suggestion," the system's responses become accurate and natural.

[0520] The generated text-based response is converted back into audio data via speech synthesis software. This audio is then played back to the user through the device, resulting in a more engaging conversation compared to text-based feedback.

[0521] For example, a user might say "I want some tea" aloud or indicate their intention with a gesture. In this case, the device sends this data to the server, which generates and speaks an appropriate response, such as "Would you like some tea?", and plays it back to the user, thereby achieving direct and natural interaction.

[0522] This system provides a comprehensive solution for users to communicate smoothly using voice and gestures. The convenience offered by the invention will support a better everyday communication experience.

[0523] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0524] Step 1:

[0525] The device uses a microphone and camera to acquire audio and video data from the user in real time. This input data includes the user's speech and nonverbal gestures. The device converts this data into a digital format and prepares it for transmission over the network to the server.

[0526] Step 2:

[0527] The server receives audio data transmitted from the terminal. Using the received audio data as input, the speech recognition software on the server converts the audio into text data using a generative AI model. In doing so, it removes noise and analyzes the audio signal to obtain a meaningful string of characters.

[0528] Step 3:

[0529] The server processes video data as input. A facial recognition algorithm extracts facial features from the video data and estimates the user's emotional state based on their facial expressions. Simultaneously, a gesture recognition algorithm analyzes hand and body movements and identifies the intentions behind the gestures. Through these processes, the server interprets the user's nonverbal information.

[0530] Step 4:

[0531] The server integrates text data obtained from audio with emotion and gesture information from video. Using generative AI, it generates contextually appropriate responses based on prompts. Specifically, it outputs an appropriate response to prompts such as, "If the user is asking for a drink, please make the best suggestion." The response generated in this step is in text format.

[0532] Step 5:

[0533] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. During this process, the speed and tone are adjusted to enhance the naturalness of the synthesized speech. The synthesized speech data is then output to the user in a playable format.

[0534] Step 6:

[0535] The terminal receives audio data transmitted from the server and plays it back to the user. Using the speaker, it provides the user with a synthesized speech response, completing two-way communication. The played audio is adjusted to a volume and quality that is easy for the user to understand.

[0536] (Application Example 1)

[0537] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0538] Modern industrial activities require smooth communication with all users. In particular, providing appropriate communication even to users who have difficulty communicating is a social demand and an important challenge. This invention aims to facilitate dialogue with such diverse users and provide a more natural experience in customer service and support in real-world settings.

[0539] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0540] In this invention, the server includes means for acquiring data from a user, means for converting the data into text information, and means for utilizing the proposed response in dialogue in industrial activities. This facilitates smooth dialogue with all users and enables appropriate and natural dialogue even in industrial settings.

[0541] "Means of acquiring data from users" refers to devices or processes that collect users' voice and video in real time.

[0542] "Means for converting the data into text information" refers to a technology or process for converting an audio signal into text format using speech recognition technology.

[0543] "Means for extracting identification information and estimating a state" refers to technologies that recognize a user's face and emotions from video data and determine their psychological or physical state.

[0544] "Means for analyzing motion information and interpreting user intentions" refers to technologies that analyze user gestures and hand movements and understand the user's intentions based on them.

[0545] "Means for generating responses" refers to a process or algorithm for automatically generating appropriate responses for the user based on acquired data.

[0546] "Means for utilizing proposed responses in dialogue in industrial activities" refers to technologies that apply generated responses to communication in multiple industrial scenarios in real society.

[0547] This invention is a system for facilitating smooth communication with users who have difficulty communicating in industrial settings. This system includes a process for collecting and analyzing audio and video data in real time. The main components of this system and their operation are described below.

[0548] The server receives audio and video data collected from users and converts it into text using speech recognition technology. By utilizing a highly accurate AI model for speech recognition, accurate text transcription is possible even when the user's speech is unclear. For example, Google Cloud Speech-to-Text can be used for speech recognition, and OpenAI GPT-based models can be used for AI.

[0549] Next, the server uses facial recognition technology to estimate emotions and intentions from the video data. Libraries such as OpenCV are utilized for facial recognition and gesture analysis, analyzing the user's facial expressions and movements. Based on these results, the movements are interpreted to understand the user's underlying intentions.

[0550] Next, the server generates an appropriate response for the user through a generative AI model. This response is then adapted for natural communication specific to the industry. The generated response is then presented to the user as natural speech using a speech synthesis engine such as Amazon Polly.

[0551] As a concrete example, consider a scenario in a store where a customer asks, "What does this product taste like?" The customer's actions and voice are collected by sensors in smart glasses and sent to a server. The server analyzes this data and generates a response such as, "This tea has a matcha flavor and is very mild," which is then made available for staff to hear directly.

[0552] An example of a prompt is given as follows: "The customer is pointing to a product and saying, 'What does this product taste like?' Based on this information, the generating AI should respond with a product description." This approach allows for accurate understanding of the user's intent and the provision of information in a natural manner.

[0553] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0554] Step 1:

[0555] The user makes requests using voice and gestures. The device's microphone and camera capture this information in real time. Specifically, the camera captures the user's voice and actions such as pointing.

[0556] Step 2:

[0557] The terminal transmits the acquired audio and video data to the server. During this process, the data is compressed as needed for efficient transmission. The input consists of audio signals and video frames, while the output is the transmitted digital data.

[0558] Step 3:

[0559] The server uses a speech recognition module to analyze the audio data and convert it into text. This process analyzes the phonemes within the audio and converts them into text. For example, the audio "What does this product taste like?" is converted into text.

[0560] Step 4:

[0561] The server analyzes the video data and extracts the user's facial features and gesture information. It processes the video frames and recognizes the user's facial expressions and movements. Using a facial recognition algorithm, it can recognize actions such as pointing.

[0562] Step 5:

[0563] The server uses a generative AI model to generate an appropriate response based on the speech recognition and gesture recognition results. Using the "generative AI model and prompt sentence," for example, a response such as "This tea has a matcha flavor and is very mild." is generated.

[0564] Step 6:

[0565] The generated response is converted into audio data and sent to the terminal. A speech synthesis engine is used to convert the text information into a natural-sounding audio signal. This audio signal is then returned to the terminal.

[0566] Step 7:

[0567] The device uses the played audio to present responses to the user. This allows the user to receive information in a natural conversational format. As a result, the user finds it easier to understand, and the conversation proceeds smoothly.

[0568] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0569] This invention supports smooth communication for aphasia patients by using a system that incorporates an emotion engine that recognizes the user's emotions. This system generates responses that correspond to the user's emotions during dialogue, thereby achieving natural and realistic communication.

[0570] System Configuration

[0571] The device is equipped with a microphone and camera to collect the user's voice signals and video data. This allows it to acquire non-verbal information such as the user's facial expressions and body movements.

[0572] The server processes the acquired audio data using a speech recognition engine and converts it into text data. Advanced AI technology is used for speech recognition to extract appropriate context even in situations where vocalization is difficult.

[0573] The video data is processed on the server using face recognition and gesture recognition algorithms. The face recognition module analyzes the user's facial expressions and works with the emotion engine to determine their emotional state. The gesture recognition module recognizes the user's intentions from their hand and body movements.

[0574] The emotion engine integrates facial and gesture information to more accurately estimate the user's emotions. This estimation result is then reflected in the generated responses, providing responses that align with the user's emotions.

[0575] The server uses generative AI to create natural-sounding responses based on text data, gesture information, and emotion recognition results. These responses are then synthesized into audio data using speech synthesis technology and transmitted to the user.

[0576] Specific example

[0577] The user smiles and says, "I'm happy today."

[0578] The device captures audio and video data and sends it to the server.

[0579] The server converts the audio into text, and the emotion engine determines that a smile on the face indicates "joy."

[0580] The server generates and speaks a joyful response saying, "Today is a wonderful day."

[0581] The device presents the generated audio to the user, facilitating dialogue that allows for the sharing of emotions.

[0582] This system enables dialogue that is sensitive to the user's emotions, allowing aphasia patients to deepen their social connections. By responding quickly to changes in emotions, it is expected to provide more harmonious communication and improve the user's quality of life.

[0583] The following describes the processing flow.

[0584] Step 1:

[0585] The user initiates communication by speaking or making gestures towards the device.

[0586] Step 2:

[0587] The device uses a microphone and camera to collect audio and video data in real time. This data is sent to a server for processing.

[0588] Step 3:

[0589] The server activates its speech recognition engine and converts the incoming audio data into text data. During this process, it also analyzes the intonation and tone of speech to aim for more accurate text output.

[0590] Step 4:

[0591] The server analyzes the video data and uses a facial recognition algorithm to analyze the user's facial expressions. This estimates the user's emotional state, and the emotion engine determines that emotion.

[0592] Step 5:

[0593] Simultaneously, the server performs gesture recognition to analyze the user's hand and body movements, interpreting the user's intentions and requests from their actions. Based on this information, it prepares a more contextually appropriate response.

[0594] Step 6:

[0595] The emotion engine integrates facial and gesture information to determine the user's emotion and sends the result to the server. This becomes a crucial element in the content of the generated response.

[0596] Step 7:

[0597] The server uses generative AI to generate natural, emotion-optimized responses by integrating text, gesture information, and recognized emotions from the speech.

[0598] Step 8:

[0599] The generated text response is converted into speech data through a speech synthesis engine.

[0600] Step 9:

[0601] The device plays the generated voice response to the user and, if necessary, displays a text response on the screen to visually supplement it. This enables smooth communication with the user.

[0602] (Example 2)

[0603] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0604] The goal is to provide a system that enables users with aphasia or other conditions that make verbal communication difficult to engage in natural, emotionally charged conversations. Furthermore, it requires the system to more accurately recognize the user's complex emotions and intentions and generate appropriate responses based on that understanding. Ensuring the security of audio and video data is also a key challenge.

[0605] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0606] In this invention, the server includes means for acquiring audio and video data from the user, means for converting the audio data into text data, and means for integrating facial information and motion information to estimate the user's emotions. This makes it possible to accurately grasp the user's emotions and provide natural dialogue that corresponds to them.

[0607] A "user" refers to an individual who interacts with a system, provides audio and video data, and communicates with it.

[0608] "Audio data" refers to digital signals that include sound information such as the user's voice and tone.

[0609] "Video data" refers to digital video that includes visual information such as the user's facial expressions and movements.

[0610] "Text data" refers to data that represents the result of converting audio data into text information.

[0611] "Facial features" refers to information extracted from video data to identify the user's facial features and expressions.

[0612] "Motion information" refers to the results of analyzing the user's hand movements and body movements.

[0613] "Emotional state" refers to the state of mind determined from the user's facial expressions and actions.

[0614] "Intention" refers to the purpose or will interpreted based on the user's actions.

[0615] A "generative AI model" refers to an algorithm that uses artificial intelligence to generate natural language responses.

[0616] A "prompt sentence" refers to a sentence that, when input into a generative AI model, prompts the model to generate an appropriate response.

[0617] "Means of converting to audio data" refers to technologies or devices that have the function of converting text data into audio information.

[0618] "Protection" refers to security measures that prevent the unauthorized use or alteration of audio and video data.

[0619] This invention provides a system that supports communication, particularly for patients with aphasia, by recognizing the user's emotions and generating a corresponding response. The following describes embodiments for carrying out the invention.

[0620] First, the device is equipped with a microphone and camera to collect the user's voice, facial expressions, and gestures. This allows for the acquisition of both audio and video data from the user. The audio data includes the user's speech, and its intonation and tone can also be considered. The video data includes non-verbal information such as the user's facial expressions and hand movements.

[0621] Next, the server converts the received audio data into text data using speech recognition software (for example, a general-purpose speech recognition engine). This conversion utilizes advanced speech processing technology to accurately convert even unclear speech into text.

[0622] Furthermore, the server uses facial recognition technology (e.g., a basic facial recognition algorithm) to analyze facial features in the video data and determine the user's emotional state. It also uses gesture recognition technology to analyze the user's movements and determine their intentions.

[0623] Based on this information, the server's emotion engine integrates facial and motion data to estimate the user's emotions. Using this estimation, the server inputs prompt sentences into a generative AI model (for example, a general natural language generation algorithm) to generate a natural response.

[0624] As a concrete example, consider a case where a user smiles and says, "I'm happy today." In this case, the device sends audio and video data to the server. After converting the audio to text, the server uses an emotion engine to determine that the emotion is "joy." An example of a prompt is: "The user's emotion is joy. Please generate an appropriate message for a happy day." Based on this prompt, the generative AI model generates the response, "Today is a wonderful day."

[0625] Subsequently, the generated text responses are converted into audio data by a speech synthesis engine (e.g., basic speech synthesis technology) and provided to the user through the terminal. This process enables users to engage in natural, emotionally charged conversations, thereby improving the quality of communication.

[0626] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0627] Step 1:

[0628] The user speaks aloud to the system and simultaneously makes facial expressions and gestures in front of the camera. This generates audio and video data.

[0629] Step 2:

[0630] The device captures audio data from the user using a microphone and simultaneously captures video data using a camera. The input for this step is the user's raw audio and visual information, and the output is digital audio and video data.

[0631] Step 3:

[0632] The terminal sends the acquired audio and video data to the server. At this stage, the terminal performs the data transfer and prepares to wait for processing on the server side.

[0633] Step 4:

[0634] The server inputs the received audio data into the speech recognition engine and converts it into text data. The input is digital audio data, and the output is text data. Specifically, it performs frequency analysis of the audio and replaces the audio patterns with characters.

[0635] Step 5:

[0636] The server uses a facial recognition module based on video data to analyze the user's facial expressions. The input is video data, and the output is the emotion estimation result based on the facial expressions. The process involves extracting facial feature points and matching them against known emotion patterns.

[0637] Step 6:

[0638] The server analyzes user movement information using a gesture recognition module and determines the user's intention. Video data is input, and the user's intentions based on their movements are output. Specifically, the trajectory of hand movements is tracked, and corresponding movement labels are assigned.

[0639] Step 7:

[0640] The server inputs the results of face recognition and gesture recognition into an emotion engine that estimates the overall emotional state. The input consists of individual emotion estimates and intention information from actions, and the output is an integrated emotion estimate. As an operation, statistical estimation is performed that takes into account different emotional elements.

[0641] Step 8:

[0642] The server inputs a prompt sentence into a generative AI model and generates an appropriate response. The input consists of text data and sentiment estimation results, and the generated response text is output. Specifically, it is a process in which the AI ​​understands the context based on the prompt sentence and generates a new response.

[0643] Step 9:

[0644] The server inputs the response text into a speech synthesis engine and converts it into speech data. The generated response text is then output as natural-sounding speech data. This process involves synthesizing speech waveforms to produce fluent speech.

[0645] Step 10:

[0646] The device presents the generated audio data to the user. The user receives the generated response through the audio emitted from the device. In this step, conversation becomes possible through the audio output device.

[0647] (Application Example 2)

[0648] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0649] Modern brick-and-mortar stores demand quick and appropriate communication with customers. However, it is difficult for employees to instantly understand the diverse emotions of customers and provide appropriate service. Therefore, improving customer satisfaction is a challenge.

[0650] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0651] In this invention, the server includes a device for acquiring audio and video data from a user, a device for converting the audio data into text data, a device for extracting person information and estimating emotional state, a device for analyzing behavioral information and interpreting the customer's intentions, and a device for displaying the emotion recognition results and recommended responses on a wearable visual display device. This enables store employees to grasp the customer's emotions in real time and provide customer service that is appropriate to those emotions.

[0652] "Audio data" refers to sound information obtained from the user, including conversations and ambient sounds.

[0653] "Video data" refers to visual information acquired using cameras or other devices, such as recordings of the user's face and movements.

[0654] "Personal information" refers to information extracted from video data, such as facial features and expressions, used to identify an individual and estimate their condition.

[0655] "Emotional state" refers to the result of estimating the user's emotions based on personal information, and includes emotions such as joy and anger.

[0656] "Motion information" refers to information obtained by analyzing the user's gestures and body movements, used to interpret their intentions and will.

[0657] A "visual presentation device" is a device that can be worn by an individual and is used to visually present video information to the user.

[0658] "Emotion recognition results" refer to the estimated emotions obtained by analyzing emotional states, behavioral information, and other factors.

[0659] A "recommended response" is a suggested response generated based on emotion recognition results, designed to facilitate smoother interactions with the user.

[0660] The system that realizes this invention is configured as follows to support interaction with customers. First, a visual presentation device that can be worn by the user acquires audio and video. This device is equipped with a microphone and a camera and collects audio and video data in real time.

[0661] On the server, AI-based speech recognition software converts the collected audio data into text data. Video data is analyzed to extract person information, and facial recognition technology is used to estimate emotional states. In this process, emotions are analyzed based on the user's facial expressions and movements, and movement information is extracted simultaneously. The emotion recognition engine integrates the facial and movement information to generate emotion recognition results, more accurately estimating the user's emotions.

[0662] Based on these results, the generative AI model automatically generates appropriate recommended responses. These responses are displayed on a visual presentation device and shown to the user along with the emotion recognition results. This device allows the user to smoothly engage in conversations with customers.

[0663] For example, if a customer is smiling while looking at a particular product in a store, the server analyzes the image and recognizes it as "joy." The visual display device then presents the staff with a recommendation, such as, "We highly recommend that product."

[0664] An example of a prompt message is, "The customer is speaking with a smile. Please generate a positive comment."

[0665] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0666] Step 1:

[0667] The terminal acquires audio and video data. The user interacts with the customer through a visual presentation device, during which the microphone and camera capture audio and video information. At this stage, audio signals and visual images are inputs, and they are output as acquired data.

[0668] Step 2:

[0669] The server receives the audio data transmitted from the terminal and converts it into text data using speech recognition software. AI technology analyzes the audio waveform and converts it into words and sentences. This converted text data is the output.

[0670] Step 3:

[0671] The server analyzes person information based on the received video data. Using a face recognition engine, it extracts facial features from the video and estimates the emotional state based on this. The input is video data, and the output is an emotional state such as "joy" or "surprise."

[0672] Step 4:

[0673] The server analyzes video data using gesture recognition software to extract movement information and interpret the user's intentions. It extracts features such as hand and body movements and infers intentions based on them. The input for this step is video data, and the output is the interpreted intention information.

[0674] Step 5:

[0675] The server integrates emotional state and interpreted intention information to generate an emotion recognition result, and then uses a generative AI model to generate a recommended response. Prompts are used to allow the AI ​​to create a natural, emotion-based response. The input is the emotion recognition result and intention information, and the output is the generated recommended response.

[0676] Step 6:

[0677] The server generates a recommended response, which is sent to the terminal and displayed on the visual display device. The user then uses this to facilitate smooth interaction with the customer. The input is the recommended response, and the output is the content displayed on the visual display device.

[0678] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0679] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0680] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0681] [Fourth Embodiment]

[0682] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0683] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0684] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0685] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0686] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0687] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0688] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0689] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0690] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0691] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0692] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0693] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0694] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0695] The system for implementing the present invention constructs a personal assistant as a program to facilitate smooth communication with the user. This system analyzes a large amount of data obtained from the user's voice and video and generates appropriate responses, thereby enabling aphasic patients to engage in natural conversations with others.

[0696] System Configuration

[0697] The device is equipped with a microphone and camera, which capture the user's voice signals and video in real time. This allows for detailed capture of the user's speech, facial expressions, and gestures.

[0698] The server is equipped with a speech recognition module to analyze the acquired audio data and instantly convert it into text data. The speech recognition technology utilizes the latest AI, enabling accurate extraction of meaning even if the user's speech is unclear.

[0699] The video data is processed on the server using face recognition and gesture recognition algorithms. The face recognition module analyzes the user's facial expressions and features to estimate their emotions. The gesture recognition module understands the content of gestures and sign language from the user's hand movements and overall body movements.

[0700] The server integrates the analyzed audio-text data with gesture information obtained from the video, and uses generative AI to generate natural and contextual responses. The generated responses are output as text data.

[0701] The text data of the response is converted into a speech signal using a speech synthesis engine. This synthesized speech is played back at a speed and tone that is easy for the user to understand, enabling natural conversation.

[0702] Specific example

[0703] The user makes a mouth gesture indicating they want to drink tea.

[0704] The device captures voice and gestures and sends them to the server.

[0705] The server converts the speech into text and generates the response, "Would you like some tea?"

[0706] The server synthesizes the generated response into speech, and the terminal plays the audio for the user.

[0707] In this way, this system can provide appropriate responses based on various nonverbal information provided by the user, making it possible to maintain and strengthen social connections for aphasia patients and provide them with new enjoyment and a sense of security in their daily lives.

[0708] The following describes the processing flow.

[0709] Step 1:

[0710] To initiate communication, the user speaks to the device or makes gestures.

[0711] Step 2:

[0712] The device acquires user voice data via the microphone and video data via the camera. This data is transmitted to the server in real time.

[0713] Step 3:

[0714] The server inputs the received audio data into a speech recognition engine, which converts the audio into text data. At this stage, it uses the latest speech analysis technology to understand the context even if the speech is unclear.

[0715] Step 4:

[0716] The server extracts facial information from video data and estimates emotional states using a facial recognition algorithm. Simultaneously, it analyzes the user's hand and body movements using a gesture recognition algorithm to understand their intentions and requests.

[0717] Step 5:

[0718] The server integrates voice-to-text and gesture information and uses generative AI to generate appropriate and natural responses. This makes it possible to create conversations that match the user's intentions.

[0719] Step 6:

[0720] The generated response is sent from the server to the terminal and stored as text data. This text data is then converted into speech by a speech synthesis engine.

[0721] Step 7:

[0722] The device plays synthesized speech data to the user through its speaker and displays it as text on the screen when visual feedback is needed. This feedback allows the user to recognize the system's response and take the next step in communication.

[0723] (Example 1)

[0724] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0725] When users communicate smoothly and naturally with others using voice and gestures, their speech may be unclear or their intentions may not be properly conveyed. Furthermore, information security and privacy protection are also important issues. In this context, there is a need to develop systems that enable people with aphasia and other speech disorders to connect more smoothly with society and reduce the difficulties they face in daily life.

[0726] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0727] In this invention, the server includes means for acquiring audio and video information, means for converting acquired audio information into text information, means for analyzing facial information from video information and estimating emotional states, means for analyzing gestures and inferring the user's intentions, means for generating a response based on the text information and the user's intentions, means for converting the generated response into audio information, and means for providing audio or text information to the user. This makes it possible to correctly understand the user's ambiguous utterances and nonverbal information, realize natural dialogue, and protect audio and video information.

[0728] A "device for acquiring audio and video information" is a device used to record the user's voice and body movements in real time and to process that data.

[0729] A "device that converts acquired audio information into text information" is a processing device that analyzes audio data and converts its content into an appropriate text format.

[0730] A "device that analyzes facial information from video data to estimate emotional state" is an analytical device that extracts the user's facial features from acquired video data and infers emotions based on those facial expressions.

[0731] A "device that analyzes gestures and infers the user's intentions" is a device that interprets the movements of a user's hands and body from video data and identifies the intentions that those movements convey.

[0732] A "device that generates responses based on textual information and user intent" is a device that creates an appropriate response based on transcribed audio data and the user's nonverbal expressions of intent.

[0733] A "device that converts generated responses into audio information" is a device that synthesizes responses created in text into a format that can be played back as audio.

[0734] A "device that provides audio or textual information to a user" is a device used to present a generated response to a user through auditory or visual means.

[0735] A "device that performs encryption and access restriction" is a system that protects data to ensure the privacy of audio and video information and prevent unauthorized access.

[0736] This invention is a system for supporting interaction with users, and in particular, provides a mechanism for achieving natural communication that does not rely on voice or gestures. The core function of the system is to utilize voice and video data.

[0737] The device is responsible for acquiring user audio and video data in real time using a microphone and camera. This makes it possible to capture not only the user's speech but also non-verbal information such as facial expressions and gestures. The data acquired from the device is transmitted to a server via the network.

[0738] The server uses speech recognition software to convert audio data into text data. This speech recognition utilizes the latest AI technology, specifically generative AI models, which can accurately understand and transcribe even unclear speech. The server also runs facial recognition and gesture recognition algorithms to analyze video data and infer the user's emotional state and intentions.

[0739] Based on the analysis results, the server uses generative AI to generate appropriate responses for the user. This response generation process is configured with prompts designed to deeply understand the user's intent. By setting prompts such as, "If the user is asking for a drink, please make the best suggestion," the system's responses become accurate and natural.

[0740] The generated text-based response is converted back into audio data via speech synthesis software. This audio is then played back to the user through the device, resulting in a more engaging conversation compared to text-based feedback.

[0741] For example, a user might say "I want some tea" aloud or indicate their intention with a gesture. In this case, the device sends this data to the server, which generates and speaks an appropriate response, such as "Would you like some tea?", and plays it back to the user, thereby achieving direct and natural interaction.

[0742] This system provides a comprehensive solution for users to communicate smoothly using voice and gestures. The convenience offered by the invention will support a better everyday communication experience.

[0743] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0744] Step 1:

[0745] The device uses a microphone and camera to acquire audio and video data from the user in real time. This input data includes the user's speech and nonverbal gestures. The device converts this data into a digital format and prepares it for transmission over the network to the server.

[0746] Step 2:

[0747] The server receives audio data transmitted from the terminal. Using the received audio data as input, the speech recognition software on the server converts the audio into text data using a generative AI model. In doing so, it removes noise and analyzes the audio signal to obtain a meaningful string of characters.

[0748] Step 3:

[0749] The server processes video data as input. A facial recognition algorithm extracts facial features from the video data and estimates the user's emotional state based on their facial expressions. Simultaneously, a gesture recognition algorithm analyzes hand and body movements and identifies the intentions behind the gestures. Through these processes, the server interprets the user's nonverbal information.

[0750] Step 4:

[0751] The server integrates text data obtained from audio with emotion and gesture information from video. Using generative AI, it generates contextually appropriate responses based on prompts. Specifically, it outputs an appropriate response to prompts such as, "If the user is asking for a drink, please make the best suggestion." The response generated in this step is in text format.

[0752] Step 5:

[0753] The server passes the generated text response to the speech synthesis engine, which converts it into speech data. During this process, the speed and tone are adjusted to enhance the naturalness of the synthesized speech. The synthesized speech data is then output to the user in a playable format.

[0754] Step 6:

[0755] The terminal receives audio data transmitted from the server and plays it back to the user. Using the speaker, it provides the user with a synthesized speech response, completing two-way communication. The played audio is adjusted to a volume and quality that is easy for the user to understand.

[0756] (Application Example 1)

[0757] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0758] Modern industrial activities require smooth communication with all users. In particular, providing appropriate communication even to users who have difficulty communicating is a social demand and an important challenge. This invention aims to facilitate dialogue with such diverse users and provide a more natural experience in customer service and support in real-world settings.

[0759] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0760] In this invention, the server includes means for acquiring data from a user, means for converting the data into text information, and means for utilizing the proposed response in dialogue in industrial activities. This facilitates smooth dialogue with all users and enables appropriate and natural dialogue even in industrial settings.

[0761] "Means of acquiring data from users" refers to devices or processes that collect users' voice and video in real time.

[0762] "Means for converting the data into text information" refers to a technology or process for converting an audio signal into text format using speech recognition technology.

[0763] "Means for extracting identification information and estimating a state" refers to technologies that recognize a user's face and emotions from video data and determine their psychological or physical state.

[0764] "Means for analyzing motion information and interpreting user intentions" refers to technologies that analyze user gestures and hand movements and understand the user's intentions based on them.

[0765] "Means for generating responses" refers to a process or algorithm for automatically generating appropriate responses for the user based on acquired data.

[0766] "Means for utilizing proposed responses in dialogue in industrial activities" refers to technologies that apply generated responses to communication in multiple industrial scenarios in real society.

[0767] This invention is a system for facilitating smooth communication with users who have difficulty communicating in industrial settings. This system includes a process for collecting and analyzing audio and video data in real time. The main components of this system and their operation are described below.

[0768] The server receives audio and video data collected from users and converts it into text using speech recognition technology. By utilizing a highly accurate AI model for speech recognition, accurate text transcription is possible even when the user's speech is unclear. For example, Google Cloud Speech-to-Text can be used for speech recognition, and OpenAI GPT-based models can be used for AI.

[0769] Next, the server uses facial recognition technology to estimate emotions and intentions from the video data. Libraries such as OpenCV are utilized for facial recognition and gesture analysis, analyzing the user's facial expressions and movements. Based on these results, the movements are interpreted to understand the user's underlying intentions.

[0770] Next, the server generates an appropriate response for the user through a generative AI model. This response is then adapted for natural communication specific to the industry. The generated response is then presented to the user as natural speech using a speech synthesis engine such as Amazon Polly.

[0771] As a concrete example, consider a scenario in a store where a customer asks, "What does this product taste like?" The customer's actions and voice are collected by sensors in smart glasses and sent to a server. The server analyzes this data and generates a response such as, "This tea has a matcha flavor and is very mild," which is then made available for staff to hear directly.

[0772] An example of a prompt is given as follows: "The customer is pointing to a product and saying, 'What does this product taste like?' Based on this information, the generating AI should respond with a product description." This approach allows for accurate understanding of the user's intent and the provision of information in a natural manner.

[0773] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0774] Step 1:

[0775] The user makes requests using voice and gestures. The device's microphone and camera capture this information in real time. Specifically, the camera captures the user's voice and actions such as pointing.

[0776] Step 2:

[0777] The terminal transmits the acquired audio and video data to the server. During this process, the data is compressed as needed for efficient transmission. The input consists of audio signals and video frames, while the output is the transmitted digital data.

[0778] Step 3:

[0779] The server uses a speech recognition module to analyze the audio data and convert it into text. This process analyzes the phonemes within the audio and converts them into text. For example, the audio "What does this product taste like?" is converted into text.

[0780] Step 4:

[0781] The server analyzes the video data and extracts the user's facial features and gesture information. It processes the video frames and recognizes the user's facial expressions and movements. Using a facial recognition algorithm, it can recognize actions such as pointing.

[0782] Step 5:

[0783] The server uses a generative AI model to generate an appropriate response based on the speech recognition and gesture recognition results. Using the "generative AI model and prompt sentence," for example, a response such as "This tea has a matcha flavor and is very mild." is generated.

[0784] Step 6:

[0785] The generated response is converted into audio data and sent to the terminal. A speech synthesis engine is used to convert the text information into a natural-sounding audio signal. This audio signal is then returned to the terminal.

[0786] Step 7:

[0787] The device uses the played audio to present responses to the user. This allows the user to receive information in a natural conversational format. As a result, the user finds it easier to understand, and the conversation proceeds smoothly.

[0788] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0789] This invention supports smooth communication for aphasia patients by using a system that incorporates an emotion engine that recognizes the user's emotions. This system generates responses that correspond to the user's emotions during dialogue, thereby achieving natural and realistic communication.

[0790] System Configuration

[0791] The device is equipped with a microphone and camera to collect the user's voice signals and video data. This allows it to acquire non-verbal information such as the user's facial expressions and body movements.

[0792] The server processes the acquired audio data using a speech recognition engine and converts it into text data. Advanced AI technology is used for speech recognition to extract appropriate context even in situations where vocalization is difficult.

[0793] The video data is processed on the server using face recognition and gesture recognition algorithms. The face recognition module analyzes the user's facial expressions and works with the emotion engine to determine their emotional state. The gesture recognition module recognizes the user's intentions from their hand and body movements.

[0794] The emotion engine integrates facial and gesture information to more accurately estimate the user's emotions. This estimation result is then reflected in the generated responses, providing responses that align with the user's emotions.

[0795] The server uses generative AI to create natural-sounding responses based on text data, gesture information, and emotion recognition results. These responses are then synthesized into audio data using speech synthesis technology and transmitted to the user.

[0796] Specific example

[0797] The user smiles and says, "I'm happy today."

[0798] The device captures audio and video data and sends it to the server.

[0799] The server converts the audio into text, and the emotion engine determines that a smile on the face indicates "joy."

[0800] The server generates and speaks a joyful response saying, "Today is a wonderful day."

[0801] The device presents the generated audio to the user, facilitating dialogue that allows for the sharing of emotions.

[0802] This system enables dialogue that is sensitive to the user's emotions, allowing aphasia patients to deepen their social connections. By responding quickly to changes in emotions, it is expected to provide more harmonious communication and improve the user's quality of life.

[0803] The following describes the processing flow.

[0804] Step 1:

[0805] The user initiates communication by speaking or making gestures towards the device.

[0806] Step 2:

[0807] The device uses a microphone and camera to collect audio and video data in real time. This data is sent to a server for processing.

[0808] Step 3:

[0809] The server activates its speech recognition engine and converts the incoming audio data into text data. During this process, it also analyzes the intonation and tone of speech to aim for more accurate text output.

[0810] Step 4:

[0811] The server analyzes the video data and uses a facial recognition algorithm to analyze the user's facial expressions. This estimates the user's emotional state, and the emotion engine determines that emotion.

[0812] Step 5:

[0813] Simultaneously, the server performs gesture recognition to analyze the user's hand and body movements, interpreting the user's intentions and requests from their actions. Based on this information, it prepares a more contextually appropriate response.

[0814] Step 6:

[0815] The emotion engine integrates facial and gesture information to determine the user's emotion and sends the result to the server. This becomes a crucial element in the content of the generated response.

[0816] Step 7:

[0817] The server uses generative AI to generate natural, emotion-optimized responses by integrating text, gesture information, and recognized emotions from the speech.

[0818] Step 8:

[0819] The generated text response is converted into speech data through a speech synthesis engine.

[0820] Step 9:

[0821] The device plays the generated voice response to the user and, if necessary, displays a text response on the screen to visually supplement it. This enables smooth communication with the user.

[0822] (Example 2)

[0823] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0824] The goal is to provide a system that enables users with aphasia or other conditions that make verbal communication difficult to engage in natural, emotionally charged conversations. Furthermore, it requires the system to more accurately recognize the user's complex emotions and intentions and generate appropriate responses based on that understanding. Ensuring the security of audio and video data is also a key challenge.

[0825] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0826] In this invention, the server includes means for acquiring audio and video data from the user, means for converting the audio data into text data, and means for integrating facial information and motion information to estimate the user's emotions. This makes it possible to accurately grasp the user's emotions and provide natural dialogue that corresponds to them.

[0827] A "user" refers to an individual who interacts with a system, provides audio and video data, and communicates with it.

[0828] "Audio data" refers to digital signals that include sound information such as the user's voice and tone.

[0829] "Video data" refers to digital video that includes visual information such as the user's facial expressions and movements.

[0830] "Text data" refers to data that represents the result of converting audio data into text information.

[0831] "Facial features" refers to information extracted from video data to identify the user's facial features and expressions.

[0832] "Motion information" refers to the results of analyzing the user's hand movements and body movements.

[0833] "Emotional state" refers to the state of mind determined from the user's facial expressions and actions.

[0834] "Intention" refers to the purpose or will interpreted based on the user's actions.

[0835] A "generative AI model" refers to an algorithm that uses artificial intelligence to generate natural language responses.

[0836] A "prompt sentence" refers to a sentence that, when input into a generative AI model, prompts the model to generate an appropriate response.

[0837] "Means of converting to audio data" refers to technologies or devices that have the function of converting text data into audio information.

[0838] "Protection" refers to security measures that prevent the unauthorized use or alteration of audio and video data.

[0839] This invention provides a system that supports communication, particularly for patients with aphasia, by recognizing the user's emotions and generating a corresponding response. The following describes embodiments for carrying out the invention.

[0840] First, the device is equipped with a microphone and camera to collect the user's voice, facial expressions, and gestures. This allows for the acquisition of both audio and video data from the user. The audio data includes the user's speech, and its intonation and tone can also be considered. The video data includes non-verbal information such as the user's facial expressions and hand movements.

[0841] Next, the server converts the received audio data into text data using speech recognition software (for example, a general-purpose speech recognition engine). This conversion utilizes advanced speech processing technology to accurately convert even unclear speech into text.

[0842] Furthermore, the server uses facial recognition technology (e.g., a basic facial recognition algorithm) to analyze facial features in the video data and determine the user's emotional state. It also uses gesture recognition technology to analyze the user's movements and determine their intentions.

[0843] Based on this information, the server's emotion engine integrates facial and motion data to estimate the user's emotions. Using this estimation, the server inputs prompt sentences into a generative AI model (for example, a general natural language generation algorithm) to generate a natural response.

[0844] As a concrete example, consider a case where a user smiles and says, "I'm happy today." In this case, the device sends audio and video data to the server. After converting the audio to text, the server uses an emotion engine to determine that the emotion is "joy." An example of a prompt is: "The user's emotion is joy. Please generate an appropriate message for a happy day." Based on this prompt, the generative AI model generates the response, "Today is a wonderful day."

[0845] Subsequently, the generated text responses are converted into audio data by a speech synthesis engine (e.g., basic speech synthesis technology) and provided to the user through the terminal. This process enables users to engage in natural, emotionally charged conversations, thereby improving the quality of communication.

[0846] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0847] Step 1:

[0848] The user speaks aloud to the system and simultaneously makes facial expressions and gestures in front of the camera. This generates audio and video data.

[0849] Step 2:

[0850] The device captures audio data from the user using a microphone and simultaneously captures video data using a camera. The input for this step is the user's raw audio and visual information, and the output is digital audio and video data.

[0851] Step 3:

[0852] The terminal sends the acquired audio and video data to the server. At this stage, the terminal performs the data transfer and prepares to wait for processing on the server side.

[0853] Step 4:

[0854] The server inputs the received audio data into the speech recognition engine and converts it into text data. The input is digital audio data, and the output is text data. Specifically, it performs frequency analysis of the audio and replaces the audio patterns with characters.

[0855] Step 5:

[0856] The server uses a facial recognition module based on video data to analyze the user's facial expressions. The input is video data, and the output is the emotion estimation result based on the facial expressions. The process involves extracting facial feature points and matching them against known emotion patterns.

[0857] Step 6:

[0858] The server analyzes user movement information using a gesture recognition module and determines the user's intention. Video data is input, and the user's intentions based on their movements are output. Specifically, the trajectory of hand movements is tracked, and corresponding movement labels are assigned.

[0859] Step 7:

[0860] The server inputs the results of face recognition and gesture recognition into an emotion engine that estimates the overall emotional state. The input consists of individual emotion estimates and intention information from actions, and the output is an integrated emotion estimate. As an operation, statistical estimation is performed that takes into account different emotional elements.

[0861] Step 8:

[0862] The server inputs a prompt sentence into a generative AI model and generates an appropriate response. The input consists of text data and sentiment estimation results, and the generated response text is output. Specifically, it is a process in which the AI ​​understands the context based on the prompt sentence and generates a new response.

[0863] Step 9:

[0864] The server inputs the response text into a speech synthesis engine and converts it into speech data. The generated response text is then output as natural-sounding speech data. This process involves synthesizing speech waveforms to produce fluent speech.

[0865] Step 10:

[0866] The device presents the generated audio data to the user. The user receives the generated response through the audio emitted from the device. In this step, conversation becomes possible through the audio output device.

[0867] (Application Example 2)

[0868] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0869] Modern brick-and-mortar stores demand quick and appropriate communication with customers. However, it is difficult for employees to instantly understand the diverse emotions of customers and provide appropriate service. Therefore, improving customer satisfaction is a challenge.

[0870] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0871] In this invention, the server includes a device for acquiring audio and video data from a user, a device for converting the audio data into text data, a device for extracting person information and estimating emotional state, a device for analyzing behavioral information and interpreting the customer's intentions, and a device for displaying the emotion recognition results and recommended responses on a wearable visual display device. This enables store employees to grasp the customer's emotions in real time and provide customer service that is appropriate to those emotions.

[0872] "Audio data" refers to sound information obtained from the user, including conversations and ambient sounds.

[0873] "Video data" refers to visual information acquired using cameras or other devices, such as recordings of the user's face and movements.

[0874] "Personal information" refers to information extracted from video data, such as facial features and expressions, used to identify an individual and estimate their condition.

[0875] "Emotional state" refers to the result of estimating the user's emotions based on personal information, and includes emotions such as joy and anger.

[0876] "Motion information" refers to information obtained by analyzing the user's gestures and body movements, used to interpret their intentions and will.

[0877] A "visual presentation device" is a device that can be worn by an individual and is used to visually present video information to the user.

[0878] "Emotion recognition results" refer to the estimated emotions obtained by analyzing emotional states, behavioral information, and other factors.

[0879] A "recommended response" is a suggested response generated based on emotion recognition results, designed to facilitate smoother interactions with the user.

[0880] The system that realizes this invention is configured as follows to support interaction with customers. First, a visual presentation device that can be worn by the user acquires audio and video. This device is equipped with a microphone and a camera and collects audio and video data in real time.

[0881] On the server, AI-based speech recognition software converts the collected audio data into text data. Video data is analyzed to extract person information, and facial recognition technology is used to estimate emotional states. In this process, emotions are analyzed based on the user's facial expressions and movements, and movement information is extracted simultaneously. The emotion recognition engine integrates the facial and movement information to generate emotion recognition results, more accurately estimating the user's emotions.

[0882] Based on these results, the generative AI model automatically generates appropriate recommended responses. These responses are displayed on a visual presentation device and shown to the user along with the emotion recognition results. This device allows the user to smoothly engage in conversations with customers.

[0883] For example, if a customer is smiling while looking at a particular product in a store, the server analyzes the image and recognizes it as "joy." The visual display device then presents the staff with a recommendation, such as "We highly recommend that product."

[0884] An example of a prompt message is, "The customer is speaking with a smile. Please generate a positive comment."

[0885] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0886] Step 1:

[0887] The terminal acquires audio and video data. The user interacts with the customer through a visual presentation device, during which the microphone and camera capture audio and video information. At this stage, audio signals and visual images are inputs, and they are output as acquired data.

[0888] Step 2:

[0889] The server receives the audio data transmitted from the terminal and converts it into text data using speech recognition software. AI technology analyzes the audio waveform and converts it into words and sentences. This converted text data is the output.

[0890] Step 3:

[0891] The server analyzes person information based on the received video data. Using a face recognition engine, it extracts facial features from the video and estimates the emotional state based on this. The input is video data, and the output is an emotional state such as "joy" or "surprise."

[0892] Step 4:

[0893] The server analyzes video data using gesture recognition software to extract movement information and interpret the user's intentions. It extracts features such as hand and body movements and infers intentions based on them. The input for this step is video data, and the output is the interpreted intention information.

[0894] Step 5:

[0895] The server integrates emotional state and interpreted intention information to generate an emotion recognition result, and then uses a generative AI model to generate a recommended response. Prompts are used to allow the AI ​​to create a natural, emotion-based response. The input is the emotion recognition result and intention information, and the output is the generated recommended response.

[0896] Step 6:

[0897] The server generates a recommended response, which is sent to the terminal and displayed on the visual display device. The user then uses this to facilitate smooth interaction with the customer. The input is the recommended response, and the output is the content displayed on the visual display device.

[0898] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0899] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0900] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0901] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0902] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0903] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0904] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0905] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0906] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0907] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0908] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0909] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0910] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0911] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0912] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0913] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0914] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0915] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0916] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0917] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0918] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0919] The following is further disclosed regarding the embodiments described above.

[0920] (Claim 1)

[0921] A means of obtaining audio and video data from the user,

[0922] Means for converting the aforementioned audio data into text data,

[0923] A means for extracting facial information from the aforementioned video data and estimating emotional state,

[0924] A means for analyzing gesture information from the aforementioned video data and interpreting the user's intentions,

[0925] Means for generating a response based on the aforementioned text data and the user's intent,

[0926] means for converting the aforementioned response into audio data,

[0927] Means for presenting the aforementioned audio data or text data to the user,

[0928] A system that includes this.

[0929] (Claim 2)

[0930] The system according to claim 1, wherein the means for generating the response constructs a natural dialogue based on the generated text data.

[0931] (Claim 3)

[0932] The system according to claim 1, further comprising means for encrypting and accessing the aforementioned audio and video data to ensure its security.

[0933] "Example 1"

[0934] (Claim 1)

[0935] A device for acquiring audio and video information,

[0936] A device that converts acquired audio information into text information,

[0937] A device that analyzes facial information from video data to estimate emotional states,

[0938] A device that analyzes gestures and infers the user's intentions,

[0939] A device that generates a response based on textual information and the user's intent,

[0940] A device that converts the generated response into audio information,

[0941] A device that provides audio or text information to the user,

[0942] A system that includes this.

[0943] (Claim 2)

[0944] The system according to claim 1, comprising a device that generates responses that form a natural dialogue based on textual information.

[0945] (Claim 3)

[0946] The system according to claim 1, comprising a device for encrypting and restricting access to audio and video information for the purpose of protecting that information.

[0947] "Application Example 1"

[0948] (Claim 1)

[0949] Means of obtaining data from users,

[0950] Means for converting the aforementioned data into character information,

[0951] A means for extracting identification information from the aforementioned data and estimating the state,

[0952] A means for analyzing the operation information from the aforementioned data and interpreting the user's intentions,

[0953] Means for generating a response based on the aforementioned textual information and the user's intent,

[0954] Means for converting the aforementioned response into data,

[0955] Means for presenting the aforementioned data to the user,

[0956] A system that includes means for utilizing proposed responses in dialogue within industrial activities.

[0957] (Claim 2)

[0958] The system according to claim 1, wherein the means for generating the response constructs a structured dialogue based on the generated character information.

[0959] (Claim 3)

[0960] The system according to claim 1, comprising means for performing information processing and management to ensure the protection of the aforementioned data.

[0961] "Example 2 of combining an emotion engine"

[0962] (Claim 1)

[0963] A means of obtaining audio and video data from the user,

[0964] Means for converting the aforementioned audio data into text data,

[0965] A means for extracting facial features from the aforementioned video data and determining the emotional state,

[0966] A means for analyzing motion information from the aforementioned video data and determining the user's intent,

[0967] A means of integrating facial information and motion information to estimate the user's emotions,

[0968] A means for generating a response using a generative AI model based on the aforementioned text data, user intent, and sentiment estimation results,

[0969] means for converting the generated response into audio data,

[0970] Means for presenting the aforementioned voice data or response data to the user,

[0971] A system that includes this.

[0972] (Claim 2)

[0973] The system according to claim 1, comprising means for constructing a natural dialogue based on prompt sentences using a generative AI model.

[0974] (Claim 3)

[0975] The system according to claim 1, further comprising means for encryption and access control to ensure the protection of the aforementioned audio and video data.

[0976] "Application example 2 when combining with an emotional engine"

[0977] (Claim 1)

[0978] A device that acquires audio and video data from the user,

[0979] A device for converting the aforementioned audio data into text data,

[0980] A device for extracting person information from the aforementioned video data and estimating their emotional state,

[0981] A device that analyzes motion information from the aforementioned video data and interprets the user's intentions,

[0982] A device that generates a response based on the aforementioned text data and the user's intentions,

[0983] A device that converts the aforementioned response into audio data,

[0984] A device that presents the aforementioned audio data or text data to the user,

[0985] A device that displays emotion recognition results and recommended responses on a wearable visual presentation device,

[0986] A system that includes this.

[0987] (Claim 2)

[0988] The system according to claim 1, wherein the device that generates the response constructs a natural dialogue based on the generated text data and displays an emotion-based response on the visual presentation device.

[0989] (Claim 3)

[0990] The system according to claim 1, comprising a device that performs encryption and access control to ensure the security of the aforementioned audio data and video data, and further ensuring the security of the displayed content to the visual display device through personal authentication. [Explanation of Symbols]

[0991] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of obtaining audio and video data from the user, Means for converting the aforementioned audio data into text data, A means for extracting facial information from the aforementioned video data and estimating emotional state, A means for analyzing gesture information from the aforementioned video data and interpreting the user's intentions, Means for generating a response based on the aforementioned text data and the user's intent, means for converting the aforementioned response into audio data, Means for presenting the aforementioned audio data or text data to the user, A system that includes this.

2. The system according to claim 1, wherein the means for generating the response constructs a natural dialogue based on the generated text data.

3. The system according to claim 1, further comprising means for encrypting and accessing the aforementioned audio and video data to ensure its security.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A