system

A system that analyzes voice and images to understand the needs of the elderly, generating personalized responses and incorporating feedback, addresses the challenge of providing continuous and flexible support, reducing isolation and caregiver burden.

JP2026070294APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing systems fail to provide continuous and flexible support for the elderly, especially in understanding their changing needs and incorporating user feedback to improve services, leading to isolation and increased caregiver burden.

Method used

A system that analyzes voice and images to identify user needs, generates appropriate responses, collects feedback, and iteratively optimizes support, using speech and image recognition technologies with an emotion engine to provide personalized and context-aware interactions.

Benefits of technology

Enables flexible and continuous support for the elderly, reducing isolation and caregiver burden by accurately understanding user needs and emotions, and improving service quality through feedback integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070294000001_ABST
    Figure 2026070294000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of analyzing the user's state using voice and images and identifying their needs, A means of generating and providing responses to users based on identified needs, A means of collecting user feedback after providing a service and incorporating it into the next service, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In an aging society, it is required to support the daily lives of the elderly, promote their independence, and reduce the burden on caregivers. In particular, for the elderly to live a comfortable life without feeling isolated, continuous communication and appropriate support are important. Furthermore, since the needs of the elderly sometimes change, a support system that can flexibly respond to them is necessary.

Means for Solving the Problems

[0005] This invention provides a system that analyzes the voice and images of elderly individuals to understand their condition and identify their individual needs. This system automatically generates appropriate responses based on the identified needs and provides them to the user. Furthermore, it collects feedback from users after service provision and incorporates it into subsequent services, enabling continuously optimized support. In this way, it becomes possible to flexibly respond to the changing needs of users.

[0006] "Voice" refers to words and sounds uttered by the user's mouth, and is a means of transmitting information through them.

[0007] An "image" is visual information acquired by a camera or other photographic device, and includes nonverbal communication such as the user's facial expressions and gestures.

[0008] "Analysis" is the act of using collected data to analyze its content and characteristics, thereby clarifying the user's state and needs.

[0009] "Needs" refer to the user's requests, wishes, and the type of support they require, and these change depending on their condition and circumstances.

[0010] A "response" refers to the information, advice, or instructions that a system provides in response to a user's needs, and is a means of establishing communication.

[0011] "Feedback" refers to the opinions and evaluations that users provide regarding the services offered by the system, and the quality of the service is improved based on this feedback.

[0012] A "system" is a collective term for a set of devices and programs designed to analyze voice and images and provide optimal support to the user. [Brief explanation of the drawing]

[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is an intelligent system for supporting the lives of the elderly, enabling communication with users through voice and images. It consists of three main components: a server, a terminal (robot), and a user, and is implemented in the following manner.

[0035] The server, located in the cloud, is responsible for advanced processing of speech recognition and image analysis. It converts the user's spoken audio data into text using speech recognition technology, and then analyzes the user's intent and emotional state based on that text data. It also recognizes the user's facial expressions and environmental elements from image data to understand the context. This allows the server to identify the user's needs and generate responses appropriate to those needs.

[0036] The terminal is placed in the living space of the elderly and functions as a user interface. This terminal is equipped with a camera and microphone, which acquire and transmit audio and image data to a server. Furthermore, it can present responses from the server to the user via speech synthesis and provide physical assistance as needed.

[0037] Specifically, when a user asks a question about cooking or everyday life, they speak to the robot via voice, and the robot receives the voice and transmits it to a server. The server converts the voice data into text and generates a contextually appropriate response. For example, if a user says, "I want to make soup," the robot will reply by voice, "What ingredients do you have?" Based on the user's response, the robot then receives more detailed support, such as recipe information, from the server and provides it to the user.

[0038] In this way, this system qualitatively supports the lives of the elderly and promotes independent living, thereby contributing to a reduction in the burden of caregiving. Furthermore, it provides users with an environment where they can live at their own pace with peace of mind through interaction with robots.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The device acquires voice input from the user via a microphone and simultaneously captures image data of the user's facial expressions and environment using a camera. This collected data is then transmitted to the server in real time.

[0042] Step 2:

[0043] The server converts the received audio data into text data using speech recognition technology. During this process, it analyzes the user's emotional state based on the tone and speed of their voice, using this information as foundational data to identify their needs.

[0044] Step 3:

[0045] The server processes the acquired image data using image analysis technology to analyze the user's facial expressions and gestures. Based on this information, it helps understand the user's context and identify their needs.

[0046] Step 4:

[0047] The server integrates information obtained from voice and images to identify the user's needs. Based on the identified needs, it internally determines the appropriate response and support method.

[0048] Step 5:

[0049] The server uses natural language generation technology to create a response that is easy for the user to understand, based on the determined response content. This generated response is then sent to the terminal.

[0050] Step 6:

[0051] The terminal presents the received response to the user as audio using speech synthesis technology. This allows the user to hear and understand the robot's response.

[0052] Step 7:

[0053] The user acts on the robot's responses and instructions, and asks additional questions or requests if necessary. This information is collected by the terminal as feedback that will be used again in the next cycle.

[0054] Step 8:

[0055] The device then sends the feedback information back to the server, which uses this feedback to adjust the system's response or as data to improve the AI ​​model. This process improves the accuracy of the service.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] There is a need for intelligent support systems that enable elderly people to live independently while gaining a sense of security. Conventional systems rely solely on audio or image information, which has the problem of not being able to fully understand the user's intentions and needs. Furthermore, there has been insufficient mechanism for accurately incorporating user feedback into future services.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes means for analyzing the user's state using voice and images and identifying the user's needs; means for generating and providing a response to the user based on the identified needs; and means for recognizing voice and converting it to text and generating a context-based response. This makes it possible to more accurately understand the user's intentions and emotions and to engage in appropriate dialogue. Furthermore, it enables the efficient collection of feedback and its incorporation into future services, thereby providing sustainable support.

[0061] "Speech recognition technology" is a technology that converts speech data into text data, and it extracts linguistic information by analyzing speech signals.

[0062] "Image recognition technology" is a technology that analyzes image data to recognize specific patterns or objects and makes decisions based on that information.

[0063] "Speech synthesis technology" is a technology that generates natural-sounding speech from text data, producing audio with a texture similar to human speech.

[0064] "Feedback collection" is the process of gathering opinions and reactions from users and using them to improve future services.

[0065] "Contextual understanding" is the process of grasping the background and intent of a conversation based on information obtained from audio and images, and deriving an appropriate response.

[0066] This invention provides an intelligent system to support the daily lives of elderly people. The system consists of three main components: a server, a terminal, and a user. Details are described below.

[0067] server:

[0068] The servers are located in a cloud environment and handle speech recognition and image analysis processing. Speech recognition uses technology to convert audio data into text, and this process utilizes common speech recognition software. For example, technologies such as Google® Cloud Speech-to-Text and Amazon Transcribe are used. Image data is analyzed using image recognition technology to understand the user's facial expressions and environment. The servers then understand the user's intentions and needs and generate appropriate responses based on that information. The generated responses are delivered as natural-sounding speech using speech synthesis technology (for example, Amazon Polly or Google Cloud Text-to-Speech).

[0069] Terminal:

[0070] The terminal is placed within the user's living space and acquires audio and image data through its camera and microphone. This data is immediately transmitted to the server. A secure and efficient communication protocol is used to enable real-time data processing. The terminal also has the ability to play back responses from the server using speech synthesis technology and display relevant information on the display as needed.

[0071] Specific user examples:

[0072] When a user asks the robot on their device, "What should I make for dinner tonight?", the device sends the voice message to a server. The server transcribes the voice into text and analyzes the user's intent and emotions. For example, it can offer dinner ideas. The device then plays back the server's response and makes a suggestion to the user, such as, "How about pasta?" This specific dialogue supports the user in their daily activities.

[0073] Examples of generative AI models and prompt statements:

[0074] When using an AI model to generate responses in response to user requests, enter a prompt such as: "The user is asking for dinner ideas. Generate appropriate dish suggestions." Using specific prompts like this ensures that the generated responses align with the user's intent.

[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0076] Step 1:

[0077] The device acquires the user's voice and images using the camera and microphone. The input consists of audio signals and image data, and the output is the digital format of that data. This operation collects fundamental data about the user's intentions and state.

[0078] Step 2:

[0079] The terminal transmits the acquired audio and image data to the server. The input is digitized audio and image data, and the output is data transfer to a server in the cloud.

[0080] Step 3:

[0081] The server receives audio data and converts it into text data using speech recognition technology. The input is audio data, and the output is the corresponding text data. This process makes the user's speech interpretable as text.

[0082] Step 4:

[0083] The server analyzes text data and uses a generative AI model to analyze the user's intent and emotions. The input is text data, and the output is the analysis result. This analysis forms the basis for responses optimized to user needs.

[0084] Step 5:

[0085] The server analyzes image data to recognize the user's facial expressions and surrounding environment. The input is image data, and the output is the result of the recognition of the environment and state. This allows for an understanding of the context of the conversation.

[0086] Step 6:

[0087] The server generates appropriate responses using a generative AI model based on the analysis results of the audio and images. The input is the analyzed intent, emotion, and context, and the output is a specific response.

[0088] Step 7:

[0089] The server sends the generated response to the terminal. The input is the generated response, and the output is data in a format suitable for speech synthesis.

[0090] Step 8:

[0091] The device uses speech synthesis technology to play the generated response as sound to the user. The input is audio data received from the server, and the output is human-readable speech. This allows the user to receive support.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] In industrial settings, workers need to prepare necessary tools and materials quickly and accurately in order to carry out their work efficiently. However, if workers have to go and find the necessary tools themselves, not only will work efficiency decrease, but the risk of errors due to incorrect tool selection and safety risks will also increase. This invention aims to solve these problems and improve work efficiency and safety.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for analyzing the user's state using voice and images and identifying their needs; means for generating and providing a response to the user based on the identified needs; means for collecting feedback from the user after the service has been provided and incorporating it into the next service; means for analyzing the voice instructions of workers in an industrial setting and selecting appropriate tools and materials; and means for providing the selected tools and materials to the workplace. This automates the selection and supply of items to enable workers to perform their tasks efficiently, thereby improving work efficiency and safety.

[0097] "Audio and images" refers to audio and image data obtained from users, and serves as a source of information for analyzing the user's state and needs.

[0098] "Means of identifying needs" refers to the process of analyzing acquired audio and image data to clarify the requests and requirements that users have.

[0099] "Means for generating and providing responses to users" refers to a system for creating appropriate information and support based on identified needs and communicating them to users.

[0100] "Means for collecting feedback and reflecting it in future services" refers to the process of collecting evaluations and attitudes from users after a service has been provided, and using that information to improve future services.

[0101] "Means for analyzing voice instructions given by workers in industrial settings" refers to technologies and processes for understanding and analyzing the content of voice instructions given by workers in industrial settings.

[0102] "Means for selecting appropriate tools and materials" refers to the process of determining and selecting the necessary tools and materials based on the worker's voice instructions.

[0103] "Means of providing selected tools and materials to the workshop" refers to a system for delivering selected tools and materials to the workshop at the appropriate time.

[0104] The system implementing this invention aims to efficiently support work in industrial settings. The system understands the needs of workers through the analysis of voice and image data and provides the necessary tools and materials.

[0105] The server is located in the cloud and has advanced speech recognition and image analysis capabilities. Speech recognition APIs (e.g., Google Speech-to-Text) are used to convert speech data into text and determine the worker's intent. Furthermore, OpenCV is used for image recognition technology to accurately identify necessary tools and materials. Based on the acquired information, the server generates appropriate responses and provides them to the user.

[0106] The terminal is installed in the industrial site and functions as an interface to assist with work. This terminal is equipped with a camera and microphone, which acquire and transmit audio and image data to a server. It also uses speech synthesis to communicate responses from the server to the worker and provides physical assistance, such as transporting tools and materials, as needed.

[0107] The user, or worker, gives voice instructions to the system. Based on these instructions, the system analyzes the worker's requests and prepares and provides the necessary items. This process allows the worker to concentrate on their work, improving efficiency while also ensuring safety.

[0108] For example, if a worker says, "Bring me the wrench that fits this bolt," the server can identify the appropriate wrench, send an instruction to the terminal, and have that wrench delivered to the worker.

[0109] An example of a prompt message could be, "List the most frequently used tools in the factory and use that to optimize word selection for speech recognition." This would further improve the accuracy of voice commands and enable more efficient work support.

[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0111] Step 1:

[0112] The terminal receives voice instructions from the user (worker) via a microphone. This voice data becomes the input. The terminal temporarily stores this voice data in a buffer and prepares to send it to the server in the next step.

[0113] Step 2:

[0114] The terminal sends the audio data acquired in step 1 to the server. The server uses a speech recognition API to convert the audio data into text. This converted text is the output and is used as the basis for determining the worker's intent.

[0115] Step 3:

[0116] The server uses a generative AI model to analyze the worker's intent from the text data in step 2. This process understands the context of words and phrases and determines the necessary tools and materials. As a result of this analysis, information on the selected tools and materials is output.

[0117] Step 4:

[0118] The server uses image recognition technology to generate specific instructions for selecting particular tools and materials based on the information from step 3. Here, location information and identification tags for tools and materials are extracted from the image data, and the information of the selected items is organized.

[0119] Step 5:

[0120] The terminal receives the specific tool and material information generated in step 4 and provides feedback to the user via speech synthesis. It provides the worker with responses such as "I'll bring the XX wrench," building trust in the system.

[0121] Step 6:

[0122] The terminal actually transports the identified tools and materials based on the worker's instructions. In this process, the selected tools are moved safely and accurately by the mounted arm or transport device. Once everything is complete, the worker is notified and the terminal awaits further instructions.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention is a support system that enables interaction with users using voice and image data, and incorporates an emotion engine that recognizes the user's emotions. The system consists of three main elements: a server, a terminal (robot), and a user, and is implemented in the following manner.

[0125] The device is placed in the living space of elderly people and acquires voice and image data in real time. When a user speaks to the robot, the voice input is captured by the device's microphone. The device also captures the user's facial expressions and surrounding environment through its camera. This data is immediately transmitted to a server.

[0126] The server converts received audio data into text using speech recognition technology, while simultaneously analyzing the user's emotional state from their speech using an emotion engine. Furthermore, image data is processed using image recognition technology to analyze the user's facial expressions and gestures. In this way, the server integrates audio and image data to understand the context and identify the user's needs and emotions.

[0127] The emotion engine uses a multifaceted approach, including voice tone analysis and facial expression analysis, to determine the user's emotions and decide on a response based on those emotions. Based on these analysis results, the server generates an appropriate response and sends it to the terminal.

[0128] The terminal presents responses from the server to the user in voice using speech synthesis technology. The responses are tailored to the user's emotions and designed to provide a sense of reassurance. For example, if the user says, "I'm lonely today," the robot will gently respond, "Shall we talk?" to cheer the user up.

[0129] In this way, this system understands users' needs and emotions and provides personalized support, thereby creating an environment where seniors can live independent and happy lives. Furthermore, by utilizing emotional information, the content of subsequent service provision can be made more individualized, leading to continuous improvement in satisfaction.

[0130] The following describes the processing flow.

[0131] Step 1:

[0132] The device acquires voice input from the user via the microphone and captures image data of the user's facial expressions and environment using the camera. The acquired data is immediately sent to the server.

[0133] Step 2:

[0134] The server receives the audio data and converts it into text using speech recognition technology. During this process, it also analyzes the tone and speed of the speech to infer the user's emotional state.

[0135] Step 3:

[0136] The server processes image data using image analysis technology to analyze the user's facial expressions and gestures. This deepens the understanding of the user's emotions and situation.

[0137] Step 4:

[0138] The server inputs information obtained from voice and images into the emotion engine to identify the user's emotions. Based on this, it integrates the user's needs and emotions to determine a response.

[0139] Step 5:

[0140] The server uses natural language generation technology to create responses that match the user's emotions. The generated responses are designed to be polite, considerate, and thoughtful.

[0141] Step 6:

[0142] The server sends the generated response data to the terminal. The response is sent as text data, but the server also sends additional data for audio playback as needed.

[0143] Step 7:

[0144] The terminal converts the response received from the server into speech using speech synthesis technology and presents it to the user. This allows the user to receive an intuitive and emotionally empathetic response.

[0145] Step 8:

[0146] The user decides on their next action based on the response provided and provides feedback to the device as needed. This feedback is also acquired by the device and sent to the server for future service improvements.

[0147] Step 9:

[0148] The server analyzes the collected feedback information and adjusts the parameters of the AI ​​model and emotion engine to improve the accuracy and personalization of future responses.

[0149] (Example 2)

[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0151] There is a need to understand the emotional state of individuals who require support in their daily lives, such as the elderly and those who feel lonely, in real time and to respond appropriately. With conventional technology, it was difficult to effectively integrate voice and image data and provide personalized responses that were tailored to the user, which severely limited the improvement in satisfaction.

[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0153] In this invention, the server includes means for acquiring the user's voice and image using a terminal and transmitting the data to the server; means for analyzing the data using speech recognition and image recognition technologies on the server and creating a response using generation technology; and means for outputting the generated response to the user as voice using the terminal. This makes it possible to quickly and accurately grasp the user's emotions, provide individually appropriate support, and improve user satisfaction.

[0154] A "terminal" refers to a device that has the function of acquiring the user's voice and images in real time and transmitting them to a server.

[0155] A "server" refers to a central computing unit that analyzes audio and image data transmitted from terminals and creates responses using generation technology.

[0156] "Speech recognition technology" is a technology that converts speech data into text and is used to analyze the content of a user's speech.

[0157] "Image recognition technology" refers to technology that analyzes a user's facial expressions and gestures from image data to understand their situation.

[0158] "Generative technology" is a technology that generates appropriate responses to provide to the user based on analysis results, and is used to realize personalized interactions.

[0159] "Emotional state" refers to the internal psychological state inferred from the user's tone of voice, facial expressions, and gestures.

[0160] This invention relates to a system that utilizes voice and image data to enable interaction with users and incorporates an emotion engine that recognizes the user's emotions. This system consists of three main elements: a server, a terminal (robot), and a user.

[0161] The terminals are placed in the living spaces of elderly individuals and are equipped with hardware and software for acquiring voice and image data in real time. Specifically, they are equipped with high-performance microphones and cameras that capture the voice of the user speaking, as well as the user's facial expressions and surrounding environment. This acquired data is transmitted to a server via secure wireless communication.

[0162] The server employs various technologies to analyze the received data. Audio data is converted to text using speech recognition software. This process utilizes services such as the Google Speech-to-Text API to achieve high-precision speech recognition. Furthermore, an emotion engine analyzes the text data and uses natural language processing techniques to determine the user's emotional state. Image data is analyzed using image recognition technologies such as OpenCV and TENSORFLOW (registered trademark) to read detailed emotional states from the user's facial expressions and gestures.

[0163] Based on these analysis results, the server uses a generative AI model (e.g., GPT-4®) to generate an appropriate response for the user. An example of a specific prompt might be, "The user appears depressed. Please generate a kind comment to cheer them up." The generated response is then sent back to the terminal.

[0164] The device uses speech synthesis technology to convert responses from the server into speech and presents them to the user. To achieve this, it utilizes speech synthesis systems such as Amazon Polly and Google Text-to-Speech to communicate in a natural and friendly voice. The responses are designed to be empathetic to the user's emotions and provide a sense of security. For example, if a user says, "I'm lonely today," the device will gently ask, "Shall we talk?" In this way, the entire system supports the user's life and helps them achieve a happy and independent life.

[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0166] Step 1:

[0167] The device acquires the user's voice and images. It receives the user's spoken audio and video as input, and outputs digital audio and image data. Specifically, it converts sound waves into digital audio using the microphone and captures video with the camera. This data is then compressed and transmitted to the server wirelessly.

[0168] Step 2:

[0169] The server converts audio data transmitted from the terminal into text using speech recognition technology. It receives digital audio data as input and obtains text data as output. Specifically, it uses speech recognition software to perform feature extraction and phoneme matching to convert the audio into accurate text.

[0170] Step 3:

[0171] The server processes the text data obtained through speech recognition into an emotion analysis engine to determine the user's emotions. It receives text data as input and obtains an emotion state as output. Specifically, it uses a natural language processing engine to analyze keywords and context, and assigns emotion labels using an emotion dictionary and a machine learning model.

[0172] Step 4:

[0173] The server analyzes image data using image recognition technology to evaluate the user's facial expressions and gestures. It receives image data as input and outputs facial expression data and gesture data. Specifically, it uses a face detection algorithm to identify facial expressions and performs gesture recognition to extract features of hand movements and posture.

[0174] Step 5:

[0175] The server integrates emotional state and facial expression data obtained from audio and images, and uses a generative AI model to generate an appropriate response. It receives emotional state and facial expression data as input and obtains a text response as output. Specifically, it creates a prompt for the generative AI model and constructs a text response based on that prompt.

[0176] Step 6:

[0177] The server sends the generated text response to the terminal. It receives the generated text as input and sends a response to the terminal in an appropriate format as output. Specifically, it performs the process of packetizing digital data and using transmission control protocols to safely deliver the data to the terminal.

[0178] Step 7:

[0179] The terminal converts received text responses into speech using speech synthesis technology and presents them to the user. It receives text responses as input and obtains speech output as output. Specifically, it uses a speech synthesis system to play the text in a natural-sounding voice and speaks it aloud to the user through a speaker.

[0180] (Application Example 2)

[0181] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0182] In supporting the independent living of the elderly, it is crucial to provide appropriate support that responds to their emotions and needs, fostering a sense of security. However, conventional support systems have struggled to respond immediately to changes in users' emotions or emergencies, and have not been able to adequately guarantee the safety and security of elderly people living alone. There is a need to address these challenges and further improve the safety and security of the elderly.

[0183] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0184] In this invention, the server includes means for analyzing the user's state using voice and images and identifying their needs; means for generating and providing a response to the user based on the identified needs; means for collecting feedback from the user after the service has been provided and incorporating it into the next service; and means for monitoring the user's safety status in real time and notifying a remote information terminal. This enables immediate responses to anxieties and safety concerns felt by the elderly, allowing family members and related parties in remote locations to quickly recognize the situation and take appropriate measures.

[0185] "Voice" refers to the words and sounds spoken by the user to the robot or system, and is important data for understanding the user's intentions and emotions.

[0186] "Images" refer to visual information obtained through a camera, such as the user's facial expressions and surroundings, and are used as materials to understand the user's emotions and situation.

[0187] "Analyzing the user's state" refers to the process of analyzing voice and image data to determine the user's current emotions, health, or safety.

[0188] "Means of identifying needs" refers to methodologies for identifying the services and support that users require from analyzed data.

[0189] "Means of generating and providing responses" refers to the process by which a system determines an appropriate response or action based on the user's needs and emotions, and communicates it to the user.

[0190] "Collecting feedback" refers to the process of gathering user reactions and opinions on the services and responses provided, which is used to improve future services.

[0191] "Real-time monitoring of safety conditions" refers to the process of constantly monitoring users' living spaces and activities to detect abnormalities and emergencies.

[0192] "Notification to remote information terminals" refers to a communication method that sends important information about the user's status to devices of designated family members or related parties located remotely.

[0193] In an embodiment of the present invention, the system is configured as follows: The server, terminal, and user each fulfill their respective roles, providing safe and secure support through voice and image data.

[0194] The terminals are installed in the user's living space and function as robotic devices equipped with microphones and cameras. These devices continuously collect the user's voice and facial expressions and transmit them to a server.

[0195] The server converts this audio data into text data using a speech recognition API such as Google Speech-to-Text. Furthermore, it analyzes the image using image processing techniques such as OpenCV to detect the user's facial expressions and movements. In this process, it utilizes an emotion engine (e.g., Affectiva SDK) to make a multifaceted assessment of the user's emotional state.

[0196] Based on these analysis results, the server sends important notifications regarding the user's status to information terminals, such as those of family members in remote locations. It also generates reassuring voice responses and provides them to the user via the terminal. These voice responses are generated using Text-to-Speech (TTS) technology and are designed to be emotionally resonant to the user.

[0197] For example, if an elderly person feels lonely, the system identifies that emotion from the user's voice and facial expressions, and the server generates an adaptive message such as "Shall we talk?" while simultaneously notifying family members that "Mom seems lonely." Useful prompt phrases in this scenario could include "generating real-time notifications when the elderly person's emotions change and communicating them to family members in friendly language," or "creating responses that can immediately address elderly people who feel anxious or lonely."

[0198] Thus, this system makes it possible to quickly detect anxieties felt by users and take appropriate measures, thereby contributing to improved user safety and security.

[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0200] Step 1:

[0201] The device uses microphones and cameras placed around the user to acquire audio and image data in real time. During this process, the user's speech and facial expressions are obtained as input, converted into digital format, and sent to the server.

[0202] Step 2:

[0203] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology such as the Google Speech-to-Text API. The input is audio data, and the output is the corresponding text information. The audio data is also analyzed concurrently during this process.

[0204] Step 3:

[0205] The server receives image data sent from the terminal and analyzes the user's facial expressions and gestures using image processing techniques such as OpenCV. The input is image data, and the output is data indicating the identified emotional state. In this step, the data is processed using a facial recognition algorithm.

[0206] Step 4:

[0207] The server utilizes an emotion engine to comprehensively analyze the text data and emotion data obtained in steps 2 and 3 to determine the user's emotional state. The input is text data and facial expression data, and the output is an integrated emotion evaluation. At this stage, the tone of voice and facial expression patterns are analyzed.

[0208] Step 5:

[0209] The server generates an appropriate response based on the emotional state. A generative AI model is used to produce the most empathetic response for the user. The input is the previously obtained emotional assessment, and the output is the generated text-formatted response.

[0210] Step 6:

[0211] The server sends an instruction to the terminal to play back the generated response using text-to-speech (TTS) technology. The terminal receives this instruction and performs audio output. The input is a text-based response, and the output is in audio format.

[0212] Step 7:

[0213] The server notifies remote information terminals of the user's status and emotional assessment details, allowing family members and stakeholders to understand the situation in real time. Inputs include emotional assessments and generated notification messages, and output is the sending of notification messages.

[0214] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0215] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0216] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0217] [Second Embodiment]

[0218] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0219] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0220] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0221] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0222] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0223] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0224] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0225] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0226] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0227] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0228] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0229] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0230] This invention is an intelligent system for supporting the lives of the elderly, enabling communication with users through voice and images. It consists of three main components: a server, a terminal (robot), and a user, and is implemented in the following manner.

[0231] The server, located in the cloud, is responsible for advanced processing of speech recognition and image analysis. It converts the user's spoken audio data into text using speech recognition technology, and then analyzes the user's intent and emotional state based on that text data. It also recognizes the user's facial expressions and environmental elements from image data to understand the context. This allows the server to identify the user's needs and generate responses appropriate to those needs.

[0232] The terminal is placed in the living space of the elderly and functions as a user interface. This terminal is equipped with a camera and microphone, which acquire and transmit audio and image data to a server. Furthermore, it can present responses from the server to the user via speech synthesis and provide physical assistance as needed.

[0233] Specifically, when a user asks a question about cooking or everyday life, they speak to the robot via voice, and the robot receives the voice and transmits it to a server. The server converts the voice data into text and generates a contextually appropriate response. For example, if a user says, "I want to make soup," the robot will reply by voice, "What ingredients do you have?" Based on the user's response, the robot then receives more detailed support, such as recipe information, from the server and provides it to the user.

[0234] In this way, this system qualitatively supports the lives of the elderly and promotes independent living, thereby contributing to a reduction in the burden of caregiving. Furthermore, it provides users with an environment where they can live at their own pace with peace of mind through interaction with the robot.

[0235] The following describes the processing flow.

[0236] Step 1:

[0237] The device acquires voice input from the user via a microphone and simultaneously captures image data of the user's facial expressions and environment using a camera. This collected data is then transmitted to the server in real time.

[0238] Step 2:

[0239] The server converts the received audio data into text data using speech recognition technology. During this process, it analyzes the user's emotional state based on the tone and speed of their voice, using this information as foundational data to identify their needs.

[0240] Step 3:

[0241] The server processes the acquired image data using image analysis technology to analyze the user's facial expressions and gestures. Based on this information, it helps understand the user's context and identify their needs.

[0242] Step 4:

[0243] The server integrates information obtained from voice and images to identify the user's needs. Based on the identified needs, it internally determines the appropriate response and support method.

[0244] Step 5:

[0245] The server uses natural language generation technology to create a response that is easy for the user to understand, based on the determined response content. This generated response is then sent to the terminal.

[0246] Step 6:

[0247] The terminal presents the received response to the user as audio using speech synthesis technology. This allows the user to hear and understand the robot's response.

[0248] Step 7:

[0249] The user acts on the robot's responses and instructions, and asks additional questions or requests if necessary. This information is collected by the terminal as feedback that will be used again in the next cycle.

[0250] Step 8:

[0251] The device sends the feedback information back to the server, which then uses this feedback to adjust the system's response or as data to improve the AI ​​model. This process improves the accuracy of the service.

[0252] (Example 1)

[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0254] There is a need for intelligent support systems that enable elderly people to live independently while gaining a sense of security. Conventional systems rely solely on audio or image information, which has the problem of not being able to fully understand the user's intentions and needs. Furthermore, there has been insufficient mechanism for accurately incorporating user feedback into future services.

[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0256] In this invention, the server includes means for analyzing the user's state using voice and images and identifying the user's needs; means for generating and providing a response to the user based on the identified needs; and means for recognizing voice and converting it to text and generating a context-based response. This makes it possible to more accurately understand the user's intentions and emotions and to engage in appropriate dialogue. Furthermore, it enables the efficient collection of feedback and its incorporation into future services, thereby providing sustainable support.

[0257] "Speech recognition technology" is a technology that converts speech data into text data, and it extracts linguistic information by analyzing speech signals.

[0258] "Image recognition technology" is a technology that analyzes image data to recognize specific patterns or objects and makes decisions based on that information.

[0259] "Speech synthesis technology" is a technology that generates natural-sounding speech from text data, producing audio with a texture similar to human speech.

[0260] "Feedback collection" is the process of gathering opinions and reactions from users and using them to improve future services.

[0261] "Contextual understanding" is the process of grasping the background and intent of a conversation based on information obtained from audio and images, and deriving an appropriate response.

[0262] This invention provides an intelligent system to support the daily lives of elderly people. The system consists of three main components: a server, a terminal, and a user. Details are described below.

[0263] server:

[0264] The servers are located in a cloud environment and handle speech recognition and image analysis processing. Speech recognition uses technology to convert audio data into text, and this process utilizes common speech recognition software. For example, technologies such as Google Cloud Speech-to-Text and Amazon Transcribe are used. Image data is analyzed using image recognition technology to understand the user's facial expressions and environment. The servers then understand the user's intentions and needs and generate appropriate responses based on that information. The generated responses are delivered as natural-sounding speech using speech synthesis technology (for example, Amazon Polly or Google Cloud Text-to-Speech).

[0265] Terminal:

[0266] The terminal is placed within the user's living space and acquires audio and image data through its camera and microphone. This data is immediately transmitted to the server. A secure and efficient communication protocol is used to enable real-time data processing. The terminal also has the ability to play back responses from the server using speech synthesis technology and display relevant information on the display as needed.

[0267] Specific user examples:

[0268] When a user asks the robot on their device, "What should I make for dinner tonight?", the device sends the voice message to a server. The server transcribes the voice into text and analyzes the user's intent and emotions. For example, it can offer dinner ideas. The device then plays back the server's response and makes a suggestion to the user, such as, "How about pasta?" This specific dialogue supports the user in their daily activities.

[0269] Examples of generative AI models and prompt statements:

[0270] When using an AI model to generate responses in response to user requests, enter a prompt such as: "The user is asking for dinner ideas. Generate appropriate dish suggestions." Using specific prompts like this ensures that the generated responses align with the user's intent.

[0271] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0272] Step 1:

[0273] The device acquires the user's voice and images using the camera and microphone. The input consists of audio signals and image data, and the output is the digital format of that data. This operation collects fundamental data about the user's intentions and state.

[0274] Step 2:

[0275] The terminal transmits the acquired voice and image data to the server. The input is digitized voice and image data, and the output is data transfer to the server on the cloud.

[0276] Step 3:

[0277] The server receives the voice data and converts it into text data using voice recognition technology. The input is voice data, and the output is the corresponding text data. Through this process, the content of the user's speech can be interpreted as text.

[0278] Step 4:

[0279] The server analyzes the text data and analyzes the user's intention and emotion using the generated AI model. The input is text data, and the output is the analysis result. This analysis forms the basis for a response optimized for user needs.

[0280] Step 5:

[0281] The server analyzes the image data and recognizes the user's expression and the surrounding environment. The input is image data, and the output is the recognition result of the environment and state. Thereby, the context of the conversation is understood.

[0282] Step 6:

[0283] Based on the analysis results of voice and image, the server generates an appropriate response using the generated AI model. The input is the analyzed intention, emotion, and context, and the output is a specific response.

[0284] Step 7:

[0285] The server transmits the generated response to the terminal. The input is the generated response, and the output is data in a format suitable for speech synthesis.

[0286] Step 8:

[0287] The device uses speech synthesis technology to play the generated response as sound to the user. The input is audio data received from the server, and the output is human-readable speech. This allows the user to receive support.

[0288] (Application Example 1)

[0289] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0290] In industrial settings, workers need to prepare necessary tools and materials quickly and accurately in order to carry out their work efficiently. However, if workers have to go and find the necessary tools themselves, not only will work efficiency decrease, but the risk of errors due to incorrect tool selection and safety risks will also increase. This invention aims to solve these problems and improve work efficiency and safety.

[0291] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0292] In this invention, the server includes means for analyzing the user's state using voice and images and identifying their needs; means for generating and providing a response to the user based on the identified needs; means for collecting feedback from the user after the service has been provided and incorporating it into the next service; means for analyzing the voice instructions of workers in an industrial setting and selecting appropriate tools and materials; and means for providing the selected tools and materials to the workplace. This automates the selection and supply of items to enable workers to perform their tasks efficiently, thereby improving work efficiency and safety.

[0293] "Audio and images" refers to audio and image data obtained from users, and serves as a source of information for analyzing the user's state and needs.

[0294] "Means of identifying needs" refers to the process of analyzing acquired audio and image data to clarify the requests and requirements that users have.

[0295] "Means for generating and providing responses to users" refers to a system for creating appropriate information and support based on identified needs and communicating them to users.

[0296] "Means for collecting feedback and reflecting it in future services" refers to the process of collecting evaluations and attitudes from users after a service has been provided, and using that information to improve future services.

[0297] "Means for analyzing voice instructions given by workers in industrial settings" refers to technologies and processes for understanding and analyzing the content of voice instructions given by workers in industrial settings.

[0298] "Means for selecting appropriate tools and materials" refers to the process of determining and selecting the necessary tools and materials based on the worker's voice instructions.

[0299] "Means of providing selected tools and materials to the workshop" refers to a system for delivering selected tools and materials to the workshop at the appropriate time.

[0300] The system implementing this invention aims to efficiently support work in industrial settings. The system understands the needs of workers through the analysis of voice and image data and provides the necessary tools and materials.

[0301] The server is located in the cloud and has advanced speech recognition and image analysis capabilities. Speech recognition APIs (e.g., Google Speech-to-Text) are used to convert speech data into text and determine the worker's intent. Furthermore, OpenCV is used for image recognition technology to accurately identify necessary tools and materials. Based on the acquired information, the server generates appropriate responses and provides them to the user.

[0302] The terminal is installed at an industrial site and functions as an interface to assist with work. This terminal is equipped with a camera and a microphone, which acquire audio and image data and transmit it to the server. In addition, it conveys the response from the server to the operator by voice synthesis and provides physical assistance such as transporting tools and materials as needed.

[0303] The user, that is, the operator, gives work instructions to the system verbally. Based on this, the system analyzes the operator's request, prepares and provides the necessary items. Through this process, the operator can concentrate on the work, improving efficiency and ensuring safety at the same time.

[0304] As a specific example, when the operator says "Bring me a wrench that fits this bolt", the server can identify which wrench it is, send an instruction to the terminal, and deliver the wrench to the operator.

[0305] As an example of a prompt sentence, "List the tools most frequently used in the factory and optimize the word selection for voice recognition based on that" can be considered. This can further improve the accuracy of voice instructions and enable efficient work support.

[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0307] [[ID=二十一]]Step 1:

[0308] The terminal acquires the voice instruction from the user (operator) with the microphone. This voice data serves as the input. The terminal temporarily stores this voice data in the buffer and prepares to transmit it to the server in the next step.

[0309] Step 2:

[0310] The terminal sends the audio data acquired in step 1 to the server. The server uses a speech recognition API to convert the audio data into text. This converted text is the output and is used as the basis for determining the worker's intent.

[0311] Step 3:

[0312] The server uses a generative AI model to analyze the worker's intent from the text data in step 2. This process understands the context of words and phrases and determines the necessary tools and materials. As a result of this analysis, information on the selected tools and materials is output.

[0313] Step 4:

[0314] The server uses image recognition technology to generate specific instructions for selecting particular tools and materials based on the information from step 3. Here, location information and identification tags for tools and materials are extracted from the image data, and the information of the selected items is organized.

[0315] Step 5:

[0316] The terminal receives the specific tool and material information generated in step 4 and provides feedback to the user via speech synthesis. It provides the worker with responses such as "I'll bring the XX wrench," building trust in the system.

[0317] Step 6:

[0318] The terminal actually transports the identified tools and materials based on the worker's instructions. In this process, the selected tools are moved safely and accurately by the mounted arm or transport device. Once everything is complete, the worker is notified and the terminal awaits further instructions.

[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0320] This invention is a support system that enables interaction with users using voice and image data, and incorporates an emotion engine that recognizes the user's emotions. The system consists of three main elements: a server, a terminal (robot), and a user, and is implemented in the following manner.

[0321] The device is placed in the living space of elderly people and acquires voice and image data in real time. When a user speaks to the robot, the voice input is captured by the device's microphone. The device also captures the user's facial expressions and surrounding environment through its camera. This data is immediately transmitted to a server.

[0322] The server converts received audio data into text using speech recognition technology, while simultaneously analyzing the user's emotional state from their speech using an emotion engine. Furthermore, image data is processed using image recognition technology to analyze the user's facial expressions and gestures. In this way, the server integrates audio and image data to understand the context and identify the user's needs and emotions.

[0323] The emotion engine uses a multifaceted approach, including voice tone analysis and facial expression analysis, to determine the user's emotions and decide on a response based on those emotions. Based on these analysis results, the server generates an appropriate response and sends it to the terminal.

[0324] The terminal presents responses from the server to the user in voice using speech synthesis technology. The responses are tailored to the user's emotions and designed to provide a sense of reassurance. For example, if the user says, "I'm lonely today," the robot will gently respond, "Shall we talk?" to cheer the user up.

[0325] In this way, this system understands users' needs and emotions and provides personalized support, thereby creating an environment where seniors can live independent and happy lives. Furthermore, by utilizing emotional information, the content of subsequent service provision can be made more individualized, leading to continuous improvement in satisfaction.

[0326] The following describes the processing flow.

[0327] Step 1:

[0328] The device acquires voice input from the user via the microphone and captures image data of the user's facial expressions and environment using the camera. The acquired data is immediately sent to the server.

[0329] Step 2:

[0330] The server receives the audio data and converts it into text using speech recognition technology. During this process, it also analyzes the tone and speed of the speech to infer the user's emotional state.

[0331] Step 3:

[0332] The server processes image data using image analysis technology to analyze the user's facial expressions and gestures. This deepens the understanding of the user's emotions and situation.

[0333] Step 4:

[0334] The server inputs information obtained from voice and images into the emotion engine to identify the user's emotions. Based on this, it integrates the user's needs and emotions to determine a response.

[0335] Step 5:

[0336] The server uses natural language generation technology to create responses that match the user's emotions. The generated responses are designed to be polite, considerate, and thoughtful.

[0337] Step 6:

[0338] The server sends the generated response data to the terminal. The response is sent as text data, but the server also sends additional data for audio playback as needed.

[0339] Step 7:

[0340] The terminal converts the response received from the server into speech using speech synthesis technology and presents it to the user. This allows the user to receive an intuitive and emotionally empathetic response.

[0341] Step 8:

[0342] The user decides on their next action based on the response provided and provides feedback to the device as needed. This feedback is also acquired by the device and sent to the server for future service improvements.

[0343] Step 9:

[0344] The server analyzes the collected feedback information and adjusts the parameters of the AI ​​model and emotion engine to improve the accuracy and personalization of future responses.

[0345] (Example 2)

[0346] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0347] There is a need to understand the emotional state of individuals who require support in their daily lives, such as the elderly and those who feel lonely, in real time and to respond appropriately. With conventional technology, it was difficult to effectively integrate voice and image data and provide personalized responses that were tailored to the user, which severely limited the improvement in satisfaction.

[0348] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0349] In this invention, the server includes means for acquiring the user's voice and image using a terminal and transmitting the data to the server; means for analyzing the data using speech recognition and image recognition technologies on the server and creating a response using generation technology; and means for outputting the generated response to the user as voice using the terminal. This makes it possible to quickly and accurately grasp the user's emotions, provide individually appropriate support, and improve user satisfaction.

[0350] A "terminal" refers to a device that has the function of acquiring the user's voice and images in real time and transmitting them to a server.

[0351] A "server" refers to a central computing unit that analyzes audio and image data transmitted from terminals and creates responses using generation technology.

[0352] "Speech recognition technology" is a technology that converts speech data into text and is used to analyze the content of a user's speech.

[0353] "Image recognition technology" refers to technology that analyzes a user's facial expressions and gestures from image data to understand their situation.

[0354] "Generative technology" is a technology that generates appropriate responses to provide to the user based on analysis results, and is used to realize personalized interactions.

[0355] "Emotional state" refers to the internal psychological state inferred from the user's tone of voice, facial expressions, and gestures.

[0356] This invention relates to a system that utilizes voice and image data to enable interaction with users and incorporates an emotion engine that recognizes the user's emotions. This system consists of three main elements: a server, a terminal (robot), and a user.

[0357] The terminals are placed in the living spaces of elderly individuals and are equipped with hardware and software for acquiring voice and image data in real time. Specifically, they are equipped with high-performance microphones and cameras that capture the voice of the user speaking, as well as the user's facial expressions and surrounding environment. This acquired data is transmitted to a server via secure wireless communication.

[0358] The server employs various technologies to analyze the received data. Audio data is converted to text using speech recognition software. This process utilizes services such as the Google Speech-to-Text API to achieve high-precision speech recognition. Furthermore, an emotion engine analyzes the text data and uses natural language processing techniques to determine the user's emotional state. Image data is analyzed using image recognition technologies such as OpenCV and TensorFlow to read detailed emotional states from the user's facial expressions and gestures.

[0359] Based on these analysis results, the server uses a generative AI model (e.g., GPT-4) to generate an appropriate response for the user. An example of a specific prompt might be, "The user appears depressed. Please generate a kind comment to cheer them up." The generated response is then sent back to the terminal.

[0360] The device uses speech synthesis technology to convert responses from the server into speech and presents them to the user. To achieve this, it utilizes speech synthesis systems such as Amazon Polly and Google Text-to-Speech to communicate in a natural and friendly voice. The responses are designed to be empathetic to the user's emotions and provide a sense of security. For example, if a user says, "I'm lonely today," the device will gently ask, "Shall we talk?" In this way, the entire system supports the user's life and helps them achieve a happy and independent life.

[0361] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0362] Step 1:

[0363] The device acquires the user's voice and images. It receives the user's spoken audio and video as input, and outputs digital audio and image data. Specifically, it converts sound waves into digital audio using the microphone and captures video with the camera. This data is then compressed and transmitted to the server wirelessly.

[0364] Step 2:

[0365] The server converts audio data transmitted from the terminal into text using speech recognition technology. It receives digital audio data as input and obtains text data as output. Specifically, it uses speech recognition software to perform feature extraction and phoneme matching to convert the audio into accurate text.

[0366] Step 3:

[0367] The server processes the text data obtained through speech recognition into an emotion analysis engine to determine the user's emotions. It receives text data as input and obtains an emotion state as output. Specifically, it uses a natural language processing engine to analyze keywords and context, and assigns emotion labels using an emotion dictionary and a machine learning model.

[0368] Step 4:

[0369] The server analyzes image data using image recognition technology to evaluate the user's facial expressions and gestures. It receives image data as input and outputs facial expression data and gesture data. Specifically, it uses a face detection algorithm to identify facial expressions and performs gesture recognition to extract features of hand movements and posture.

[0370] Step 5:

[0371] The server integrates emotional state and facial expression data obtained from audio and images, and uses a generative AI model to generate an appropriate response. It receives emotional state and facial expression data as input and obtains a text response as output. Specifically, it creates a prompt for the generative AI model and constructs a text response based on that prompt.

[0372] Step 6:

[0373] The server sends the generated text response to the terminal. It receives the generated text as input and sends a response to the terminal in an appropriate format as output. Specifically, it performs the process of packetizing digital data and using transmission control protocols to safely deliver the data to the terminal.

[0374] Step 7:

[0375] The terminal converts received text responses into speech using speech synthesis technology and presents them to the user. It receives text responses as input and obtains speech output as output. Specifically, it uses a speech synthesis system to play the text in a natural-sounding voice and speaks it aloud to the user through a speaker.

[0376] (Application Example 2)

[0377] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0378] In supporting the independent living of the elderly, it is crucial to provide appropriate support that responds to their emotions and needs, fostering a sense of security. However, conventional support systems have struggled to respond immediately to changes in users' emotions or emergencies, and have not been able to adequately guarantee the safety and security of elderly people living alone. There is a need to address these challenges and further improve the safety and security of the elderly.

[0379] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0380] In this invention, the server includes means for analyzing the user's state using voice and images and identifying their needs; means for generating and providing a response to the user based on the identified needs; means for collecting feedback from the user after the service has been provided and incorporating it into the next service; and means for monitoring the user's safety status in real time and notifying a remote information terminal. This enables immediate responses to anxieties and safety concerns felt by the elderly, allowing family members and related parties in remote locations to quickly recognize the situation and take appropriate measures.

[0381] "Voice" refers to the words and sounds spoken by the user to the robot or system, and is important data for understanding the user's intentions and emotions.

[0382] "Images" refer to visual information obtained through a camera, such as the user's facial expressions and surroundings, and are used as materials to understand the user's emotions and situation.

[0383] "Analyzing the user's state" refers to the process of analyzing voice and image data to determine the user's current emotions, health, or safety.

[0384] "Means of identifying needs" refers to methodologies for identifying the services and support that users require from analyzed data.

[0385] "Means of generating and providing responses" refers to the process by which a system determines an appropriate response or action based on the user's needs and emotions, and communicates it to the user.

[0386] "Collecting feedback" refers to the process of gathering user reactions and opinions on the services and responses provided, which is used to improve future services.

[0387] "Real-time monitoring of safety conditions" refers to the process of constantly monitoring users' living spaces and activities to detect abnormalities and emergencies.

[0388] "Notification to remote information terminals" refers to a communication method that sends important information about the user's status to devices of designated family members or related parties located remotely.

[0389] In an embodiment of the present invention, the system is configured as follows: The server, terminal, and user each fulfill their respective roles, providing safe and secure support through voice and image data.

[0390] The terminals are installed in the user's living space and function as robotic devices equipped with microphones and cameras. These devices continuously collect the user's voice and facial expressions and transmit them to a server.

[0391] The server converts this audio data into text data using a speech recognition API such as Google Speech-to-Text. Furthermore, it analyzes the image using image processing techniques such as OpenCV to detect the user's facial expressions and movements. In this process, it utilizes an emotion engine (e.g., Affectiva SDK) to make a multifaceted assessment of the user's emotional state.

[0392] Based on these analysis results, the server sends important notifications regarding the user's status to information terminals, such as those of family members in remote locations. It also generates reassuring voice responses and provides them to the user via the terminal. These voice responses are generated using Text-to-Speech (TTS) technology and are designed to be emotionally resonant to the user.

[0393] For example, if an elderly person feels lonely, the system identifies that emotion from the user's voice and facial expressions, and the server generates an adaptive message such as "Shall we talk?" while simultaneously notifying family members that "Mom seems lonely." Useful prompt phrases in this scenario could include "generating real-time notifications when the elderly person's emotions change and communicating them to family members in friendly language," or "creating responses that can immediately address elderly people who feel anxious or lonely."

[0394] Thus, this system makes it possible to quickly detect anxieties felt by users and take appropriate measures, thereby contributing to improved user safety and security.

[0395] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0396] Step 1:

[0397] The device uses microphones and cameras placed around the user to acquire audio and image data in real time. During this process, the user's speech and facial expressions are obtained as input, converted into digital format, and sent to the server.

[0398] Step 2:

[0399] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology such as the Google Speech-to-Text API. The input is audio data, and the output is the corresponding text information. The audio data is also analyzed concurrently during this process.

[0400] Step 3:

[0401] The server receives image data sent from the terminal and analyzes the user's facial expressions and gestures using image processing techniques such as OpenCV. The input is image data, and the output is data indicating the identified emotional state. In this step, the data is processed using a facial recognition algorithm.

[0402] Step 4:

[0403] The server utilizes an emotion engine to comprehensively analyze the text data and emotion data obtained in steps 2 and 3 to determine the user's emotional state. The input is text data and facial expression data, and the output is an integrated emotion evaluation. At this stage, the tone of voice and facial expression patterns are analyzed.

[0404] Step 5:

[0405] The server generates an appropriate response based on the emotional state. A generative AI model is used to produce the most empathetic response for the user. The input is the previously obtained emotional assessment, and the output is the generated text-formatted response.

[0406] Step 6:

[0407] The server sends an instruction to the terminal to play back the generated response using text-to-speech (TTS) technology. The terminal receives this instruction and performs audio output. The input is a text-based response, and the output is in audio format.

[0408] Step 7:

[0409] The server notifies remote information terminals of the user's status and emotional assessment details, allowing family members and stakeholders to understand the situation in real time. Inputs include emotional assessments and generated notification messages, and output is the sending of notification messages.

[0410] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0411] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0412] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0413] [Third Embodiment]

[0414] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0415] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0416] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0417] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0418] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0420] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0421] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0422] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0423] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0424] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0425] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0426] This invention is an intelligent system for supporting the lives of the elderly, enabling communication with users through voice and images. It consists of three main components: a server, a terminal (robot), and a user, and is implemented in the following manner.

[0427] The server, located in the cloud, is responsible for advanced processing of speech recognition and image analysis. It converts the user's spoken audio data into text using speech recognition technology, and then analyzes the user's intent and emotional state based on that text data. It also recognizes the user's facial expressions and environmental elements from image data to understand the context. This allows the server to identify the user's needs and generate responses appropriate to those needs.

[0428] The terminal is placed in the living space of the elderly and functions as a user interface. This terminal is equipped with a camera and microphone, which acquire and transmit audio and image data to a server. Furthermore, it can present responses from the server to the user via speech synthesis and provide physical assistance as needed.

[0429] Specifically, when a user asks a question about cooking or everyday life, they speak to the robot via voice, and the robot receives the voice and transmits it to a server. The server converts the voice data into text and generates a contextually appropriate response. For example, if a user says, "I want to make soup," the robot will reply by voice, "What ingredients do you have?" Based on the user's response, the robot then receives more detailed support, such as recipe information, from the server and provides it to the user.

[0430] In this way, this system qualitatively supports the lives of the elderly and promotes independent living, thereby contributing to a reduction in the burden of caregiving. Furthermore, it provides users with an environment where they can live at their own pace with peace of mind through interaction with the robot.

[0431] The following describes the processing flow.

[0432] Step 1:

[0433] The device acquires voice input from the user via a microphone and simultaneously captures image data of the user's facial expressions and environment using a camera. This collected data is then transmitted to the server in real time.

[0434] Step 2:

[0435] The server converts the received audio data into text data using speech recognition technology. During this process, it analyzes the user's emotional state based on the tone and speed of their voice, using this information as foundational data to identify their needs.

[0436] Step 3:

[0437] The server processes the acquired image data using image analysis technology to analyze the user's facial expressions and gestures. Based on this information, it helps understand the user's context and identify their needs.

[0438] Step 4:

[0439] The server integrates information obtained from voice and images to identify the user's needs. Based on the identified needs, it internally determines the appropriate response and support method.

[0440] Step 5:

[0441] The server uses natural language generation technology to create a response that is easy for the user to understand, based on the determined response content. This generated response is then sent to the terminal.

[0442] Step 6:

[0443] The terminal presents the received response to the user as audio using speech synthesis technology. This allows the user to hear and understand the robot's response.

[0444] Step 7:

[0445] The user acts on the robot's responses and instructions, and asks additional questions or requests if necessary. This information is collected by the terminal as feedback that will be used again in the next cycle.

[0446] Step 8:

[0447] The device sends the feedback information back to the server, which then uses this feedback to adjust the system's response or as data to improve the AI ​​model. This process improves the accuracy of the service.

[0448] (Example 1)

[0449] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0450] There is a need for intelligent support systems that enable elderly people to live independently while gaining a sense of security. Conventional systems rely solely on audio or image information, which has the problem of not being able to fully understand the user's intentions and needs. Furthermore, there has been insufficient mechanism for accurately incorporating user feedback into future services.

[0451] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0452] In this invention, the server includes means for analyzing the user's state using voice and images and identifying the user's needs; means for generating and providing a response to the user based on the identified needs; and means for recognizing voice and converting it to text and generating a context-based response. This makes it possible to more accurately understand the user's intentions and emotions and to engage in appropriate dialogue. Furthermore, it enables the efficient collection of feedback and its incorporation into future services, thereby providing sustainable support.

[0453] "Speech recognition technology" is a technology that converts speech data into text data, and it extracts linguistic information by analyzing speech signals.

[0454] "Image recognition technology" is a technology that analyzes image data to recognize specific patterns or objects and makes decisions based on that information.

[0455] "Speech synthesis technology" is a technology that generates natural-sounding speech from text data, producing audio with a texture similar to human speech.

[0456] "Feedback collection" is the process of gathering opinions and reactions from users and using them to improve future services.

[0457] "Contextual understanding" is the process of grasping the background and intent of a conversation based on information obtained from audio and images, and deriving an appropriate response.

[0458] This invention provides an intelligent system to support the daily lives of elderly people. The system consists of three main components: a server, a terminal, and a user. Details are described below.

[0459] server:

[0460] The servers are located in a cloud environment and handle speech recognition and image analysis processing. Speech recognition uses technology to convert audio data into text, and this process utilizes common speech recognition software. For example, technologies such as Google Cloud Speech-to-Text and Amazon Transcribe are used. Image data is analyzed using image recognition technology to understand the user's facial expressions and environment. The servers then understand the user's intentions and needs and generate appropriate responses based on that information. The generated responses are delivered as natural-sounding speech using speech synthesis technology (for example, Amazon Polly or Google Cloud Text-to-Speech).

[0461] Terminal:

[0462] The terminal is placed within the user's living space and acquires audio and image data through its camera and microphone. This data is immediately transmitted to the server. A secure and efficient communication protocol is used to enable real-time data processing. The terminal also has the ability to play back responses from the server using speech synthesis technology and display relevant information on the display as needed.

[0463] Specific user examples:

[0464] When a user asks the robot on their device, "What should I make for dinner tonight?", the device sends the voice message to a server. The server transcribes the voice into text and analyzes the user's intent and emotions. For example, it can offer dinner ideas. The device then plays back the server's response and makes a suggestion to the user, such as, "How about pasta?" This specific dialogue supports the user in their daily activities.

[0465] Examples of generative AI models and prompt statements:

[0466] When using an AI model to generate responses in response to user requests, enter a prompt such as: "The user is asking for dinner ideas. Generate appropriate dish suggestions." Using specific prompts like this ensures that the generated responses align with the user's intent.

[0467] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0468] Step 1:

[0469] The device acquires the user's voice and images using the camera and microphone. The input consists of audio signals and image data, and the output is the digital format of that data. This operation collects fundamental data about the user's intentions and state.

[0470] Step 2:

[0471] The terminal transmits the acquired audio and image data to the server. The input is digitized audio and image data, and the output is data transfer to a server in the cloud.

[0472] Step 3:

[0473] The server receives audio data and converts it into text data using speech recognition technology. The input is audio data, and the output is the corresponding text data. This process makes the user's speech interpretable as text.

[0474] Step 4:

[0475] The server analyzes text data and uses a generative AI model to analyze the user's intent and emotions. The input is text data, and the output is the analysis result. This analysis forms the basis for responses optimized to user needs.

[0476] Step 5:

[0477] The server analyzes image data to recognize the user's facial expressions and surrounding environment. The input is image data, and the output is the result of the recognition of the environment and state. This allows for an understanding of the context of the conversation.

[0478] Step 6:

[0479] The server generates appropriate responses using a generative AI model based on the analysis results of the audio and images. The input is the analyzed intent, emotion, and context, and the output is a specific response.

[0480] Step 7:

[0481] The server sends the generated response to the terminal. The input is the generated response, and the output is data in a format suitable for speech synthesis.

[0482] Step 8:

[0483] The device uses speech synthesis technology to play the generated response as sound to the user. The input is audio data received from the server, and the output is human-readable speech. This allows the user to receive support.

[0484] (Application Example 1)

[0485] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0486] In industrial settings, workers need to prepare necessary tools and materials quickly and accurately in order to carry out their work efficiently. However, if workers have to go and find the necessary tools themselves, not only will work efficiency decrease, but the risk of errors due to incorrect tool selection and safety risks will also increase. This invention aims to solve these problems and improve work efficiency and safety.

[0487] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0488] In this invention, the server includes means for analyzing the user's state using voice and images and identifying their needs; means for generating and providing a response to the user based on the identified needs; means for collecting feedback from the user after the service has been provided and incorporating it into the next service; means for analyzing the voice instructions of workers in an industrial setting and selecting appropriate tools and materials; and means for providing the selected tools and materials to the workplace. This automates the selection and supply of items to enable workers to perform their tasks efficiently, thereby improving work efficiency and safety.

[0489] "Audio and images" refers to audio and image data obtained from users, and serves as a source of information for analyzing the user's state and needs.

[0490] "Means of identifying needs" refers to the process of analyzing acquired audio and image data to clarify the requests and requirements that users have.

[0491] "Means for generating and providing responses to users" refers to a system for creating appropriate information and support based on identified needs and communicating them to users.

[0492] "Means for collecting feedback and reflecting it in future services" refers to the process of collecting evaluations and attitudes from users after a service has been provided, and using that information to improve future services.

[0493] "Means for analyzing voice instructions given by workers in industrial settings" refers to technologies and processes for understanding and analyzing the content of voice instructions given by workers in industrial settings.

[0494] "Means for selecting appropriate tools and materials" refers to the process of determining and selecting the necessary tools and materials based on the worker's voice instructions.

[0495] "Means of providing selected tools and materials to the workshop" refers to a system for delivering selected tools and materials to the workshop at the appropriate time.

[0496] The system implementing this invention aims to efficiently support work in industrial settings. The system understands the needs of workers through the analysis of voice and image data and provides the necessary tools and materials.

[0497] The server is located in the cloud and has advanced speech recognition and image analysis capabilities. Speech recognition APIs (e.g., Google Speech-to-Text) are used to convert speech data into text and determine the worker's intent. Furthermore, OpenCV is used for image recognition technology to accurately identify necessary tools and materials. Based on the acquired information, the server generates appropriate responses and provides them to the user.

[0498] The terminal is installed in the industrial site and functions as an interface to assist with work. This terminal is equipped with a camera and microphone, which acquire and transmit audio and image data to a server. It also uses speech synthesis to communicate responses from the server to the worker and provides physical assistance, such as transporting tools and materials, as needed.

[0499] The user, or worker, gives voice instructions to the system. Based on these instructions, the system analyzes the worker's requests and prepares and provides the necessary items. This process allows the worker to concentrate on their work, improving efficiency while also ensuring safety.

[0500] For example, if a worker says, "Bring me the wrench that fits this bolt," the server can identify the appropriate wrench, send an instruction to the terminal, and have that wrench delivered to the worker.

[0501] An example of a prompt message could be, "List the most frequently used tools in the factory and use that to optimize word selection for speech recognition." This would further improve the accuracy of voice commands and enable more efficient work support.

[0502] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0503] Step 1:

[0504] The terminal receives voice instructions from the user (worker) via a microphone. This voice data becomes the input. The terminal temporarily stores this voice data in a buffer and prepares to send it to the server in the next step.

[0505] Step 2:

[0506] The terminal sends the audio data acquired in step 1 to the server. The server uses a speech recognition API to convert the audio data into text. This converted text is the output and is used as the basis for determining the worker's intent.

[0507] Step 3:

[0508] The server uses a generative AI model to analyze the worker's intent from the text data in step 2. This process understands the context of words and phrases and determines the necessary tools and materials. As a result of this analysis, information on the selected tools and materials is output.

[0509] Step 4:

[0510] The server uses image recognition technology to generate specific instructions for selecting particular tools and materials based on the information from step 3. Here, location information and identification tags for tools and materials are extracted from the image data, and the information of the selected items is organized.

[0511] Step 5:

[0512] The terminal receives the specific tool and material information generated in step 4 and provides feedback to the user via speech synthesis. It provides the worker with responses such as "I'll bring the XX wrench," building trust in the system.

[0513] Step 6:

[0514] The terminal actually transports the identified tools and materials based on the worker's instructions. In this process, the selected tools are moved safely and accurately by the mounted arm or transport device. Once everything is complete, the worker is notified and the terminal awaits further instructions.

[0515] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0516] This invention is a support system that enables interaction with users using voice and image data, and incorporates an emotion engine that recognizes the user's emotions. The system consists of three main elements: a server, a terminal (robot), and a user, and is implemented in the following manner.

[0517] The device is placed in the living space of elderly people and acquires voice and image data in real time. When a user speaks to the robot, the voice input is captured by the device's microphone. The device also captures the user's facial expressions and surrounding environment through its camera. This data is immediately transmitted to a server.

[0518] The server converts received audio data into text using speech recognition technology, while simultaneously analyzing the user's emotional state from their speech using an emotion engine. Furthermore, image data is processed using image recognition technology to analyze the user's facial expressions and gestures. In this way, the server integrates audio and image data to understand the context and identify the user's needs and emotions.

[0519] The emotion engine uses a multifaceted approach, including voice tone analysis and facial expression analysis, to determine the user's emotions and decide on a response based on those emotions. Based on these analysis results, the server generates an appropriate response and sends it to the terminal.

[0520] The terminal presents responses from the server to the user in voice using speech synthesis technology. The responses are tailored to the user's emotions and designed to provide a sense of reassurance. For example, if the user says, "I'm lonely today," the robot will gently respond, "Shall we talk?" to cheer the user up.

[0521] In this way, this system understands users' needs and emotions and provides personalized support, thereby creating an environment where seniors can live independent and happy lives. Furthermore, by utilizing emotional information, the content of subsequent service provision can be made more individualized, leading to continuous improvement in satisfaction.

[0522] The following describes the processing flow.

[0523] Step 1:

[0524] The device acquires voice input from the user via the microphone and captures image data of the user's facial expressions and environment using the camera. The acquired data is immediately sent to the server.

[0525] Step 2:

[0526] The server receives the audio data and converts it into text using speech recognition technology. During this process, it also analyzes the tone and speed of the speech to infer the user's emotional state.

[0527] Step 3:

[0528] The server processes image data using image analysis technology to analyze the user's facial expressions and gestures. This deepens the understanding of the user's emotions and situation.

[0529] Step 4:

[0530] The server inputs information obtained from voice and images into the emotion engine to identify the user's emotions. Based on this, it integrates the user's needs and emotions to determine a response.

[0531] Step 5:

[0532] The server uses natural language generation technology to create responses that match the user's emotions. The generated responses are designed to be polite, considerate, and thoughtful.

[0533] Step 6:

[0534] The server sends the generated response data to the terminal. The response is sent as text data, but the server also sends additional data for audio playback as needed.

[0535] Step 7:

[0536] The terminal converts the response received from the server into speech using speech synthesis technology and presents it to the user. This allows the user to receive an intuitive and emotionally empathetic response.

[0537] Step 8:

[0538] The user decides on their next action based on the response provided and provides feedback to the device as needed. This feedback is also acquired by the device and sent to the server for future service improvements.

[0539] Step 9:

[0540] The server analyzes the collected feedback information and adjusts the parameters of the AI ​​model and emotion engine to improve the accuracy and personalization of future responses.

[0541] (Example 2)

[0542] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0543] There is a need to understand the emotional state of individuals who require support in their daily lives, such as the elderly and those who feel lonely, in real time and to respond appropriately. With conventional technology, it was difficult to effectively integrate voice and image data and provide personalized responses that were tailored to the user, which severely limited the improvement in satisfaction.

[0544] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0545] In this invention, the server includes means for acquiring the user's voice and image using a terminal and transmitting the data to the server; means for analyzing the data using speech recognition and image recognition technologies on the server and creating a response using generation technology; and means for outputting the generated response to the user as voice using the terminal. This makes it possible to quickly and accurately grasp the user's emotions, provide individually appropriate support, and improve user satisfaction.

[0546] A "terminal" refers to a device that has the function of acquiring the user's voice and images in real time and transmitting them to a server.

[0547] A "server" refers to a central computing unit that analyzes audio and image data transmitted from terminals and creates responses using generation technology.

[0548] "Speech recognition technology" is a technology that converts speech data into text and is used to analyze the content of a user's speech.

[0549] "Image recognition technology" refers to technology that analyzes a user's facial expressions and gestures from image data to understand their situation.

[0550] "Generative technology" is a technology that generates appropriate responses to provide to the user based on analysis results, and is used to realize personalized interactions.

[0551] "Emotional state" refers to the internal psychological state inferred from the user's tone of voice, facial expressions, and gestures.

[0552] This invention relates to a system that utilizes voice and image data to enable interaction with users and incorporates an emotion engine that recognizes the user's emotions. This system consists of three main elements: a server, a terminal (robot), and a user.

[0553] The terminals are placed in the living spaces of elderly individuals and are equipped with hardware and software for acquiring voice and image data in real time. Specifically, they are equipped with high-performance microphones and cameras that capture the voice of the user speaking, as well as the user's facial expressions and surrounding environment. This acquired data is transmitted to a server via secure wireless communication.

[0554] The server employs various technologies to analyze the received data. Audio data is converted to text using speech recognition software. This process utilizes services such as the Google Speech-to-Text API to achieve high-precision speech recognition. Furthermore, an emotion engine analyzes the text data and uses natural language processing techniques to determine the user's emotional state. Image data is analyzed using image recognition technologies such as OpenCV and TensorFlow to read detailed emotional states from the user's facial expressions and gestures.

[0555] Based on these analysis results, the server uses a generative AI model (e.g., GPT-4) to generate an appropriate response for the user. An example of a specific prompt might be, "The user appears depressed. Please generate a kind comment to cheer them up." The generated response is then sent back to the terminal.

[0556] The device uses speech synthesis technology to convert responses from the server into speech and presents them to the user. To achieve this, it utilizes speech synthesis systems such as Amazon Polly and Google Text-to-Speech to communicate in a natural and friendly voice. The responses are designed to be empathetic to the user's emotions and provide a sense of security. For example, if a user says, "I'm lonely today," the device will gently ask, "Shall we talk?" In this way, the entire system supports the user's life and helps them achieve a happy and independent life.

[0557] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0558] Step 1:

[0559] The device acquires the user's voice and images. It receives the user's spoken audio and video as input, and outputs digital audio and image data. Specifically, it converts sound waves into digital audio using the microphone and captures video with the camera. This data is then compressed and transmitted to the server wirelessly.

[0560] Step 2:

[0561] The server converts audio data transmitted from the terminal into text using speech recognition technology. It receives digital audio data as input and obtains text data as output. Specifically, it uses speech recognition software to perform feature extraction and phoneme matching to convert the audio into accurate text.

[0562] Step 3:

[0563] The server processes the text data obtained through speech recognition into an emotion analysis engine to determine the user's emotions. It receives text data as input and obtains an emotion state as output. Specifically, it uses a natural language processing engine to analyze keywords and context, and assigns emotion labels using an emotion dictionary and a machine learning model.

[0564] Step 4:

[0565] The server analyzes image data using image recognition technology to evaluate the user's facial expressions and gestures. It receives image data as input and outputs facial expression data and gesture data. Specifically, it uses a face detection algorithm to identify facial expressions and performs gesture recognition to extract features of hand movements and posture.

[0566] Step 5:

[0567] The server integrates emotional state and facial expression data obtained from audio and images, and uses a generative AI model to generate an appropriate response. It receives emotional state and facial expression data as input and obtains a text response as output. Specifically, it creates a prompt for the generative AI model and constructs a text response based on that prompt.

[0568] Step 6:

[0569] The server sends the generated text response to the terminal. It receives the generated text as input and sends a response to the terminal in an appropriate format as output. Specifically, it performs the process of packetizing digital data and using transmission control protocols to safely deliver the data to the terminal.

[0570] Step 7:

[0571] The terminal converts received text responses into speech using speech synthesis technology and presents them to the user. It receives text responses as input and obtains speech output as output. Specifically, it uses a speech synthesis system to play the text in a natural-sounding voice and speaks it aloud to the user through a speaker.

[0572] (Application Example 2)

[0573] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0574] In supporting the independent living of the elderly, it is crucial to provide appropriate support that responds to their emotions and needs, fostering a sense of security. However, conventional support systems have struggled to respond immediately to changes in users' emotions or emergencies, and have not been able to adequately guarantee the safety and security of elderly people living alone. There is a need to address these challenges and further improve the safety and security of the elderly.

[0575] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0576] In this invention, the server includes means for analyzing the user's state using voice and images and identifying their needs; means for generating and providing a response to the user based on the identified needs; means for collecting feedback from the user after the service has been provided and incorporating it into the next service; and means for monitoring the user's safety status in real time and notifying a remote information terminal. This enables immediate responses to anxieties and safety concerns felt by the elderly, allowing family members and related parties in remote locations to quickly recognize the situation and take appropriate measures.

[0577] "Voice" refers to the words and sounds spoken by the user to the robot or system, and is important data for understanding the user's intentions and emotions.

[0578] "Images" refer to visual information obtained through a camera, such as the user's facial expressions and surroundings, and are used as materials to understand the user's emotions and situation.

[0579] "Analyzing the user's state" refers to the process of analyzing voice and image data to determine the user's current emotions, health, or safety.

[0580] "Means of identifying needs" refers to methodologies for identifying the services and support that users require from analyzed data.

[0581] "Means of generating and providing responses" refers to the process by which a system determines an appropriate response or action based on the user's needs and emotions, and communicates it to the user.

[0582] "Collecting feedback" refers to the process of gathering user reactions and opinions on the services and responses provided, which is used to improve future services.

[0583] "Real-time monitoring of safety conditions" refers to the process of constantly monitoring users' living spaces and activities to detect abnormalities and emergencies.

[0584] "Notification to remote information terminals" refers to a communication method that sends important information about the user's status to devices of designated family members or related parties located remotely.

[0585] In an embodiment of the present invention, the system is configured as follows: The server, terminal, and user each fulfill their respective roles, providing safe and secure support through voice and image data.

[0586] The terminals are installed in the user's living space and function as robotic devices equipped with microphones and cameras. These devices continuously collect the user's voice and facial expressions and transmit them to a server.

[0587] The server converts this audio data into text data using a speech recognition API such as Google Speech-to-Text. Furthermore, it analyzes the image using image processing techniques such as OpenCV to detect the user's facial expressions and movements. In this process, it utilizes an emotion engine (e.g., Affectiva SDK) to make a multifaceted assessment of the user's emotional state.

[0588] Based on these analysis results, the server sends important notifications regarding the user's status to information terminals, such as those of family members in remote locations. It also generates reassuring voice responses and provides them to the user via the terminal. These voice responses are generated using Text-to-Speech (TTS) technology and are designed to be emotionally resonant to the user.

[0589] For example, if an elderly person feels lonely, the system identifies that emotion from the user's voice and facial expressions, and the server generates an adaptive message such as "Shall we talk?" while simultaneously notifying family members that "Mom seems lonely." Useful prompt phrases in this scenario could include "generating real-time notifications when the elderly person's emotions change and communicating them to family members in friendly language," or "creating responses that can immediately address elderly people who feel anxious or lonely."

[0590] Thus, this system makes it possible to quickly detect anxieties felt by users and take appropriate measures, thereby contributing to improved user safety and security.

[0591] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0592] Step 1:

[0593] The device uses microphones and cameras placed around the user to acquire audio and image data in real time. During this process, the user's speech and facial expressions are obtained as input, converted into digital format, and sent to the server.

[0594] Step 2:

[0595] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology such as the Google Speech-to-Text API. The input is audio data, and the output is the corresponding text information. The audio data is also analyzed concurrently during this process.

[0596] Step 3:

[0597] The server receives image data sent from the terminal and analyzes the user's facial expressions and gestures using image processing techniques such as OpenCV. The input is image data, and the output is data indicating the identified emotional state. In this step, the data is processed using a facial recognition algorithm.

[0598] Step 4:

[0599] The server utilizes an emotion engine to comprehensively analyze the text data and emotion data obtained in steps 2 and 3 to determine the user's emotional state. The input is text data and facial expression data, and the output is an integrated emotion evaluation. At this stage, the tone of voice and facial expression patterns are analyzed.

[0600] Step 5:

[0601] The server generates an appropriate response based on the emotional state. A generative AI model is used to produce the most empathetic response for the user. The input is the previously obtained emotional assessment, and the output is the generated text-formatted response.

[0602] Step 6:

[0603] The server sends an instruction to the terminal to play back the generated response using text-to-speech (TTS) technology. The terminal receives this instruction and performs audio output. The input is a text-based response, and the output is in audio format.

[0604] Step 7:

[0605] The server notifies remote information terminals of the user's status and emotional assessment details, allowing family members and stakeholders to understand the situation in real time. Inputs include emotional assessments and generated notification messages, and output is the sending of notification messages.

[0606] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0607] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0608] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0609] [Fourth Embodiment]

[0610] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0611] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0612] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0613] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0614] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0615] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0616] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0617] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0618] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0619] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0620] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0621] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0622] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0623] This invention is an intelligent system for supporting the lives of the elderly, enabling communication with users through voice and images. It consists of three main components: a server, a terminal (robot), and a user, and is implemented in the following manner.

[0624] The server, located in the cloud, is responsible for advanced processing of speech recognition and image analysis. It converts the user's spoken audio data into text using speech recognition technology, and then analyzes the user's intent and emotional state based on that text data. It also recognizes the user's facial expressions and environmental elements from image data to understand the context. This allows the server to identify the user's needs and generate responses appropriate to those needs.

[0625] The terminal is placed in the living space of the elderly and functions as a user interface. This terminal is equipped with a camera and microphone, which acquire and transmit audio and image data to a server. Furthermore, it can present responses from the server to the user via speech synthesis and provide physical assistance as needed.

[0626] Specifically, when a user asks a question about cooking or everyday life, they speak to the robot via voice, and the robot receives the voice and transmits it to a server. The server converts the voice data into text and generates a contextually appropriate response. For example, if a user says, "I want to make soup," the robot will reply by voice, "What ingredients do you have?" Based on the user's response, the robot then receives more detailed support, such as recipe information, from the server and provides it to the user.

[0627] In this way, this system qualitatively supports the lives of the elderly and promotes independent living, thereby contributing to a reduction in the burden of caregiving. Furthermore, it provides users with an environment where they can live at their own pace with peace of mind through interaction with the robot.

[0628] The following describes the processing flow.

[0629] Step 1:

[0630] The device acquires voice input from the user via a microphone and simultaneously captures image data of the user's facial expressions and environment using a camera. This collected data is then transmitted to the server in real time.

[0631] Step 2:

[0632] The server converts the received audio data into text data using speech recognition technology. During this process, it analyzes the user's emotional state based on the tone and speed of their voice, using this information as foundational data to identify their needs.

[0633] Step 3:

[0634] The server processes the acquired image data using image analysis technology to analyze the user's facial expressions and gestures. Based on this information, it helps understand the user's context and identify their needs.

[0635] Step 4:

[0636] The server integrates information obtained from voice and images to identify the user's needs. Based on the identified needs, it internally determines the appropriate response and support method.

[0637] Step 5:

[0638] The server uses natural language generation technology to create a response that is easy for the user to understand, based on the determined response content. This generated response is then sent to the terminal.

[0639] Step 6:

[0640] The terminal presents the received response to the user as audio using speech synthesis technology. This allows the user to hear and understand the robot's response.

[0641] Step 7:

[0642] The user acts on the robot's responses and instructions, and asks additional questions or requests if necessary. This information is collected by the terminal as feedback that will be used again in the next cycle.

[0643] Step 8:

[0644] The device sends the feedback information back to the server, which then uses this feedback to adjust the system's response or as data to improve the AI ​​model. This process improves the accuracy of the service.

[0645] (Example 1)

[0646] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0647] There is a need for intelligent support systems that enable elderly people to live independently while gaining a sense of security. Conventional systems rely solely on audio or image information, which has the problem of not being able to fully understand the user's intentions and needs. Furthermore, there has been insufficient mechanism for accurately incorporating user feedback into future services.

[0648] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0649] In this invention, the server includes means for analyzing the user's state using voice and images and identifying the user's needs; means for generating and providing a response to the user based on the identified needs; and means for recognizing voice and converting it to text and generating a context-based response. This makes it possible to more accurately understand the user's intentions and emotions and to engage in appropriate dialogue. Furthermore, it enables the efficient collection of feedback and its incorporation into future services, thereby providing sustainable support.

[0650] "Speech recognition technology" is a technology that converts speech data into text data, and it extracts linguistic information by analyzing speech signals.

[0651] "Image recognition technology" is a technology that analyzes image data to recognize specific patterns or objects and makes decisions based on that information.

[0652] "Speech synthesis technology" is a technology that generates natural-sounding speech from text data, producing audio with a texture similar to human speech.

[0653] "Feedback collection" is the process of gathering opinions and reactions from users and using them to improve future services.

[0654] "Contextual understanding" is the process of grasping the background and intent of a conversation based on information obtained from audio and images, and deriving an appropriate response.

[0655] This invention provides an intelligent system to support the daily lives of elderly people. The system consists of three main components: a server, a terminal, and a user. Details are described below.

[0656] server:

[0657] The servers are located in a cloud environment and handle speech recognition and image analysis processing. Speech recognition uses technology to convert audio data into text, and this process utilizes common speech recognition software. For example, technologies such as Google Cloud Speech-to-Text and Amazon Transcribe are used. Image data is analyzed using image recognition technology to understand the user's facial expressions and environment. The servers then understand the user's intentions and needs and generate appropriate responses based on that information. The generated responses are delivered as natural-sounding speech using speech synthesis technology (for example, Amazon Polly or Google Cloud Text-to-Speech).

[0658] Terminal:

[0659] The terminal is placed within the user's living space and acquires audio and image data through its camera and microphone. This data is immediately transmitted to the server. A secure and efficient communication protocol is used to enable real-time data processing. The terminal also has the ability to play back responses from the server using speech synthesis technology and display relevant information on the display as needed.

[0660] Specific user examples:

[0661] When a user asks the robot on their device, "What should I make for dinner tonight?", the device sends the voice message to a server. The server transcribes the voice into text and analyzes the user's intent and emotions. For example, it can offer dinner ideas. The device then plays back the server's response and makes a suggestion to the user, such as, "How about pasta?" This specific dialogue supports the user in their daily activities.

[0662] Examples of generative AI models and prompt statements:

[0663] When using an AI model to generate responses in response to user requests, enter a prompt such as: "The user is asking for dinner ideas. Generate appropriate dish suggestions." Using specific prompts like this ensures that the generated responses align with the user's intent.

[0664] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0665] Step 1:

[0666] The device acquires the user's voice and images using the camera and microphone. The input consists of audio signals and image data, and the output is the digital format of that data. This operation collects fundamental data about the user's intentions and state.

[0667] Step 2:

[0668] The terminal transmits the acquired audio and image data to the server. The input is digitized audio and image data, and the output is data transfer to a server in the cloud.

[0669] Step 3:

[0670] The server receives audio data and converts it into text data using speech recognition technology. The input is audio data, and the output is the corresponding text data. This process makes the user's speech interpretable as text.

[0671] Step 4:

[0672] The server analyzes text data and uses a generative AI model to analyze the user's intent and emotions. The input is text data, and the output is the analysis result. This analysis forms the basis for responses optimized to user needs.

[0673] Step 5:

[0674] The server analyzes image data to recognize the user's facial expressions and surrounding environment. The input is image data, and the output is the result of the recognition of the environment and state. This allows for an understanding of the context of the conversation.

[0675] Step 6:

[0676] The server generates appropriate responses using a generative AI model based on the analysis results of the audio and images. The input is the analyzed intent, emotion, and context, and the output is a specific response.

[0677] Step 7:

[0678] The server sends the generated response to the terminal. The input is the generated response, and the output is data in a format suitable for speech synthesis.

[0679] Step 8:

[0680] The device uses speech synthesis technology to play the generated response as sound to the user. The input is audio data received from the server, and the output is human-readable speech. This allows the user to receive support.

[0681] (Application Example 1)

[0682] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0683] In industrial settings, workers need to prepare necessary tools and materials quickly and accurately in order to carry out their work efficiently. However, if workers have to go and find the necessary tools themselves, not only will work efficiency decrease, but the risk of errors due to incorrect tool selection and safety risks will also increase. This invention aims to solve these problems and improve work efficiency and safety.

[0684] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0685] In this invention, the server includes means for analyzing the user's state using voice and images and identifying their needs; means for generating and providing a response to the user based on the identified needs; means for collecting feedback from the user after the service has been provided and incorporating it into the next service; means for analyzing the voice instructions of workers in an industrial setting and selecting appropriate tools and materials; and means for providing the selected tools and materials to the workplace. This automates the selection and supply of items to enable workers to perform their tasks efficiently, thereby improving work efficiency and safety.

[0686] "Audio and images" refers to audio and image data obtained from users, and serves as a source of information for analyzing the user's state and needs.

[0687] "Means of identifying needs" refers to the process of analyzing acquired audio and image data to clarify the requests and requirements that users have.

[0688] "Means for generating and providing responses to users" refers to a system for creating appropriate information and support based on identified needs and communicating them to users.

[0689] "Means for collecting feedback and reflecting it in future services" refers to the process of collecting evaluations and attitudes from users after a service has been provided, and using that information to improve future services.

[0690] "Means for analyzing voice instructions given by workers in industrial settings" refers to technologies and processes for understanding and analyzing the content of voice instructions given by workers in industrial settings.

[0691] "Means for selecting appropriate tools and materials" refers to the process of determining and selecting the necessary tools and materials based on the worker's voice instructions.

[0692] "Means of providing selected tools and materials to the workshop" refers to a system for delivering selected tools and materials to the workshop at the appropriate time.

[0693] The system implementing this invention aims to efficiently support work in industrial settings. The system understands the needs of workers through the analysis of voice and image data and provides the necessary tools and materials.

[0694] The server is located in the cloud and has advanced speech recognition and image analysis capabilities. Speech recognition APIs (e.g., Google Speech-to-Text) are used to convert speech data into text and determine the worker's intent. Furthermore, OpenCV is used for image recognition technology to accurately identify necessary tools and materials. Based on the acquired information, the server generates appropriate responses and provides them to the user.

[0695] The terminal is installed in the industrial site and functions as an interface to assist with work. This terminal is equipped with a camera and microphone, which acquire and transmit audio and image data to a server. It also uses speech synthesis to communicate responses from the server to the worker and provides physical assistance, such as transporting tools and materials, as needed.

[0696] The user, or worker, gives voice instructions to the system. Based on these instructions, the system analyzes the worker's requests and prepares and provides the necessary items. This process allows the worker to concentrate on their work, improving efficiency while also ensuring safety.

[0697] For example, if a worker says, "Bring me the wrench that fits this bolt," the server can identify the appropriate wrench, send an instruction to the terminal, and have that wrench delivered to the worker.

[0698] An example of a prompt message could be, "List the most frequently used tools in the factory and use that to optimize word selection for speech recognition." This would further improve the accuracy of voice commands and enable more efficient work support.

[0699] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0700] Step 1:

[0701] The terminal receives voice instructions from the user (worker) via a microphone. This voice data becomes the input. The terminal temporarily stores this voice data in a buffer and prepares to send it to the server in the next step.

[0702] Step 2:

[0703] The terminal sends the audio data acquired in step 1 to the server. The server uses a speech recognition API to convert the audio data into text. This converted text is the output and is used as the basis for determining the worker's intent.

[0704] Step 3:

[0705] The server uses a generative AI model to analyze the worker's intent from the text data in step 2. This process understands the context of words and phrases and determines the necessary tools and materials. As a result of this analysis, information on the selected tools and materials is output.

[0706] Step 4:

[0707] The server uses image recognition technology to generate specific instructions for selecting particular tools and materials based on the information from step 3. Here, location information and identification tags for tools and materials are extracted from the image data, and the information of the selected items is organized.

[0708] Step 5:

[0709] The terminal receives the specific tool and material information generated in step 4 and provides feedback to the user via speech synthesis. It provides the worker with responses such as "I'll bring the XX wrench," building trust in the system.

[0710] Step 6:

[0711] The terminal actually transports the identified tools and materials based on the worker's instructions. In this process, the selected tools are moved safely and accurately by the mounted arm or transport device. Once everything is complete, the worker is notified and the terminal awaits further instructions.

[0712] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0713] This invention is a support system that enables interaction with users using voice and image data, and incorporates an emotion engine that recognizes the user's emotions. The system consists of three main elements: a server, a terminal (robot), and a user, and is implemented in the following manner.

[0714] The device is placed in the living space of elderly people and acquires voice and image data in real time. When a user speaks to the robot, the voice input is captured by the device's microphone. The device also captures the user's facial expressions and surrounding environment through its camera. This data is immediately transmitted to a server.

[0715] The server converts received audio data into text using speech recognition technology, while simultaneously analyzing the user's emotional state from their speech using an emotion engine. Furthermore, image data is processed using image recognition technology to analyze the user's facial expressions and gestures. In this way, the server integrates audio and image data to understand the context and identify the user's needs and emotions.

[0716] The emotion engine uses a multifaceted approach, including voice tone analysis and facial expression analysis, to determine the user's emotions and decide on a response based on those emotions. Based on these analysis results, the server generates an appropriate response and sends it to the terminal.

[0717] The terminal presents responses from the server to the user in voice using speech synthesis technology. The responses are tailored to the user's emotions and designed to provide a sense of reassurance. For example, if the user says, "I'm lonely today," the robot will gently respond, "Shall we talk?" to cheer the user up.

[0718] In this way, this system understands users' needs and emotions and provides personalized support, thereby creating an environment where seniors can live independent and happy lives. Furthermore, by utilizing emotional information, the content of subsequent service provision can be made more individualized, leading to continuous improvement in satisfaction.

[0719] The following describes the processing flow.

[0720] Step 1:

[0721] The device acquires voice input from the user via the microphone and captures image data of the user's facial expressions and environment using the camera. The acquired data is immediately sent to the server.

[0722] Step 2:

[0723] The server receives the audio data and converts it into text using speech recognition technology. During this process, it also analyzes the tone and speed of the speech to infer the user's emotional state.

[0724] Step 3:

[0725] The server processes image data using image analysis technology to analyze the user's facial expressions and gestures. This deepens the understanding of the user's emotions and situation.

[0726] Step 4:

[0727] The server inputs information obtained from voice and images into the emotion engine to identify the user's emotions. Based on this, it integrates the user's needs and emotions to determine a response.

[0728] Step 5:

[0729] The server uses natural language generation technology to create responses that match the user's emotions. The generated responses are designed to be polite, considerate, and thoughtful.

[0730] Step 6:

[0731] The server sends the generated response data to the terminal. The response is sent as text data, but the server also sends additional data for audio playback as needed.

[0732] Step 7:

[0733] The terminal converts the response received from the server into speech using speech synthesis technology and presents it to the user. This allows the user to receive an intuitive and emotionally empathetic response.

[0734] Step 8:

[0735] The user decides on their next action based on the response provided and provides feedback to the device as needed. This feedback is also acquired by the device and sent to the server for future service improvements.

[0736] Step 9:

[0737] The server analyzes the collected feedback information and adjusts the parameters of the AI ​​model and emotion engine to improve the accuracy and personalization of future responses.

[0738] (Example 2)

[0739] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0740] There is a need to understand the emotional state of individuals who require support in their daily lives, such as the elderly and those who feel lonely, in real time and to respond appropriately. With conventional technology, it was difficult to effectively integrate voice and image data and provide personalized responses that were tailored to the user, which severely limited the improvement in satisfaction.

[0741] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0742] In this invention, the server includes means for acquiring the user's voice and image using a terminal and transmitting the data to the server; means for analyzing the data using speech recognition and image recognition technologies on the server and creating a response using generation technology; and means for outputting the generated response to the user as voice using the terminal. This makes it possible to quickly and accurately grasp the user's emotions, provide individually appropriate support, and improve user satisfaction.

[0743] A "terminal" refers to a device that has the function of acquiring the user's voice and images in real time and transmitting them to a server.

[0744] A "server" refers to a central computing unit that analyzes audio and image data transmitted from terminals and creates responses using generation technology.

[0745] "Speech recognition technology" is a technology that converts speech data into text and is used to analyze the content of a user's speech.

[0746] "Image recognition technology" refers to technology that analyzes a user's facial expressions and gestures from image data to understand their situation.

[0747] "Generative technology" is a technology that generates appropriate responses to provide to the user based on analysis results, and is used to realize personalized interactions.

[0748] "Emotional state" refers to the internal psychological state inferred from the user's tone of voice, facial expressions, and gestures.

[0749] This invention relates to a system that utilizes voice and image data to enable interaction with users and incorporates an emotion engine that recognizes the user's emotions. This system consists of three main elements: a server, a terminal (robot), and a user.

[0750] The terminals are placed in the living spaces of elderly individuals and are equipped with hardware and software for acquiring voice and image data in real time. Specifically, they are equipped with high-performance microphones and cameras that capture the voice of the user speaking, as well as the user's facial expressions and surrounding environment. This acquired data is transmitted to a server via secure wireless communication.

[0751] The server employs various technologies to analyze the received data. Audio data is converted to text using speech recognition software. This process utilizes services such as the Google Speech-to-Text API to achieve high-precision speech recognition. Furthermore, an emotion engine analyzes the text data and uses natural language processing techniques to determine the user's emotional state. Image data is analyzed using image recognition technologies such as OpenCV and TensorFlow to read detailed emotional states from the user's facial expressions and gestures.

[0752] Based on these analysis results, the server uses a generative AI model (e.g., GPT-4) to generate an appropriate response for the user. An example of a specific prompt might be, "The user appears depressed. Please generate a kind comment to cheer them up." The generated response is then sent back to the terminal.

[0753] The device uses speech synthesis technology to convert responses from the server into speech and presents them to the user. To achieve this, it utilizes speech synthesis systems such as Amazon Polly and Google Text-to-Speech to communicate in a natural and friendly voice. The responses are designed to be empathetic to the user's emotions and provide a sense of security. For example, if a user says, "I'm lonely today," the device will gently ask, "Shall we talk?" In this way, the entire system supports the user's life and helps them achieve a happy and independent life.

[0754] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0755] Step 1:

[0756] The device acquires the user's voice and images. It receives the user's spoken audio and video as input, and outputs digital audio and image data. Specifically, it converts sound waves into digital audio using the microphone and captures video with the camera. This data is then compressed and transmitted to the server wirelessly.

[0757] Step 2:

[0758] The server converts audio data transmitted from the terminal into text using speech recognition technology. It receives digital audio data as input and obtains text data as output. Specifically, it uses speech recognition software to perform feature extraction and phoneme matching to convert the audio into accurate text.

[0759] Step 3:

[0760] The server processes the text data obtained through speech recognition into an emotion analysis engine to determine the user's emotions. It receives text data as input and obtains an emotion state as output. Specifically, it uses a natural language processing engine to analyze keywords and context, and assigns emotion labels using an emotion dictionary and a machine learning model.

[0761] Step 4:

[0762] The server analyzes image data using image recognition technology to evaluate the user's facial expressions and gestures. It receives image data as input and outputs facial expression data and gesture data. Specifically, it uses a face detection algorithm to identify facial expressions and performs gesture recognition to extract features of hand movements and posture.

[0763] Step 5:

[0764] The server integrates emotional state and facial expression data obtained from audio and images, and uses a generative AI model to generate an appropriate response. It receives emotional state and facial expression data as input and obtains a text response as output. Specifically, it creates a prompt for the generative AI model and constructs a text response based on that prompt.

[0765] Step 6:

[0766] The server sends the generated text response to the terminal. It receives the generated text as input and sends a response to the terminal in an appropriate format as output. Specifically, it performs the process of packetizing digital data and using transmission control protocols to safely deliver the data to the terminal.

[0767] Step 7:

[0768] The terminal converts received text responses into speech using speech synthesis technology and presents them to the user. It receives text responses as input and obtains speech output as output. Specifically, it uses a speech synthesis system to play the text in a natural-sounding voice and speaks it aloud to the user through a speaker.

[0769] (Application Example 2)

[0770] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0771] In supporting the independent living of the elderly, it is crucial to provide appropriate support that responds to their emotions and needs, fostering a sense of security. However, conventional support systems have struggled to respond immediately to changes in users' emotions or emergencies, and have not been able to adequately guarantee the safety and security of elderly people living alone. There is a need to address these challenges and further improve the safety and security of the elderly.

[0772] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0773] In this invention, the server includes means for analyzing the user's state using voice and images and identifying their needs; means for generating and providing a response to the user based on the identified needs; means for collecting feedback from the user after the service has been provided and incorporating it into the next service; and means for monitoring the user's safety status in real time and notifying a remote information terminal. This enables immediate responses to anxieties and safety concerns felt by the elderly, allowing family members and related parties in remote locations to quickly recognize the situation and take appropriate measures.

[0774] "Voice" refers to the words and sounds spoken by the user to the robot or system, and is important data for understanding the user's intentions and emotions.

[0775] "Images" refer to visual information obtained through a camera, such as the user's facial expressions and surroundings, and are used as materials to understand the user's emotions and situation.

[0776] "Analyzing the user's state" refers to the process of analyzing voice and image data to determine the user's current emotions, health, or safety.

[0777] "Means of identifying needs" refers to methodologies for identifying the services and support that users require from analyzed data.

[0778] "Means of generating and providing responses" refers to the process by which a system determines an appropriate response or action based on the user's needs and emotions, and communicates it to the user.

[0779] "Collecting feedback" refers to the process of gathering user reactions and opinions on the services and responses provided, which is used to improve future services.

[0780] "Real-time monitoring of safety conditions" refers to the process of constantly monitoring users' living spaces and activities to detect abnormalities and emergencies.

[0781] "Notification to remote information terminals" refers to a communication method that sends important information about the user's status to devices of designated family members or related parties located remotely.

[0782] In an embodiment of the present invention, the system is configured as follows: The server, terminal, and user each fulfill their respective roles, providing safe and secure support through voice and image data.

[0783] The terminals are installed in the user's living space and function as robotic devices equipped with microphones and cameras. These devices continuously collect the user's voice and facial expressions and transmit them to a server.

[0784] The server converts this audio data into text data using a speech recognition API such as Google Speech-to-Text. Furthermore, it analyzes the image using image processing techniques such as OpenCV to detect the user's facial expressions and movements. In this process, it utilizes an emotion engine (e.g., Affectiva SDK) to make a multifaceted assessment of the user's emotional state.

[0785] Based on these analysis results, the server sends important notifications regarding the user's status to information terminals, such as those of family members in remote locations. It also generates reassuring voice responses and provides them to the user via the terminal. These voice responses are generated using Text-to-Speech (TTS) technology and are designed to be emotionally resonant to the user.

[0786] For example, if an elderly person feels lonely, the system identifies that emotion from the user's voice and facial expressions, and the server generates an adaptive message such as "Shall we talk?" while simultaneously notifying family members that "Mom seems lonely." Useful prompt phrases in this scenario could include "generating real-time notifications when the elderly person's emotions change and communicating them to family members in friendly language," or "creating responses that can immediately address elderly people who feel anxious or lonely."

[0787] Thus, this system makes it possible to quickly detect anxieties felt by users and take appropriate measures, thereby contributing to improved user safety and security.

[0788] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0789] Step 1:

[0790] The device uses microphones and cameras placed around the user to acquire audio and image data in real time. During this process, the user's speech and facial expressions are obtained as input, converted into digital format, and sent to the server.

[0791] Step 2:

[0792] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology such as the Google Speech-to-Text API. The input is audio data, and the output is the corresponding text information. The audio data is also analyzed concurrently during this process.

[0793] Step 3:

[0794] The server receives image data sent from the terminal and analyzes the user's facial expressions and gestures using image processing techniques such as OpenCV. The input is image data, and the output is data indicating the identified emotional state. In this step, the data is processed using a facial recognition algorithm.

[0795] Step 4:

[0796] The server utilizes an emotion engine to comprehensively analyze the text data and emotion data obtained in steps 2 and 3 to determine the user's emotional state. The input is text data and facial expression data, and the output is an integrated emotion evaluation. At this stage, the tone of voice and facial expression patterns are analyzed.

[0797] Step 5:

[0798] The server generates an appropriate response based on the emotional state. A generative AI model is used to produce the most empathetic response for the user. The input is the previously obtained emotional assessment, and the output is the generated text-formatted response.

[0799] Step 6:

[0800] The server sends an instruction to the terminal to play back the generated response using text-to-speech (TTS) technology. The terminal receives this instruction and performs audio output. The input is a text-based response, and the output is in audio format.

[0801] Step 7:

[0802] The server notifies remote information terminals of the user's status and emotional assessment details, allowing family members and stakeholders to understand the situation in real time. Inputs include emotional assessments and generated notification messages, and output is the sending of notification messages.

[0803] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0804] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0805] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0806] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0807] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0808] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0809] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0810] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0811] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0812] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0813] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0814] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0815] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0816] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0817] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0818] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0819] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0820] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0821] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0822] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0823] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0824] The following is further disclosed regarding the embodiments described above.

[0825] (Claim 1)

[0826] A means of analyzing the user's state using voice and images and identifying their needs,

[0827] A means of generating and providing responses to users based on identified needs,

[0828] A means of collecting user feedback after providing a service and incorporating it into the next service,

[0829] A system that includes this.

[0830] (Claim 2)

[0831] The system according to claim 1, comprising means for converting voice data into text using speech recognition technology and determining the user's emotional state and intentions.

[0832] (Claim 3)

[0833] The system according to claim 1, comprising means for analyzing the user's facial expressions and gestures using image recognition technology and understanding their situation.

[0834] "Example 1"

[0835] (Claim 1)

[0836] A means of analyzing the user's state using audio and images and identifying the user's needs,

[0837] A means of generating and providing responses to users based on identified needs,

[0838] A means for recognizing speech, converting it to text, and generating context-based responses,

[0839] A means of analyzing the user's emotional state and intentions and generating an appropriate response within a reproducible range,

[0840] A means of analyzing image data to understand the user's surrounding environment,

[0841] A means of providing the generated response using speech synthesis technology,

[0842] A means of collecting user feedback after providing a service and incorporating it into the next service,

[0843] A system that includes this.

[0844] (Claim 2)

[0845] The system according to claim 1, comprising means for converting voice data into text using speech recognition technology and determining the user's emotional state and intentions.

[0846] (Claim 3)

[0847] The system according to claim 1, comprising means for analyzing the user's facial expressions and gestures using image recognition technology to understand the situation, as well as means for performing environmental recognition.

[0848] "Application Example 1"

[0849] (Claim 1)

[0850] A means of analyzing the user's state using voice and images and identifying their needs,

[0851] A means of generating and providing responses to users based on identified needs,

[0852] A means of collecting user feedback after providing a service and incorporating it into the next service,

[0853] A method for analyzing workers' voice instructions in an industrial setting to select appropriate tools and materials,

[0854] Means of providing selected tools and materials to the workshop,

[0855] A system that includes this.

[0856] (Claim 2)

[0857] The system according to claim 1, comprising means for converting voice data into text using speech recognition technology and determining the user's emotional state and intentions.

[0858] (Claim 3)

[0859] The system according to claim 1, comprising means for analyzing the user's facial expressions and gestures using image recognition technology and understanding their situation.

[0860] "Example 2 of combining an emotion engine"

[0861] (Claim 1)

[0862] A means of analyzing the user's state using voice and images and identifying their needs,

[0863] A means of generating and providing responses to users based on identified needs,

[0864] A means of collecting user feedback after providing a service and incorporating it into the next service,

[0865] A means for acquiring the user's voice and images using a terminal and transmitting that data to a server,

[0866] A means for analyzing data using speech recognition and image recognition technologies on a server and creating a response using generation technology,

[0867] A means of outputting a response generated using a terminal to the user as audio,

[0868] A system that includes this.

[0869] (Claim 2)

[0870] The system according to claim 1, comprising means for converting voice data into text using speech recognition technology and determining the user's emotional state and intentions.

[0871] (Claim 3)

[0872] The system according to claim 1, comprising means for analyzing the user's facial expressions and gestures using image recognition technology and understanding their situation.

[0873] "Application example 2 when combining with an emotional engine"

[0874] (Claim 1)

[0875] A means of analyzing the user's state using voice and images and identifying their needs,

[0876] A means of generating and providing responses to users based on identified needs,

[0877] A means of collecting user feedback after providing a service and incorporating it into the next service,

[0878] A means of monitoring the user's safety status in real time and notifying information terminals in remote locations,

[0879] A system that includes this.

[0880] (Claim 2)

[0881] The system according to claim 1, comprising means for converting voice data into text using speech recognition technology and determining the user's emotional state and intentions.

[0882] (Claim 3)

[0883] The system according to claim 1, comprising means for analyzing the user's facial expressions and gestures using image recognition technology and understanding their situation. [Explanation of Symbols]

[0884] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of analyzing the user's state using voice and images and identifying their needs, A means of generating and providing responses to users based on identified needs, A means of collecting user feedback after providing a service and incorporating it into the next service, A system that includes this.

2. The system according to claim 1, comprising means for converting voice data into text using speech recognition technology and determining the user's emotional state and intentions.

3. The system according to claim 1, comprising means for analyzing the user's facial expressions and gestures using image recognition technology and understanding their situation.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A