system
A system addresses the lack of continuous care support by analyzing user data to provide personalized responses and emergency alerts, reducing caregiver burden and ensuring timely assistance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing systems fail to provide 24/7 reliable care advice and psychological support to younger generations engaged in home care, leading to psychological burden and information shortages.
A system that analyzes user input data, estimates intentions and psychological states, and generates individualized responses, with emergency reporting capabilities to external support agencies.
Reduces caregiver burden and provides timely, tailored support by analyzing user inputs through natural language and multimodal processing, with emergency notification features.
Smart Images

Figure 2026073512000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] To solve the problem that it is difficult to relieve the psychological burden and information shortage of the younger generation engaged in home care and to provide reliable care advice and psychological support 24 hours a day, 365 days a year
Means for Solving the Problems
[0005] Adopt a technology for analyzing various input data from a user, estimating the intention and psychological state thereof, and generating an individualized response based on the analysis result. In addition, when detecting an emergency, it is equipped with a function of automatically reporting to an external support agency, thereby reducing the burden on the caregiver and providing necessary support promptly.
[0006] A "communication terminal" refers to an electronic device used by a user to send and receive input data.
[0007] "Input data" refers to information transmitted by the user through a communication terminal, and includes text, audio, and images.
[0008] A "server" refers to a computer system that analyzes received input data and generates an appropriate response.
[0009] "Natural language processing" refers to the technology that enables computers to understand and analyze human language.
[0010] "Multimodal analysis technology" refers to a technology that integrates and analyzes data in multiple different formats, such as audio and images.
[0011] "Response generation" refers to the process of creating appropriate answers or instructions based on user input data.
[0012] "Speech synthesis technology" refers to the technology that converts text information into speech and presents it to the user.
[0013] "External support organizations" refer to organizations or facilities that provide assistance or care for users in emergency situations. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5]It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the language used in the following description will be explained.
[0017] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0018] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0019] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention is a system for processing information obtained from users and providing appropriate care and psychological support. The embodiments thereof are described below.
[0036] The system of the present invention mainly consists of a server, a terminal, and a user interface. The user creates input data in text, voice, or image format using a communication terminal. For example, the user can input text requesting advice on caregiving or upload images of a specific situation.
[0037] The terminal's role is to acquire user input data in real time and send this data to the server. The transmitted data is encrypted and managed securely.
[0038] Upon receiving input data from a terminal, the server analyzes the text using natural language processing (NLP) techniques and determines the content of images and audio using multimodal analysis techniques. Based on this, the server estimates the user's intentions and the support they require, and understands their psychological state through sentiment analysis. Through these analyses, the server generates an appropriate response to the user's request.
[0039] The generated response is sent back to the terminal, where it is displayed to the user. In some cases, the response may be converted from text to speech and provided verbally. In particular, if the server determines that the user is in an emotionally distressed situation, it will generate a response that takes emergency measures into consideration.
[0040] For example, if a user enters a message such as, "My father is depressed, and I don't know what to do," the server analyzes this message and determines that the user needs psychological support within the family. The server generates relevant support information as a response, providing the user with "specific support measures" and "referrals to professional organizations as needed." Furthermore, if the server determines that the situation is urgent, it can automatically notify relevant external support organizations to encourage prompt action.
[0041] In this way, the system of the present invention provides flexible and rapid support that responds to diverse needs, and reduces the psychological and informational burden on caregivers.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The user inputs questions and situation descriptions in text, voice, or image format via a communication terminal. The communication terminal converts this input data into digital signals.
[0045] Step 2:
[0046] The terminal transmits the converted digital signal to the server in real time. This communication is encrypted to ensure security.
[0047] Step 3:
[0048] The server analyzes the received data. For text, natural language processing techniques are used to perform grammatical analysis and keyword extraction. Audio data is converted to text using speech recognition technology, and image data is identified using image recognition algorithms.
[0049] Step 4:
[0050] The server performs sentiment analysis to understand the user's intentions based on the analyzed data. In this process, the user's emotional state is classified as positive, negative, or neutral, and the server adjusts the response accordingly.
[0051] Step 5:
[0052] The server generates a response based on the sentiment analysis results and the user's request. This generation process references a knowledge base and constructs a response that includes the latest care and psychological support information.
[0053] Step 6:
[0054] The server sends the generated response to the terminal. The terminal displays the received text response to the user and, if necessary, also provides a voice response using speech synthesis technology.
[0055] Step 7:
[0056] If the server determines that a user is in an emergency, it automatically notifies registered external support organizations. This notification includes concise and important information, enabling a swift response.
[0057] Step 8:
[0058] The terminal notifies the user of responses from the server and confirmation of emergency calls initiated by the server, preparing the user for their next action.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] Through the input and analysis of information, there is a challenge in providing appropriate and prompt responses when users require care or psychological support. Furthermore, providing individually customized support tailored to the emotional state of each user is also a crucial challenge.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for collecting data from users via information terminals, means for processing the data using natural language processing and various forms of analysis techniques in a data processing device, and means for evaluating the user's purpose and mental state based on the processing results. This enables the rapid provision of customized support tailored to the user's needs and allows for coordinated responses with external support organizations in emergencies.
[0064] An "information terminal" is a general term for electronic devices used by users to input and display data.
[0065] "User" refers to the entity that uses the system to input data and receives analysis results and responses.
[0066] A "data processing device" refers to a device that receives data transmitted from an information terminal and performs analysis and processing on it.
[0067] "Natural language processing" refers to the technology used by data processing devices to analyze human language and extract its meaning.
[0068] "Analysis techniques for diverse formats" refers to technologies that effectively analyze data in different formats, such as text, audio, and images.
[0069] "Evaluation" refers to the act of judging a user's purpose and mental state based on the analysis results.
[0070] "Response" refers to the response provided to the user based on the analysis results and evaluation.
[0071] "Speech synthesis technology" refers to the technology that converts text data into speech and outputs it as speech.
[0072] An "external support organization" refers to an external, specialized support organization that can integrate with the system in an emergency.
[0073] This invention is a system for providing flexible and rapid care and psychological support to users. It mainly includes information terminals, servers, and user-involved processes.
[0074] Users input information requiring support in text, voice, or image format using an information terminal. For example, they can input a question such as, "My mother hasn't been feeling well lately, and I'm worried," or images related to a specific situation. The information terminal collects this input data and securely transmits it to a data processing unit using encryption technology.
[0075] The server functions as a data processing unit, processing received data using specific analysis techniques. It analyzes text data using natural language processing (NLP) techniques and evaluates audio and image data in detail using multimodal analysis techniques.
[0076] The server utilizes a generative AI model to analyze the results and build support content optimized for the user. For example, if a user enters the sentence, "My father is depressed, and I don't know what to do," the server analyzes this and determines that psychological support is needed. It then provides appropriate support measures and information on partner organizations.
[0077] An example of a prompt message a user might use for input is, "My father is depressed, and I don't know what to do." Such input allows the system to respond to the user's individual needs. Furthermore, in emergencies, the server can assess the user's situation and automatically notify external support organizations as needed. This allows users to utilize the system with peace of mind.
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] Users create input data using an information terminal. For example, they might enter text describing the situation in which they need support, or related images. The data entered by the user is the first step toward providing accurate information services and is collected by the terminal after input.
[0081] Step 2:
[0082] The device acquires data entered by the user in real time. This data is encrypted to ensure security. The device then transmits the collected data to a server via a secure protocol. Input data can be in text, audio, or image format, while output is encrypted communication data.
[0083] Step 3:
[0084] The server receives encrypted data sent from the terminal and first decrypts it. Then, the server uses natural language processing (NLP) techniques to analyze the text data and process it to understand the type of support the user is seeking. It also performs detailed analysis of audio and image data using multimodal analysis. The input is encrypted data, and the output is information corresponding to the user's needs after the analysis.
[0085] Step 4:
[0086] The server determines the user's intentions and emotions based on the analyzed data. Using a generative AI model, it generates appropriate responses tailored to the user. At this stage, the server constructs specific support measures or information from the analysis results. The input is the analyzed user information, and the output is the generated response data.
[0087] Step 5:
[0088] The response data generated by the server is sent to the terminal. The terminal displays this response to the user. If voice output is required, the terminal also responds to the user verbally via text-to-speech synthesis. The input is the response data from the server, and the output is the displayed or audio information delivered to the user.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] There is a need to understand changes in a person's psychological and physical state in real time and provide appropriate support. However, conventional systems have had difficulty instantly detecting changes in a user's condition and providing specific support tailored to their living environment.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes means for acquiring input information from a human via a communication device, means for estimating the human's intentions and psychological state based on the analysis results, and means for generating an appropriate response based on the analyzed information. This makes it possible to monitor changes in the user's living environment and mood in real time and provide specific and appropriate support.
[0094] A "communication device" is a device that acquires input information from a human and transmits it to a central processing unit.
[0095] "Human" refers to the end-users who use the system, and their intentions and psychological states are analyzed through their input.
[0096] "Input information" refers to data provided by humans in various forms, such as text, audio, and images.
[0097] A "central processing unit" is a computer system that performs overall language processing and diverse analysis based on acquired input information and generates analysis results.
[0098] "Holistic language processing" is a technique that uses natural language processing technology to analyze input text data and understand its meaning.
[0099] "Diverse analysis technology" is a technology that analyzes image and audio data to identify and interpret its content.
[0100] "Intention" refers to the purpose or request that a person tries to convey through a system.
[0101] "Psychological state" refers to a person's emotions and mental state, and by estimating this, it becomes possible to provide appropriate support.
[0102] "Response" refers to support information and advice generated by the central processing unit based on the analysis results.
[0103] A "display device" is a device that monitors a person's living environment and circumstances, and displays a response visually or audibly as needed.
[0104] A "support organization" is an external professional organization or institution that is referred to when needed and provides support to individuals.
[0105] The system for realizing this invention consists of a communication device, a central processing unit, a display device, and cooperation with a support organization. First, the communication device acquires input information from humans, i.e., text, voice, and images, and transmits it to the central processing unit. This information is encrypted and securely sent to the central processing unit.
[0106] The central processing unit analyzes the acquired input information using holistic language processing technologies (such as TENSORFLOW® and spaCy) and diverse analysis technologies (such as OpenCV and Librosa). This allows for accurate estimation of human intentions and psychological states. For example, it can estimate emotional changes from voice input or determine everyday behavior through image analysis.
[0107] Based on the analysis results, the central processing unit generates the optimal response. This response is presented to a human using speech synthesis technology as needed. In emergencies, it can also automatically notify the support organizations that require assistance.
[0108] For example, if a user inputs "I've been feeling unwell lately and having difficulty with daily tasks," the central processing unit can analyze this data and provide suggestions for relaxation or referrals to medical institutions. This allows users to receive timely and appropriate support.
[0109] An example of a prompt might be, "Generate a program that suggests relaxation methods when the user appears tired." Such prompts allow the generating AI model to formulate processing steps to indicate appropriate intervention methods.
[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0111] Step 1:
[0112] The user provides input information using a communication device. This information can be in text, audio, or image format, and different initial processing is performed depending on the format. The user's input is temporarily stored by the communication device. Subsequently, this input information is transmitted securely to the central processing unit using encryption technology.
[0113] Step 2:
[0114] The server acquires input information received from the communication device and performs initial analysis according to the data format. In the case of audio data, it is converted to text using speech recognition technology, and for image data, key features are extracted using image processing technology. This allows the server to generate a standardized data format for subsequent processing.
[0115] Step 3:
[0116] The server uses a global language processing engine to analyze the meaning of text data. This process extracts keywords and context from the input information to understand the user's intent. For example, if keywords such as "tired" are included, the server proceeds to estimate the related psychological state based on these keywords.
[0117] Step 4:
[0118] The server uses diverse analysis techniques to detect the user's daily activities and changes in facial expressions from image data. This generates complementary data that allows for more accurate emotion estimation based on the information obtained from the images. This data makes it possible to visually understand the user's situation.
[0119] Step 5:
[0120] Based on the analysis results, the server applies a generative AI model to generate the optimal response. This model performs inference according to the prompt text and designs a response that is appropriate to the user's state. The response is generated in text format and then converted to speech using speech synthesis technology as needed.
[0121] Step 6:
[0122] The server sends the generated response to the communication device. The communication device displays the response as text or plays it back as audio. Specifically, it provides suggestions for relaxation methods and contact information for support organizations as needed. As a result, the user is encouraged to take appropriate action and is in a position to receive support.
[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0124] This invention is a novel care support system that incorporates an emotion engine to recognize the user's emotions and provide highly personalized support based on those emotions. The system includes a communication terminal, a server, and an emotion engine.
[0125] Users can input questions and information about caregiving using a communication terminal. The input data is appropriately digitized on the terminal and sent to the server. Upon receiving the input data, the server first analyzes it using natural language processing and multimodal analysis techniques. This analysis involves sentence structure analysis if the data is text, speech recognition if it is audio, and image recognition if it is images to understand the content.
[0126] After this data analysis, the server uses an emotion engine to identify emotions from the user's input. The emotion engine analyzes, for example, the tone of text, the intonation of speech, and facial expressions in images, and classifies them into emotion categories such as positive, negative, and neutral. Based on the results of the emotion engine, the user's psychological state is accurately understood.
[0127] The server then synthesizes these analysis results and proceeds to a process of generating a response that aligns with the user's intentions. This response generation process includes considering the emotions identified by the emotion engine and creating a detailed response that matches the support the user desires and their psychological state. The generated response is sent to the terminal as text and, in some cases, presented as speech using speech synthesis technology.
[0128] For example, if a user enters "My mother hasn't been feeling well lately, and I'm worried about spending time with her," the server analyzes this message and uses an emotion engine to read the user's anxiety. The server then generates emotional support and practical advice that corresponds to this emotion, providing a response such as, "I recommend you wait and see for a while, and consult your doctor if necessary. We can also think of activities to cheer her up together."
[0129] This system allows users to receive care support that is tailored to their emotional needs, thereby reducing the burden of daily caregiving.
[0130] The following describes the processing flow.
[0131] Step 1:
[0132] Users input questions or concerns in text, voice, or image format using a communication terminal. The communication terminal digitizes the input data and prepares it to be immediately transmitted to the server.
[0133] Step 2:
[0134] The terminal transmits digitized input data to the server. The communication is encrypted to ensure the security of user data.
[0135] Step 3:
[0136] The server analyzes the received data. For text, it uses natural language processing for syntactic analysis and keyword extraction; for speech, it uses speech recognition to convert it to text; and for images, it uses image recognition technology to determine visual elements.
[0137] Step 4:
[0138] The server passes the analyzed data through the emotion engine. The emotion engine analyzes the emotions contained in the user's input and assigns one of three emotion labels: positive, negative, or neutral.
[0139] Step 5:
[0140] The server comprehensively understands the user's intent and emotions based on the results of natural language processing and the emotion engine. It then moves on to a process of generating the most appropriate response for the user based on this information.
[0141] Step 6:
[0142] The server uses a generative AI model to create responses that align with the user's intentions. These responses include emotional support and practical advice based on the emotions indicated by the emotion engine.
[0143] Step 7:
[0144] The server sends the completed response to the terminal. The terminal displays the received response to the user. If necessary, the response can also be provided in voice using speech synthesis technology.
[0145] Step 8:
[0146] Under the monitoring of the emotion engine, if the server determines that user input is urgent, it will automatically notify pre-configured external support organizations and initiate procedures to request prompt assistance.
[0147] (Example 2)
[0148] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0149] In modern society, caregivers are required to respond quickly and appropriately to increasingly complex needs. However, caregivers often find it difficult to understand the underlying emotions and psychological states of users, limiting their ability to provide appropriate support. This problem highlights the need to develop systems that enable even non-professionals to provide appropriate advice and support tailored to the situation.
[0150] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0151] In this invention, the server includes means for analyzing information from the user using language processing technology and various data analysis technologies, means for identifying the user's abstract intentions and emotional state based on the analysis results, and means for creating an individualized response based on the identified information. This makes it possible to quickly provide individualized support that is in line with the user's emotions and situation, reduce the burden on caregivers, and improve the quality of care.
[0152] A "communication device" is a device that acquires information from a user and transmits it to a data processing device.
[0153] A "data processing device" is a device that analyzes received information to identify the user's intentions and emotional state.
[0154] "Language processing technology" refers to techniques for analyzing text data and understanding its structure and meaning.
[0155] "Diverse data analysis technologies" refer to technologies for analyzing data in various formats, such as text, audio, and images, and understanding their content.
[0156] A "personalized response" is a response that is generated in a way that is tailored to the user's specific situation and emotions.
[0157] "Abstract user intent" refers to the goals or needs that users have implicitly but do not explicitly express.
[0158] "Emotional state" refers to the psychological state a user exhibits, such as positive, negative, or neutral emotions.
[0159] One embodiment of the present invention is configured as a care support system incorporating emotion recognition technology. This system includes a communication terminal, a data processing device, and an emotion engine.
[0160] Users can input questions and information about caregiving using a communication terminal. The communication terminal digitizes this input and transmits it to a data processing device. Encryption and secure communication protocols (e.g., HTTPS) are used here.
[0161] The data processing device analyzes the received data using language processing techniques (e.g., natural language processing libraries) and various data analysis techniques (speech recognition software, image recognition algorithms, etc.). If it is text, structural analysis is performed; if it is audio, it is converted to text; and if it is an image, important features are extracted to understand the context.
[0162] After data analysis, the data processing unit utilizes an emotion engine to identify the user's emotional state. The emotion engine analyzes text tone, voice intonation, and facial expressions in images to categorize the user's emotions as positive, negative, neutral, etc.
[0163] For example, if a user enters "My mother hasn't been feeling well lately, and I'm worried about spending time with her," the device sends this message to the data processing unit. The data processing unit uses an emotion engine to identify the user's anxiety and generates a reassuring response. For example, it might provide a response such as, "I recommend you wait and see for a while, and consult your doctor if necessary. We can also think of activities to cheer her up together."
[0164] An example of a prompt message for a generative AI model would be: "Identify the emotions the user is feeling and suggest care support based on those emotions."
[0165] In this way, the system can provide personalized care support tailored to the user's emotions, reduce the burden on caregivers, and improve the quality of care.
[0166] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0167] Step 1:
[0168] Users input questions and information about caregiving into a communication terminal. This input can be text, voice, or images. The input information is converted into a digital format. For example, voice input is saved as an audio file.
[0169] Step 2:
[0170] The terminal transmits digitized information to a data processing unit. The data is protected using encryption technology and transmitted via a secure protocol (e.g., HTTPS). This ensures the data is transferred safely.
[0171] Step 3:
[0172] The server analyzes the received data using language processing techniques and various data analysis techniques. For text data, it analyzes the grammatical structure and extracts meaning. Speech data is converted to text using speech recognition technology, and image data has its main features extracted using image processing technology. The output of the analysis provides unified information in the form of text data.
[0173] Step 4:
[0174] The server uses an emotion engine to identify emotional states from the analyzed data. The emotion engine classifies the user's emotions as positive, negative, or neutral based on text tone, voice intonation, and image facial expressions. This allows the user's psychological state to be identified.
[0175] Step 5:
[0176] The server generates a response to the user based on the results of the emotion engine. Response generation creates a personalized message based on the identified emotion and the user's intent. For example, it might create a message containing support and advice for a worried user. This generated response is in text format and can be converted to speech format using speech synthesis technology as needed.
[0177] Step 6:
[0178] The terminal displays the response received from the server to the user. Text responses are displayed on the screen, and audio responses are played through the speaker. This allows the user to receive support based on their own emotions.
[0179] (Application Example 2)
[0180] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0181] In modern brick-and-mortar stores, it is difficult to instantly grasp customer emotions and provide appropriate service. Furthermore, while implementing systems that can sense customer needs and emotions in real time and provide services accordingly would greatly improve a store's competitiveness, such systems are not yet sufficiently available.
[0182] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0183] In this invention, the server includes means for acquiring input information from a user via a communication device, means for transmitting the acquired input information to a processing device, and means for analyzing the received input information in the processing device using natural language processing and complex modal analysis techniques. This enables real-time sensing of the customer's emotional state in a physical store and the rapid provision of services based on that understanding.
[0184] A "communication device" is a device used to acquire information from a user and transmit it to another device.
[0185] A "user" is a person or entity that uses the system and provides information.
[0186] "Input information" refers to data transmitted by the user, including text, audio, images, and other forms.
[0187] A "processing device" is a device that analyzes received information and generates instructions based on that analysis.
[0188] "Natural language processing" is a technology that uses computers to process and understand human language.
[0189] "Complex modal analysis technology" is a technology that integrates and analyzes multiple data formats such as text, audio, and images.
[0190] "Emotional state" refers to an individual's temporary emotional state and is classified into categories such as positive, negative, and neutral.
[0191] "Service improvement information" refers to advice and instructions provided in real time based on the customer's emotional state, aimed at improving the quality of service.
[0192] The system that realizes this invention mainly uses a communication device, a processing device, and related analysis software. The communication device functions as a user device such as a smartphone or tablet and acquires input information from the user. This information is acquired as text, audio, or images, converted into digital data, and then transmitted to the processing device.
[0193] The processing unit functions as a server, analyzing the received input information. This analysis involves processing text using a natural language processing library (e.g., SpaCy) and clarifying the user's emotional state using a sentiment analysis engine (e.g., Google® Cloud Natural Language). Compound modal analysis techniques are used to comprehensively analyze audio and image information.
[0194] Based on the results of sentiment analysis, the server provides real-time advice for service improvement. This advice is transmitted to a communication device and notified to staff. This enables rapid and accurate improvement of customer service within the store.
[0195] For example, in a restaurant, if a customer says they've "waited too long," the system can sense their stress and prompt staff for a quicker response. As an example of real-time service based on customer emotional state, an example of a prompt message is as follows: "Convert the customer's statement to text, determine their emotional state, and generate a service suggestion based on that."
[0196] This invention makes it possible to provide meticulous service that takes customer emotions into consideration, thereby improving the quality of service at stores.
[0197] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0198] Step 1:
[0199] Users input either voice or text using a smartphone, which is a communication device. This input is processed in real time as digital data on the device. If the input data is text, it is acquired as text data; if it is voice, it is acquired as voice data.
[0200] Step 2:
[0201] The terminal sends the acquired input data to the server, which acts as a processing unit. In the case of audio data, it is first converted to text, and the input data is then formatted into an appropriate data format such as JSON before being transferred to the server.
[0202] Step 3:
[0203] The server analyzes the received input data using natural language processing techniques. If text is input, it analyzes its structure and extracts important keywords and phrases. This process utilizes natural language processing libraries such as SpaCy.
[0204] Step 4:
[0205] The server uses complex modal analysis techniques to gain a deep understanding of the input data. In this step, it analyzes the user's emotional state through the analysis of speech intonation and frequently occurring phrases.
[0206] Step 5:
[0207] The server uses a generative AI model to determine the user's emotions based on the results obtained from the emotion analysis engine. This process utilizes emotion analysis tools such as Google Cloud Natural Language, resulting in output classified as positive, negative, neutral, etc.
[0208] Step 6:
[0209] The server generates real-time service improvement information based on the user's emotional state. This information is generated as a concrete action plan using prompts to encourage the delivery of the best possible service to the customer.
[0210] Step 7:
[0211] The terminal notifies store staff of service improvement information received from the server. The notification is provided via voice or text to the staff's dedicated terminal to prompt the next action.
[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0215] [Second Embodiment]
[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0228] This invention is a system for processing information obtained from users and providing appropriate care and psychological support. The embodiments thereof are described below.
[0229] The system of the present invention mainly consists of a server, a terminal, and a user interface. The user creates input data in text, voice, or image format using a communication terminal. For example, the user can input text requesting advice on caregiving or upload images of a specific situation.
[0230] The terminal's role is to acquire user input data in real time and send this data to the server. The transmitted data is encrypted and managed securely.
[0231] Upon receiving input data from a terminal, the server analyzes the text using natural language processing (NLP) techniques and determines the content of images and audio using multimodal analysis techniques. Based on this, the server estimates the user's intentions and the support they require, and understands their psychological state through sentiment analysis. Through these analyses, the server generates an appropriate response to the user's request.
[0232] The generated response is sent back to the terminal, where it is displayed to the user. In some cases, the response may be converted from text to speech and provided verbally. In particular, if the server determines that the user is in an emotionally distressed situation, it will generate a response that takes emergency measures into consideration.
[0233] For example, if a user enters a message such as, "My father is depressed, and I don't know what to do," the server analyzes this message and determines that the user needs psychological support within the family. The server generates relevant support information as a response, providing the user with "specific support measures" and "referrals to professional organizations as needed." Furthermore, if the server determines that the situation is urgent, it can automatically notify relevant external support organizations to encourage prompt action.
[0234] In this way, the system of the present invention provides flexible and rapid support that responds to diverse needs, and reduces the psychological and informational burden on caregivers.
[0235] The following describes the processing flow.
[0236] Step 1:
[0237] The user inputs questions and situation descriptions in text, voice, or image format via a communication terminal. The communication terminal converts this input data into digital signals.
[0238] Step 2:
[0239] The terminal transmits the converted digital signal to the server in real time. This communication is encrypted to ensure security.
[0240] Step 3:
[0241] The server analyzes the received data. For text, natural language processing techniques are used to perform grammatical analysis and keyword extraction. Audio data is converted to text using speech recognition technology, and image data is identified using image recognition algorithms.
[0242] Step 4:
[0243] The server performs sentiment analysis to understand the user's intentions based on the analyzed data. In this process, the user's emotional state is classified as positive, negative, or neutral, and the server adjusts the response accordingly.
[0244] Step 5:
[0245] The server generates a response based on the sentiment analysis results and the user's request. This generation process references a knowledge base and constructs a response that includes the latest care and psychological support information.
[0246] Step 6:
[0247] The server sends the generated response to the terminal. The terminal displays the received text response to the user and, if necessary, also provides a voice response using speech synthesis technology.
[0248] Step 7:
[0249] If the server determines that a user is in an emergency, it automatically notifies registered external support organizations. This notification includes concise and important information, enabling a swift response.
[0250] Step 8:
[0251] The terminal notifies the user of responses from the server and confirmation of emergency calls initiated by the server, preparing the user for their next action.
[0252] (Example 1)
[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0254] Through the input and analysis of information, there is a challenge in providing appropriate and prompt responses when users require care or psychological support. Furthermore, providing individually customized support tailored to the emotional state of each user is also a crucial challenge.
[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0256] In this invention, the server includes means for collecting data from users via information terminals, means for processing the data using natural language processing and various forms of analysis techniques in a data processing device, and means for evaluating the user's purpose and mental state based on the processing results. This enables the rapid provision of customized support tailored to the user's needs and allows for coordinated responses with external support organizations in emergencies.
[0257] An "information terminal" is a general term for electronic devices used by users to input and display data.
[0258] "User" refers to the entity that uses the system to input data and receives analysis results and responses.
[0259] A "data processing device" refers to a device that receives data transmitted from an information terminal and performs analysis and processing on it.
[0260] "Natural language processing" refers to the technology used by data processing devices to analyze human language and extract its meaning.
[0261] "Analysis techniques for diverse formats" refers to technologies that effectively analyze data in different formats, such as text, audio, and images.
[0262] "Evaluation" refers to the act of judging a user's purpose and mental state based on the analysis results.
[0263] "Response" refers to the response provided to the user based on the analysis results and evaluation.
[0264] "Speech synthesis technology" refers to the technology that converts text data into speech and outputs it as speech.
[0265] An "external support organization" refers to an external, specialized support organization that can integrate with the system in an emergency.
[0266] This invention is a system for providing flexible and rapid care and psychological support to users. It mainly includes information terminals, servers, and user-involved processes.
[0267] Users input information requiring support in text, voice, or image format using an information terminal. For example, they can input a question such as, "My mother hasn't been feeling well lately, and I'm worried," or images related to a specific situation. The information terminal collects this input data and securely transmits it to a data processing unit using encryption technology.
[0268] The server functions as a data processing unit, processing received data using specific analysis techniques. It analyzes text data using natural language processing (NLP) techniques and evaluates audio and image data in detail using multimodal analysis techniques.
[0269] The server utilizes a generative AI model to analyze the results and build support content optimized for the user. For example, if a user enters the sentence, "My father is depressed, and I don't know what to do," the server analyzes this and determines that psychological support is needed. It then provides appropriate support measures and information on partner organizations.
[0270] An example of a prompt message a user might use for input is, "My father is depressed, and I don't know what to do." Such input allows the system to respond to the user's individual needs. Furthermore, in emergencies, the server can assess the user's situation and automatically notify external support organizations as needed. This allows users to utilize the system with peace of mind.
[0271] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0272] Step 1:
[0273] Users create input data using an information terminal. For example, they might enter text describing the situation in which they need support, or related images. The data entered by the user is the first step toward providing accurate information services and is collected by the terminal after input.
[0274] Step 2:
[0275] The device acquires data entered by the user in real time. This data is encrypted to ensure security. The device then transmits the collected data to a server via a secure protocol. Input data can be in text, audio, or image format, while output is encrypted communication data.
[0276] Step 3:
[0277] The server receives encrypted data sent from the terminal and first decrypts it. Then, the server uses natural language processing (NLP) techniques to analyze the text data and process it to understand the type of support the user is seeking. It also performs detailed analysis of audio and image data using multimodal analysis. The input is encrypted data, and the output is information corresponding to the user's needs after the analysis.
[0278] Step 4:
[0279] The server determines the user's intentions and emotions based on the analyzed data. Using a generative AI model, it generates appropriate responses tailored to the user. At this stage, the server constructs specific support measures or information from the analysis results. The input is the analyzed user information, and the output is the generated response data.
[0280] Step 5:
[0281] The response data generated by the server is sent to the terminal. The terminal displays this response to the user. If voice output is required, the terminal also responds to the user verbally via text-to-speech synthesis. The input is the response data from the server, and the output is the displayed or audio information delivered to the user.
[0282] (Application Example 1)
[0283] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0284] There is a demand to grasp changes in the psychological and health states of humans in real time and provide appropriate support. However, in conventional systems, it has been difficult to immediately detect changes in the user's state and provide specific support according to the living environment.
[0285] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following respective means.
[0286] In this invention, the server includes means for acquiring input information from a human via a communication device, means for estimating the intention and psychological state of the human based on the analysis result, and means for generating an appropriate response based on the analyzed information. Thereby, it becomes possible to monitor changes in the user's living environment and mood in real time and specifically provide appropriate support.
[0287] The "communication device" is a device for acquiring input information from a human and transmitting it to a central processing unit.
[0288] The "human" refers to an end user who uses the system and is the target for analyzing the intention and psychological state through their input.
[0289] The "input information" is data provided by a human in the form of text, voice, image, etc.
[0290] The "central processing unit" is a computer system for performing overall language processing and various analyses based on the acquired input information and generating an analysis result.
[0291] The "overall language processing" is a technology for analyzing text data input using natural language processing technology and understanding its meaning.
[0292] The "various analysis technologies" are technologies for analyzing image and voice data to identify and judge their contents.
[0293] "Intention" refers to the purpose or request that a person tries to convey through a system.
[0294] "Psychological state" refers to a person's emotions and mental state, and by estimating this, it becomes possible to provide appropriate support.
[0295] "Response" refers to support information and advice generated by the central processing unit based on the analysis results.
[0296] A "display device" is a device that monitors a person's living environment and circumstances, and displays a response visually or audibly as needed.
[0297] A "support organization" is an external professional organization or institution that is referred to when needed and provides support to individuals.
[0298] The system for realizing this invention consists of a communication device, a central processing unit, a display device, and cooperation with a support organization. First, the communication device acquires input information from humans, i.e., text, voice, and images, and transmits it to the central processing unit. This information is encrypted and securely sent to the central processing unit.
[0299] The central processing unit analyzes the acquired input information using holistic language processing techniques (such as TensorFlow and spaCy) and diverse analysis techniques (such as OpenCV and Librosa). This allows for accurate estimation of human intentions and psychological states. For example, it can estimate emotional changes from voice input or determine everyday behavior through image analysis.
[0300] Based on the analysis results, the central processing unit generates the optimal response. This response is presented to a human using speech synthesis technology as needed. In emergencies, it can also automatically notify the support organizations that require assistance.
[0301] As a specific example, when a user inputs "I've been feeling unwell recently and having difficulty with my daily tasks", the central processing unit can analyze this data and provide recommendations for relaxation or introductions to medical institutions. This enables the user to receive timely and accurate support.
[0302] As an example of a prompt sentence, it is in the form of "Generate a program that proposes relaxation methods when the user appears tired." With such a prompt, the generative AI model can formulate a processing procedure for indicating an appropriate intervention method.
[0303] The flow of the specific processing in Application Example 1 will be described using FIG. 12.
[0304] Step 1:
[0305] The user uses a communication device to provide input information. This information is in text, voice, or image format, and different initial processing is performed depending on each format. The user's input is temporarily stored by the communication device. Subsequently, this input information is transmitted to the central processing unit in a secure form using encryption technology.
[0306] Step 2:
[0307] The server acquires the input information received from the communication device and performs initial analysis according to the data format. In the case of voice data, it is converted into text using voice recognition technology, and for image data, main features are extracted using image processing technology. Thereby, the server generates a standardized data format for subsequent processing.
[0308] Step 3:
[0309] The server uses a general language processing engine to analyze the meaning of the text data. In this process, keywords and context are extracted from the input information to grasp the user's intention. For example, if keywords such as "tired" are included, based on that, it proceeds to the estimation of related mental states.
[0310] Step 4:
[0311] The server uses diverse analysis techniques to detect the user's daily activities and changes in facial expressions from image data. This generates complementary data that allows for more accurate emotion estimation based on the information obtained from the images. This data makes it possible to visually understand the user's situation.
[0312] Step 5:
[0313] Based on the analysis results, the server applies a generative AI model to generate the optimal response. This model performs inference according to the prompt text and designs a response that is appropriate to the user's state. The response is generated in text format and then converted to speech using speech synthesis technology as needed.
[0314] Step 6:
[0315] The server sends the generated response to the communication device. The communication device displays the response as text or plays it back as audio. Specifically, it provides suggestions for relaxation methods and contact information for support organizations as needed. As a result, the user is encouraged to take appropriate action and is in a position to receive support.
[0316] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0317] This invention is a novel care support system that incorporates an emotion engine to recognize the user's emotions and provide highly personalized support based on those emotions. The system includes a communication terminal, a server, and an emotion engine.
[0318] Users can input questions and information about caregiving using a communication terminal. The input data is appropriately digitized on the terminal and sent to the server. Upon receiving the input data, the server first analyzes it using natural language processing and multimodal analysis techniques. This analysis involves sentence structure analysis if the data is text, speech recognition if it is audio, and image recognition if it is images to understand the content.
[0319] After this data analysis, the server uses an emotion engine to identify emotions from the user's input. The emotion engine analyzes, for example, the tone of text, the intonation of speech, and facial expressions in images, and classifies them into emotion categories such as positive, negative, and neutral. Based on the results of the emotion engine, the user's psychological state is accurately understood.
[0320] The server then synthesizes these analysis results and proceeds to a process of generating a response that aligns with the user's intentions. This response generation process includes considering the emotions identified by the emotion engine and creating a detailed response that matches the support the user desires and their psychological state. The generated response is sent to the terminal as text and, in some cases, presented as speech using speech synthesis technology.
[0321] For example, if a user enters "My mother hasn't been feeling well lately, and I'm worried about spending time with her," the server analyzes this message and uses an emotion engine to read the user's anxiety. The server then generates emotional support and practical advice that corresponds to this emotion, providing a response such as, "I recommend you wait and see for a while, and consult your doctor if necessary. We can also think of activities to cheer her up together."
[0322] This system allows users to receive care support that is tailored to their emotional needs, thereby reducing the burden of daily caregiving.
[0323] The following describes the processing flow.
[0324] Step 1:
[0325] Users input questions or concerns in text, voice, or image format using a communication terminal. The communication terminal digitizes the input data and prepares it to be immediately transmitted to the server.
[0326] Step 2:
[0327] The terminal transmits digitized input data to the server. The communication is encrypted to ensure the security of user data.
[0328] Step 3:
[0329] The server analyzes the received data. For text, it uses natural language processing for syntactic analysis and keyword extraction; for speech, it uses speech recognition to convert it to text; and for images, it uses image recognition technology to determine visual elements.
[0330] Step 4:
[0331] The server passes the analyzed data through the emotion engine. The emotion engine analyzes the emotions contained in the user's input and assigns one of three emotion labels: positive, negative, or neutral.
[0332] Step 5:
[0333] The server comprehensively understands the user's intent and emotions based on the results of natural language processing and the emotion engine. It then moves on to a process of generating the most appropriate response for the user based on this information.
[0334] Step 6:
[0335] The server uses a generative AI model to create responses that align with the user's intentions. These responses include emotional support and practical advice based on the emotions indicated by the emotion engine.
[0336] Step 7:
[0337] The server sends the completed response to the terminal. The terminal displays the received response to the user. If necessary, the response can also be provided in voice using speech synthesis technology.
[0338] Step 8:
[0339] Under the monitoring of the emotion engine, if the server determines that user input is urgent, it will automatically notify pre-configured external support organizations and initiate procedures to request prompt assistance.
[0340] (Example 2)
[0341] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0342] In modern society, caregivers are required to respond quickly and appropriately to increasingly complex needs. However, caregivers often find it difficult to understand the underlying emotions and psychological states of users, limiting their ability to provide appropriate support. This problem highlights the need to develop systems that enable even non-professionals to provide appropriate advice and support tailored to the situation.
[0343] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0344] In this invention, the server includes means for analyzing information from the user using language processing technology and various data analysis technologies, means for identifying the user's abstract intentions and emotional state based on the analysis results, and means for creating an individualized response based on the identified information. This makes it possible to quickly provide individualized support that is in line with the user's emotions and situation, reduce the burden on caregivers, and improve the quality of care.
[0345] A "communication device" is a device that acquires information from a user and transmits it to a data processing device.
[0346] A "data processing device" is a device that analyzes received information to identify the user's intentions and emotional state.
[0347] "Language processing technology" refers to techniques for analyzing text data and understanding its structure and meaning.
[0348] "Diverse data analysis technologies" refer to technologies for analyzing data in various formats, such as text, audio, and images, and understanding their content.
[0349] A "personalized response" is a response that is generated in a way that is tailored to the user's specific situation and emotions.
[0350] "Abstract user intent" refers to the goals or needs that users have implicitly but do not explicitly express.
[0351] "Emotional state" refers to the psychological state a user exhibits, such as positive, negative, or neutral emotions.
[0352] One embodiment of the present invention is configured as a care support system incorporating emotion recognition technology. This system includes a communication terminal, a data processing device, and an emotion engine.
[0353] Users can input questions and information about caregiving using a communication terminal. The communication terminal digitizes this input and transmits it to a data processing device. Encryption and secure communication protocols (e.g., HTTPS) are used here.
[0354] The data processing device analyzes the received data using language processing techniques (e.g., natural language processing libraries) and various data analysis techniques (speech recognition software, image recognition algorithms, etc.). If it is text, structural analysis is performed; if it is audio, it is converted to text; and if it is an image, important features are extracted to understand the context.
[0355] After data analysis, the data processing unit utilizes an emotion engine to identify the user's emotional state. The emotion engine analyzes text tone, voice intonation, and facial expressions in images to categorize the user's emotions as positive, negative, neutral, etc.
[0356] For example, if a user enters "My mother hasn't been feeling well lately, and I'm worried about spending time with her," the device sends this message to the data processing unit. The data processing unit uses an emotion engine to identify the user's anxiety and generates a reassuring response. For example, it might provide a response such as, "I recommend you wait and see for a while, and consult your doctor if necessary. We can also think of activities to cheer her up together."
[0357] An example of a prompt message for a generative AI model would be: "Identify the emotions the user is feeling and suggest care support based on those emotions."
[0358] In this way, the system can provide personalized care support tailored to the user's emotions, reduce the burden on caregivers, and improve the quality of care.
[0359] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0360] Step 1:
[0361] Users input questions and information about caregiving into a communication terminal. This input can be text, voice, or images. The input information is converted into a digital format. For example, voice input is saved as an audio file.
[0362] Step 2:
[0363] The terminal transmits digitized information to a data processing unit. The data is protected using encryption technology and transmitted via a secure protocol (e.g., HTTPS). This ensures the data is transferred safely.
[0364] Step 3:
[0365] The server analyzes the received data using language processing techniques and various data analysis techniques. For text data, it analyzes the grammatical structure and extracts meaning. Speech data is converted to text using speech recognition technology, and image data has its main features extracted using image processing technology. The output of the analysis provides unified information in the form of text data.
[0366] Step 4:
[0367] The server uses an emotion engine to identify emotional states from the analyzed data. The emotion engine classifies the user's emotions as positive, negative, or neutral based on text tone, voice intonation, and image facial expressions. This allows the user's psychological state to be identified.
[0368] Step 5:
[0369] The server generates a response to the user based on the results of the emotion engine. Response generation creates a personalized message based on the identified emotion and the user's intent. For example, it might create a message containing support and advice for a worried user. This generated response is in text format and can be converted to speech format using speech synthesis technology as needed.
[0370] Step 6:
[0371] The terminal displays the response received from the server to the user. Text responses are displayed on the screen, and audio responses are played through the speaker. This allows the user to receive support based on their own emotions.
[0372] (Application Example 2)
[0373] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0374] In modern brick-and-mortar stores, it is difficult to instantly grasp customer emotions and provide appropriate service. Furthermore, while implementing systems that can sense customer needs and emotions in real time and provide services accordingly would greatly improve a store's competitiveness, such systems are not yet sufficiently available.
[0375] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0376] In this invention, the server includes means for acquiring input information from a user via a communication device, means for transmitting the acquired input information to a processing device, and means for analyzing the received input information in the processing device using natural language processing and complex modal analysis techniques. This enables real-time sensing of the customer's emotional state in a physical store and the rapid provision of services based on that understanding.
[0377] A "communication device" is a device used to acquire information from a user and transmit it to another device.
[0378] A "user" is a person or entity that uses the system and provides information.
[0379] "Input information" refers to data transmitted by the user, including text, audio, images, and other forms.
[0380] A "processing device" is a device that analyzes received information and generates instructions based on that analysis.
[0381] "Natural language processing" is a technology that uses computers to process and understand human language.
[0382] "Complex modal analysis technology" is a technology that integrates and analyzes multiple data formats such as text, audio, and images.
[0383] "Emotional state" refers to an individual's temporary emotional state and is classified into categories such as positive, negative, and neutral.
[0384] "Service improvement information" refers to advice and instructions provided in real time based on the customer's emotional state, aimed at improving the quality of service.
[0385] The system that realizes this invention mainly uses a communication device, a processing device, and related analysis software. The communication device functions as a user device such as a smartphone or tablet and acquires input information from the user. This information is acquired as text, audio, or images, converted into digital data, and then transmitted to the processing device.
[0386] The processing unit functions as a server, analyzing the received input information. This analysis involves processing text using a natural language processing library (e.g., SpaCy) and clarifying the user's emotional state using a sentiment analysis engine (e.g., Google Cloud Natural Language). Compound modal analysis techniques are used to comprehensively analyze audio and image information.
[0387] Based on the results of sentiment analysis, the server provides real-time advice for service improvement. This advice is transmitted to a communication device and notified to staff. This enables rapid and accurate improvement of customer service within the store.
[0388] For example, in a restaurant, if a customer says they've "waited too long," the system can sense their stress and prompt staff for a quicker response. As an example of real-time service based on customer emotional state, an example of a prompt message is as follows: "Convert the customer's statement to text, determine their emotional state, and generate a service suggestion based on that."
[0389] This invention makes it possible to provide meticulous service that takes customer emotions into consideration, thereby improving the quality of service at stores.
[0390] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0391] Step 1:
[0392] Users input either voice or text using a smartphone, which is a communication device. This input is processed in real time as digital data on the device. If the input data is text, it is acquired as text data; if it is voice, it is acquired as voice data.
[0393] Step 2:
[0394] The terminal sends the acquired input data to the server, which acts as a processing unit. In the case of audio data, it is first converted to text, and the input data is then formatted into an appropriate data format such as JSON before being transferred to the server.
[0395] Step 3:
[0396] The server analyzes the received input data using natural language processing techniques. If text is input, it analyzes its structure and extracts important keywords and phrases. This process utilizes natural language processing libraries such as SpaCy.
[0397] Step 4:
[0398] The server uses complex modal analysis techniques to gain a deep understanding of the input data. In this step, it analyzes the user's emotional state through the analysis of speech intonation and frequently occurring phrases.
[0399] Step 5:
[0400] The server uses a generative AI model to determine the user's emotions based on the results obtained from the emotion analysis engine. This process utilizes emotion analysis tools such as Google Cloud Natural Language, resulting in output classified as positive, negative, neutral, etc.
[0401] Step 6:
[0402] The server generates real-time service improvement information based on the user's emotional state. This information is generated as a concrete action plan using prompts to encourage the delivery of the best possible service to the customer.
[0403] Step 7:
[0404] The terminal notifies store staff of service improvement information received from the server. The notification is provided via voice or text to the staff's dedicated terminal to prompt the next action.
[0405] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0406] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0408] [Third Embodiment]
[0409] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0410] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0412] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0416] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0417] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0418] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0419] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0420] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0421] This invention is a system for processing information obtained from users and providing appropriate care and psychological support. The embodiments thereof are described below.
[0422] The system of the present invention mainly consists of a server, a terminal, and a user interface. The user creates input data in text, voice, or image format using a communication terminal. For example, the user can input text requesting advice on caregiving or upload images of a specific situation.
[0423] The terminal's role is to acquire user input data in real time and send this data to the server. The transmitted data is encrypted and managed securely.
[0424] Upon receiving input data from a terminal, the server analyzes the text using natural language processing (NLP) techniques and determines the content of images and audio using multimodal analysis techniques. Based on this, the server estimates the user's intentions and the support they require, and understands their psychological state through sentiment analysis. Through these analyses, the server generates an appropriate response to the user's request.
[0425] The generated response is sent back to the terminal, where it is displayed to the user. In some cases, the response may be converted from text to speech and provided verbally. In particular, if the server determines that the user is in an emotionally distressed situation, it will generate a response that takes emergency measures into consideration.
[0426] For example, if a user enters a message such as, "My father is depressed, and I don't know what to do," the server analyzes this message and determines that the user needs psychological support within the family. The server generates relevant support information as a response, providing the user with "specific support measures" and "referrals to professional organizations as needed." Furthermore, if the server determines that the situation is urgent, it can automatically notify relevant external support organizations to encourage prompt action.
[0427] In this way, the system of the present invention provides flexible and rapid support that responds to diverse needs, and reduces the psychological and informational burden on caregivers.
[0428] The following describes the processing flow.
[0429] Step 1:
[0430] The user inputs questions and situation descriptions in text, voice, or image format via a communication terminal. The communication terminal converts this input data into digital signals.
[0431] Step 2:
[0432] The terminal transmits the converted digital signal to the server in real time. This communication is encrypted to ensure security.
[0433] Step 3:
[0434] The server analyzes the received data. For text, natural language processing techniques are used to perform grammatical analysis and keyword extraction. Audio data is converted to text using speech recognition technology, and image data is identified using image recognition algorithms.
[0435] Step 4:
[0436] The server performs sentiment analysis to understand the user's intentions based on the analyzed data. In this process, the user's emotional state is classified as positive, negative, or neutral, and the server adjusts the response accordingly.
[0437] Step 5:
[0438] The server generates a response based on the sentiment analysis results and the user's request. This generation process references a knowledge base and constructs a response that includes the latest care and psychological support information.
[0439] Step 6:
[0440] The server sends the generated response to the terminal. The terminal displays the received text response to the user and, if necessary, also provides a voice response using speech synthesis technology.
[0441] Step 7:
[0442] If the server determines that a user is in an emergency, it automatically notifies registered external support organizations. This notification includes concise and important information, enabling a swift response.
[0443] Step 8:
[0444] The terminal notifies the user of responses from the server and confirmation of emergency calls initiated by the server, preparing the user for their next action.
[0445] (Example 1)
[0446] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0447] Through the input and analysis of information, there is a challenge in providing appropriate and prompt responses when users require care or psychological support. Furthermore, providing individually customized support tailored to the emotional state of each user is also a crucial challenge.
[0448] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0449] In this invention, the server includes means for collecting data from users via information terminals, means for processing the data using natural language processing and various forms of analysis techniques in a data processing device, and means for evaluating the user's purpose and mental state based on the processing results. This enables the rapid provision of customized support tailored to the user's needs and allows for coordinated responses with external support organizations in emergencies.
[0450] An "information terminal" is a general term for electronic devices used by users to input and display data.
[0451] "User" refers to the entity that uses the system to input data and receives analysis results and responses.
[0452] A "data processing device" refers to a device that receives data transmitted from an information terminal and performs analysis and processing on it.
[0453] "Natural language processing" refers to the technology used by data processing devices to analyze human language and extract its meaning.
[0454] "Analysis techniques for diverse formats" refers to technologies that effectively analyze data in different formats, such as text, audio, and images.
[0455] "Evaluation" refers to the act of judging a user's purpose and mental state based on the analysis results.
[0456] "Response" refers to the response provided to the user based on the analysis results and evaluation.
[0457] "Speech synthesis technology" refers to the technology that converts text data into speech and outputs it as speech.
[0458] An "external support organization" refers to an external, specialized support organization that can integrate with the system in an emergency.
[0459] This invention is a system for providing flexible and rapid care and psychological support to users. It mainly includes information terminals, servers, and user-involved processes.
[0460] Users input information requiring support in text, voice, or image format using an information terminal. For example, they can input a question such as, "My mother hasn't been feeling well lately, and I'm worried," or images related to a specific situation. The information terminal collects this input data and securely transmits it to a data processing unit using encryption technology.
[0461] The server functions as a data processing unit, processing received data using specific analysis techniques. It analyzes text data using natural language processing (NLP) techniques and evaluates audio and image data in detail using multimodal analysis techniques.
[0462] The server utilizes a generative AI model to analyze the results and build support content optimized for the user. For example, if a user enters the sentence, "My father is depressed, and I don't know what to do," the server analyzes this and determines that psychological support is needed. It then provides appropriate support measures and information on partner organizations.
[0463] An example of a prompt message a user might use for input is, "My father is depressed, and I don't know what to do." Such input allows the system to respond to the user's individual needs. Furthermore, in emergencies, the server can assess the user's situation and automatically notify external support organizations as needed. This allows users to utilize the system with peace of mind.
[0464] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0465] Step 1:
[0466] Users create input data using an information terminal. For example, they might enter text describing the situation in which they need support, or related images. The data entered by the user is the first step toward providing accurate information services and is collected by the terminal after input.
[0467] Step 2:
[0468] The device acquires data entered by the user in real time. This data is encrypted to ensure security. The device then transmits the collected data to a server via a secure protocol. Input data can be in text, audio, or image format, while output is encrypted communication data.
[0469] Step 3:
[0470] The server receives encrypted data sent from the terminal and first decrypts it. Then, the server uses natural language processing (NLP) techniques to analyze the text data and process it to understand the type of support the user is seeking. It also performs detailed analysis of audio and image data using multimodal analysis. The input is encrypted data, and the output is information corresponding to the user's needs after the analysis.
[0471] Step 4:
[0472] The server determines the user's intentions and emotions based on the analyzed data. Using a generative AI model, it generates appropriate responses tailored to the user. At this stage, the server constructs specific support measures or information from the analysis results. The input is the analyzed user information, and the output is the generated response data.
[0473] Step 5:
[0474] The response data generated by the server is sent to the terminal. The terminal displays this response to the user. If voice output is required, the terminal also responds to the user verbally via text-to-speech synthesis. The input is the response data from the server, and the output is the displayed or audio information delivered to the user.
[0475] (Application Example 1)
[0476] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0477] There is a need to understand changes in a person's psychological and physical state in real time and provide appropriate support. However, conventional systems have had difficulty instantly detecting changes in a user's condition and providing specific support tailored to their living environment.
[0478] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0479] In this invention, the server includes means for acquiring input information from a human via a communication device, means for estimating the human's intentions and psychological state based on the analysis results, and means for generating an appropriate response based on the analyzed information. This makes it possible to monitor changes in the user's living environment and mood in real time and provide specific and appropriate support.
[0480] A "communication device" is a device that acquires input information from a human and transmits it to a central processing unit.
[0481] "Human" refers to the end-users who use the system, and their intentions and psychological states are analyzed through their input.
[0482] "Input information" refers to data provided by humans in various forms, such as text, audio, and images.
[0483] A "central processing unit" is a computer system that performs overall language processing and diverse analysis based on acquired input information and generates analysis results.
[0484] "Holistic language processing" is a technique that uses natural language processing technology to analyze input text data and understand its meaning.
[0485] "Diverse analysis technology" is a technology that analyzes image and audio data to identify and interpret its content.
[0486] "Intention" refers to the purpose or request that a person tries to convey through a system.
[0487] "Psychological state" refers to a person's emotions and mental state, and by estimating this, it becomes possible to provide appropriate support.
[0488] "Response" refers to support information and advice generated by the central processing unit based on the analysis results.
[0489] A "display device" is a device that monitors a person's living environment and circumstances, and displays a response visually or audibly as needed.
[0490] A "support organization" is an external professional organization or institution that is referred to when needed and provides support to individuals.
[0491] The system for realizing this invention consists of a communication device, a central processing unit, a display device, and cooperation with a support organization. First, the communication device acquires input information from humans, i.e., text, voice, and images, and transmits it to the central processing unit. This information is encrypted and securely sent to the central processing unit.
[0492] The central processing unit analyzes the acquired input information using holistic language processing techniques (such as TensorFlow and spaCy) and diverse analysis techniques (such as OpenCV and Librosa). This allows for accurate estimation of human intentions and psychological states. For example, it can estimate emotional changes from voice input or determine everyday behavior through image analysis.
[0493] Based on the analysis results, the central processing unit generates the optimal response. This response is presented to a human using speech synthesis technology as needed. In emergencies, it can also automatically notify the support organizations that require assistance.
[0494] For example, if a user inputs "I've been feeling unwell lately and having difficulty with daily tasks," the central processing unit can analyze this data and provide suggestions for relaxation or referrals to medical institutions. This allows users to receive timely and appropriate support.
[0495] An example of a prompt might be, "Generate a program that suggests relaxation methods when the user appears tired." Such prompts allow the generating AI model to formulate processing steps to indicate appropriate intervention methods.
[0496] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0497] Step 1:
[0498] The user provides input information using a communication device. This information can be in text, audio, or image format, and different initial processing is performed depending on the format. The user's input is temporarily stored by the communication device. Subsequently, this input information is transmitted securely to the central processing unit using encryption technology.
[0499] Step 2:
[0500] The server acquires input information received from the communication device and performs initial analysis according to the data format. In the case of audio data, it is converted to text using speech recognition technology, and for image data, key features are extracted using image processing technology. This allows the server to generate a standardized data format for subsequent processing.
[0501] Step 3:
[0502] The server uses a global language processing engine to analyze the meaning of text data. This process extracts keywords and context from the input information to understand the user's intent. For example, if keywords such as "tired" are included, the server proceeds to estimate the related psychological state based on these keywords.
[0503] Step 4:
[0504] The server uses diverse analysis techniques to detect the user's daily activities and changes in facial expressions from image data. This generates complementary data that allows for more accurate emotion estimation based on the information obtained from the images. This data makes it possible to visually understand the user's situation.
[0505] Step 5:
[0506] Based on the analysis results, the server applies a generative AI model to generate the optimal response. This model performs inference according to the prompt text and designs a response that is appropriate to the user's state. The response is generated in text format and then converted to speech using speech synthesis technology as needed.
[0507] Step 6:
[0508] The server sends the generated response to the communication device. The communication device displays the response as text or plays it back as audio. Specifically, it provides suggestions for relaxation methods and contact information for support organizations as needed. As a result, the user is encouraged to take appropriate action and is in a position to receive support.
[0509] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0510] This invention is a novel care support system that incorporates an emotion engine to recognize the user's emotions and provide highly personalized support based on those emotions. The system includes a communication terminal, a server, and an emotion engine.
[0511] Users can input questions and information about caregiving using a communication terminal. The input data is appropriately digitized on the terminal and sent to the server. Upon receiving the input data, the server first analyzes it using natural language processing and multimodal analysis techniques. This analysis involves sentence structure analysis if the data is text, speech recognition if it is audio, and image recognition if it is images to understand the content.
[0512] After this data analysis, the server uses an emotion engine to identify emotions from the user's input. The emotion engine analyzes, for example, the tone of text, the intonation of speech, and facial expressions in images, and classifies them into emotion categories such as positive, negative, and neutral. Based on the results of the emotion engine, the user's psychological state is accurately understood.
[0513] The server then synthesizes these analysis results and proceeds to a process of generating a response that aligns with the user's intentions. This response generation process includes considering the emotions identified by the emotion engine and creating a detailed response that matches the support the user desires and their psychological state. The generated response is sent to the terminal as text and, in some cases, presented as speech using speech synthesis technology.
[0514] For example, if a user enters "My mother hasn't been feeling well lately, and I'm worried about spending time with her," the server analyzes this message and uses an emotion engine to read the user's anxiety. The server then generates emotional support and practical advice that corresponds to this emotion, providing a response such as, "I recommend you wait and see for a while, and consult your doctor if necessary. We can also think of activities to cheer her up together."
[0515] This system allows users to receive care support that is tailored to their emotional needs, thereby reducing the burden of daily caregiving.
[0516] The following describes the processing flow.
[0517] Step 1:
[0518] Users input questions or concerns in text, voice, or image format using a communication terminal. The communication terminal digitizes the input data and prepares it to be immediately transmitted to the server.
[0519] Step 2:
[0520] The terminal transmits digitized input data to the server. The communication is encrypted to ensure the security of user data.
[0521] Step 3:
[0522] The server analyzes the received data. For text, it uses natural language processing for syntactic analysis and keyword extraction; for speech, it uses speech recognition to convert it to text; and for images, it uses image recognition technology to determine visual elements.
[0523] Step 4:
[0524] The server passes the analyzed data through the emotion engine. The emotion engine analyzes the emotions contained in the user's input and assigns one of three emotion labels: positive, negative, or neutral.
[0525] Step 5:
[0526] The server comprehensively understands the user's intent and emotions based on the results of natural language processing and the emotion engine. It then moves on to a process of generating the most appropriate response for the user based on this information.
[0527] Step 6:
[0528] The server uses a generative AI model to create responses that align with the user's intentions. These responses include emotional support and practical advice based on the emotions indicated by the emotion engine.
[0529] Step 7:
[0530] The server sends the completed response to the terminal. The terminal displays the received response to the user. If necessary, the response can also be provided in voice using speech synthesis technology.
[0531] Step 8:
[0532] Under the monitoring of the emotion engine, if the server determines that user input is urgent, it will automatically notify pre-configured external support organizations and initiate procedures to request prompt assistance.
[0533] (Example 2)
[0534] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0535] In modern society, caregivers are required to respond quickly and appropriately to increasingly complex needs. However, caregivers often find it difficult to understand the underlying emotions and psychological states of users, limiting their ability to provide appropriate support. This problem highlights the need to develop systems that enable even non-professionals to provide appropriate advice and support tailored to the situation.
[0536] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0537] In this invention, the server includes means for analyzing information from the user using language processing technology and various data analysis technologies, means for identifying the user's abstract intentions and emotional state based on the analysis results, and means for creating an individualized response based on the identified information. This makes it possible to quickly provide individualized support that is in line with the user's emotions and situation, reduce the burden on caregivers, and improve the quality of care.
[0538] A "communication device" is a device that acquires information from a user and transmits it to a data processing device.
[0539] A "data processing device" is a device that analyzes received information to identify the user's intentions and emotional state.
[0540] "Language processing technology" refers to techniques for analyzing text data and understanding its structure and meaning.
[0541] "Diverse data analysis technologies" refer to technologies for analyzing data in various formats, such as text, audio, and images, and understanding their content.
[0542] A "personalized response" is a response that is generated in a way that is tailored to the user's specific situation and emotions.
[0543] "Abstract user intent" refers to the goals or needs that users have implicitly but do not explicitly express.
[0544] "Emotional state" refers to the psychological state a user exhibits, such as positive, negative, or neutral emotions.
[0545] One embodiment of the present invention is configured as a care support system incorporating emotion recognition technology. This system includes a communication terminal, a data processing device, and an emotion engine.
[0546] Users can input questions and information about caregiving using a communication terminal. The communication terminal digitizes this input and transmits it to a data processing device. Encryption and secure communication protocols (e.g., HTTPS) are used here.
[0547] The data processing device analyzes the received data using language processing techniques (e.g., natural language processing libraries) and various data analysis techniques (speech recognition software, image recognition algorithms, etc.). If it is text, structural analysis is performed; if it is audio, it is converted to text; and if it is an image, important features are extracted to understand the context.
[0548] After data analysis, the data processing unit utilizes an emotion engine to identify the user's emotional state. The emotion engine analyzes text tone, voice intonation, and facial expressions in images to categorize the user's emotions as positive, negative, neutral, etc.
[0549] For example, if a user enters "My mother hasn't been feeling well lately, and I'm worried about spending time with her," the device sends this message to the data processing unit. The data processing unit uses an emotion engine to identify the user's anxiety and generates a reassuring response. For example, it might provide a response such as, "I recommend you wait and see for a while, and consult your doctor if necessary. We can also think of activities to cheer her up together."
[0550] An example of a prompt message for a generative AI model would be: "Identify the emotions the user is feeling and suggest care support based on those emotions."
[0551] In this way, the system can provide personalized care support tailored to the user's emotions, reduce the burden on caregivers, and improve the quality of care.
[0552] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0553] Step 1:
[0554] Users input questions and information about caregiving into a communication terminal. This input can be text, voice, or images. The input information is converted into a digital format. For example, voice input is saved as an audio file.
[0555] Step 2:
[0556] The terminal transmits digitized information to a data processing unit. The data is protected using encryption technology and transmitted via a secure protocol (e.g., HTTPS). This ensures the data is transferred safely.
[0557] Step 3:
[0558] The server analyzes the received data using language processing techniques and various data analysis techniques. For text data, it analyzes the grammatical structure and extracts meaning. Speech data is converted to text using speech recognition technology, and image data has its main features extracted using image processing technology. The output of the analysis provides unified information in the form of text data.
[0559] Step 4:
[0560] The server uses an emotion engine to identify emotional states from the analyzed data. The emotion engine classifies the user's emotions as positive, negative, or neutral based on text tone, voice intonation, and image facial expressions. This allows the user's psychological state to be identified.
[0561] Step 5:
[0562] The server generates a response to the user based on the results of the emotion engine. Response generation creates a personalized message based on the identified emotion and the user's intent. For example, it might create a message containing support and advice for a worried user. This generated response is in text format and can be converted to speech format using speech synthesis technology as needed.
[0563] Step 6:
[0564] The terminal displays the response received from the server to the user. Text responses are displayed on the screen, and audio responses are played through the speaker. This allows the user to receive support based on their own emotions.
[0565] (Application Example 2)
[0566] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0567] In modern brick-and-mortar stores, it is difficult to instantly grasp customer emotions and provide appropriate service. Furthermore, while implementing systems that can sense customer needs and emotions in real time and provide services accordingly would greatly improve a store's competitiveness, such systems are not yet sufficiently available.
[0568] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0569] In this invention, the server includes means for acquiring input information from a user via a communication device, means for transmitting the acquired input information to a processing device, and means for analyzing the received input information in the processing device using natural language processing and complex modal analysis techniques. This enables real-time sensing of the customer's emotional state in a physical store and the rapid provision of services based on that understanding.
[0570] A "communication device" is a device used to acquire information from a user and transmit it to another device.
[0571] A "user" is a person or entity that uses the system and provides information.
[0572] "Input information" refers to data transmitted by the user, including text, audio, images, and other forms.
[0573] A "processing device" is a device that analyzes received information and generates instructions based on that analysis.
[0574] "Natural language processing" is a technology that uses computers to process and understand human language.
[0575] "Complex modal analysis technology" is a technology that integrates and analyzes multiple data formats such as text, audio, and images.
[0576] "Emotional state" refers to an individual's temporary emotional state and is classified into categories such as positive, negative, and neutral.
[0577] "Service improvement information" refers to advice and instructions provided in real time based on the customer's emotional state, aimed at improving the quality of service.
[0578] The system that realizes this invention mainly uses a communication device, a processing device, and related analysis software. The communication device functions as a user device such as a smartphone or tablet and acquires input information from the user. This information is acquired as text, audio, or images, converted into digital data, and then transmitted to the processing device.
[0579] The processing unit functions as a server, analyzing the received input information. This analysis involves processing text using a natural language processing library (e.g., SpaCy) and clarifying the user's emotional state using a sentiment analysis engine (e.g., Google Cloud Natural Language). Compound modal analysis techniques are used to comprehensively analyze audio and image information.
[0580] Based on the results of sentiment analysis, the server provides real-time advice for service improvement. This advice is transmitted to a communication device and notified to staff. This enables rapid and accurate improvement of customer service within the store.
[0581] For example, in a restaurant, if a customer says they've "waited too long," the system can sense their stress and prompt staff for a quicker response. As an example of real-time service based on customer emotional state, an example of a prompt message is as follows: "Convert the customer's statement to text, determine their emotional state, and generate a service suggestion based on that."
[0582] This invention makes it possible to provide meticulous service that takes customer emotions into consideration, thereby improving the quality of service at stores.
[0583] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0584] Step 1:
[0585] Users input either voice or text using a smartphone, which is a communication device. This input is processed in real time as digital data on the device. If the input data is text, it is acquired as text data; if it is voice, it is acquired as voice data.
[0586] Step 2:
[0587] The terminal sends the acquired input data to the server, which acts as a processing unit. In the case of audio data, it is first converted to text, and the input data is then formatted into an appropriate data format such as JSON before being transferred to the server.
[0588] Step 3:
[0589] The server analyzes the received input data using natural language processing techniques. If text is input, it analyzes its structure and extracts important keywords and phrases. This process utilizes natural language processing libraries such as SpaCy.
[0590] Step 4:
[0591] The server uses complex modal analysis techniques to gain a deep understanding of the input data. In this step, it analyzes the user's emotional state through the analysis of speech intonation and frequently occurring phrases.
[0592] Step 5:
[0593] The server uses a generative AI model to determine the user's emotions based on the results obtained from the emotion analysis engine. This process utilizes emotion analysis tools such as Google Cloud Natural Language, resulting in output classified as positive, negative, neutral, etc.
[0594] Step 6:
[0595] The server generates real-time service improvement information based on the user's emotional state. This information is generated as a concrete action plan using prompts to encourage the delivery of the best possible service to the customer.
[0596] Step 7:
[0597] The terminal notifies store staff of service improvement information received from the server. The notification is provided via voice or text to the staff's dedicated terminal to prompt the next action.
[0598] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0599] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0600] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0601] [Fourth Embodiment]
[0602] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0603] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0604] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0605] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0606] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0607] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0608] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0609] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0610] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0611] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0612] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0613] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0614] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0615] This invention is a system for processing information obtained from users and providing appropriate care and psychological support. The embodiments thereof are described below.
[0616] The system of the present invention mainly consists of a server, a terminal, and a user interface. The user creates input data in text, voice, or image format using a communication terminal. For example, the user can input text requesting advice on caregiving or upload images of a specific situation.
[0617] The terminal's role is to acquire user input data in real time and send this data to the server. The transmitted data is encrypted and managed securely.
[0618] Upon receiving input data from a terminal, the server analyzes the text using natural language processing (NLP) techniques and determines the content of images and audio using multimodal analysis techniques. Based on this, the server estimates the user's intentions and the support they require, and understands their psychological state through sentiment analysis. Through these analyses, the server generates an appropriate response to the user's request.
[0619] The generated response is sent back to the terminal, where it is displayed to the user. In some cases, the response may be converted from text to speech and provided verbally. In particular, if the server determines that the user is in an emotionally distressed situation, it will generate a response that takes emergency measures into consideration.
[0620] For example, if a user enters a message such as, "My father is depressed, and I don't know what to do," the server analyzes this message and determines that the user needs psychological support within the family. The server generates relevant support information as a response, providing the user with "specific support measures" and "referrals to professional organizations as needed." Furthermore, if the server determines that the situation is urgent, it can automatically notify relevant external support organizations to encourage prompt action.
[0621] In this way, the system of the present invention provides flexible and rapid support that responds to diverse needs, and reduces the psychological and informational burden on caregivers.
[0622] The following describes the processing flow.
[0623] Step 1:
[0624] The user inputs questions and situation descriptions in text, voice, or image format via a communication terminal. The communication terminal converts this input data into digital signals.
[0625] Step 2:
[0626] The terminal transmits the converted digital signal to the server in real time. This communication is encrypted to ensure security.
[0627] Step 3:
[0628] The server analyzes the received data. For text, natural language processing techniques are used to perform grammatical analysis and keyword extraction. Audio data is converted to text using speech recognition technology, and image data is identified using image recognition algorithms.
[0629] Step 4:
[0630] The server performs sentiment analysis to understand the user's intentions based on the analyzed data. In this process, the user's emotional state is classified as positive, negative, or neutral, and the server adjusts the response accordingly.
[0631] Step 5:
[0632] The server generates a response based on the sentiment analysis results and the user's request. This generation process references a knowledge base and constructs a response that includes the latest care and psychological support information.
[0633] Step 6:
[0634] The server sends the generated response to the terminal. The terminal displays the received text response to the user and, if necessary, also provides a voice response using speech synthesis technology.
[0635] Step 7:
[0636] If the server determines that a user is in an emergency, it automatically notifies registered external support organizations. This notification includes concise and important information, enabling a swift response.
[0637] Step 8:
[0638] The terminal notifies the user of responses from the server and confirmation of emergency calls initiated by the server, preparing the user for their next action.
[0639] (Example 1)
[0640] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0641] Through the input and analysis of information, there is a challenge in providing appropriate and prompt responses when users require care or psychological support. Furthermore, providing individually customized support tailored to the emotional state of each user is also a crucial challenge.
[0642] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0643] In this invention, the server includes means for collecting data from users via information terminals, means for processing the data using natural language processing and various forms of analysis techniques in a data processing device, and means for evaluating the user's purpose and mental state based on the processing results. This enables the rapid provision of customized support tailored to the user's needs and allows for coordinated responses with external support organizations in emergencies.
[0644] An "information terminal" is a general term for electronic devices used by users to input and display data.
[0645] "User" refers to the entity that uses the system to input data and receives analysis results and responses.
[0646] A "data processing device" refers to a device that receives data transmitted from an information terminal and performs analysis and processing on it.
[0647] "Natural language processing" refers to the technology used by data processing devices to analyze human language and extract its meaning.
[0648] "Analysis techniques for diverse formats" refers to technologies that effectively analyze data in different formats, such as text, audio, and images.
[0649] "Evaluation" refers to the act of judging a user's purpose and mental state based on the analysis results.
[0650] "Response" refers to the response provided to the user based on the analysis results and evaluation.
[0651] "Speech synthesis technology" refers to the technology that converts text data into speech and outputs it as speech.
[0652] An "external support organization" refers to an external, specialized support organization that can integrate with the system in an emergency.
[0653] This invention is a system for providing flexible and rapid care and psychological support to users. It mainly includes information terminals, servers, and user-involved processes.
[0654] Users input information requiring support in text, voice, or image format using an information terminal. For example, they can input a question such as, "My mother hasn't been feeling well lately, and I'm worried," or images related to a specific situation. The information terminal collects this input data and securely transmits it to a data processing unit using encryption technology.
[0655] The server functions as a data processing unit, processing received data using specific analysis techniques. It analyzes text data using natural language processing (NLP) techniques and evaluates audio and image data in detail using multimodal analysis techniques.
[0656] The server utilizes a generative AI model to analyze the results and build support content optimized for the user. For example, if a user enters the sentence, "My father is depressed, and I don't know what to do," the server analyzes this and determines that psychological support is needed. It then provides appropriate support measures and information on partner organizations.
[0657] An example of a prompt message a user might use for input is, "My father is depressed, and I don't know what to do." Such input allows the system to respond to the user's individual needs. Furthermore, in emergencies, the server can assess the user's situation and automatically notify external support organizations as needed. This allows users to utilize the system with peace of mind.
[0658] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0659] Step 1:
[0660] Users create input data using an information terminal. For example, they might enter text describing the situation in which they need support, or related images. The data entered by the user is the first step toward providing accurate information services and is collected by the terminal after input.
[0661] Step 2:
[0662] The device acquires data entered by the user in real time. This data is encrypted to ensure security. The device then transmits the collected data to a server via a secure protocol. Input data can be in text, audio, or image format, while output is encrypted communication data.
[0663] Step 3:
[0664] The server receives encrypted data sent from the terminal and first decrypts it. Then, the server uses natural language processing (NLP) techniques to analyze the text data and process it to understand the type of support the user is seeking. It also performs detailed analysis of audio and image data using multimodal analysis. The input is encrypted data, and the output is information corresponding to the user's needs after the analysis.
[0665] Step 4:
[0666] The server determines the user's intentions and emotions based on the analyzed data. Using a generative AI model, it generates appropriate responses tailored to the user. At this stage, the server constructs specific support measures or information from the analysis results. The input is the analyzed user information, and the output is the generated response data.
[0667] Step 5:
[0668] The response data generated by the server is sent to the terminal. The terminal displays this response to the user. If voice output is required, the terminal also responds to the user verbally via text-to-speech synthesis. The input is the response data from the server, and the output is the displayed or audio information delivered to the user.
[0669] (Application Example 1)
[0670] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0671] There is a need to understand changes in a person's psychological and physical state in real time and provide appropriate support. However, conventional systems have had difficulty instantly detecting changes in a user's condition and providing specific support tailored to their living environment.
[0672] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0673] In this invention, the server includes means for acquiring input information from a human via a communication device, means for estimating the human's intentions and psychological state based on the analysis results, and means for generating an appropriate response based on the analyzed information. This makes it possible to monitor changes in the user's living environment and mood in real time and provide specific and appropriate support.
[0674] A "communication device" is a device that acquires input information from a human and transmits it to a central processing unit.
[0675] "Human" refers to the end-users who use the system, and their intentions and psychological states are analyzed through their input.
[0676] "Input information" refers to data provided by humans in various forms, such as text, audio, and images.
[0677] A "central processing unit" is a computer system that performs overall language processing and diverse analysis based on acquired input information and generates analysis results.
[0678] "Holistic language processing" is a technique that uses natural language processing technology to analyze input text data and understand its meaning.
[0679] "Diverse analysis technology" is a technology that analyzes image and audio data to identify and interpret its content.
[0680] "Intention" refers to the purpose or request that a person tries to convey through a system.
[0681] "Psychological state" refers to a person's emotions and mental state, and by estimating this, it becomes possible to provide appropriate support.
[0682] "Response" refers to support information and advice generated by the central processing unit based on the analysis results.
[0683] A "display device" is a device that monitors a person's living environment and circumstances, and displays a response visually or audibly as needed.
[0684] A "support organization" is an external professional organization or institution that is referred to when needed and provides support to individuals.
[0685] The system for realizing this invention consists of a communication device, a central processing unit, a display device, and cooperation with a support organization. First, the communication device acquires input information from humans, i.e., text, voice, and images, and transmits it to the central processing unit. This information is encrypted and securely sent to the central processing unit.
[0686] The central processing unit analyzes the acquired input information using holistic language processing techniques (such as TensorFlow and spaCy) and diverse analysis techniques (such as OpenCV and Librosa). This allows for accurate estimation of human intentions and psychological states. For example, it can estimate emotional changes from voice input or determine everyday behavior through image analysis.
[0687] Based on the analysis results, the central processing unit generates the optimal response. This response is presented to a human using speech synthesis technology as needed. In emergencies, it can also automatically notify the support organizations that require assistance.
[0688] For example, if a user inputs "I've been feeling unwell lately and having difficulty with daily tasks," the central processing unit can analyze this data and provide suggestions for relaxation or referrals to medical institutions. This allows users to receive timely and appropriate support.
[0689] An example of a prompt might be, "Generate a program that suggests relaxation methods when the user appears tired." Such prompts allow the generating AI model to formulate processing steps to indicate appropriate intervention methods.
[0690] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0691] Step 1:
[0692] The user provides input information using a communication device. This information can be in text, audio, or image format, and different initial processing is performed depending on the format. The user's input is temporarily stored by the communication device. Subsequently, this input information is transmitted securely to the central processing unit using encryption technology.
[0693] Step 2:
[0694] The server acquires input information received from the communication device and performs initial analysis according to the data format. In the case of audio data, it is converted to text using speech recognition technology, and for image data, key features are extracted using image processing technology. This allows the server to generate a standardized data format for subsequent processing.
[0695] Step 3:
[0696] The server uses a global language processing engine to analyze the meaning of text data. This process extracts keywords and context from the input information to understand the user's intent. For example, if keywords such as "tired" are included, the server proceeds to estimate the related psychological state based on these keywords.
[0697] Step 4:
[0698] The server uses diverse analysis techniques to detect the user's daily activities and changes in facial expressions from image data. This generates complementary data that allows for more accurate emotion estimation based on the information obtained from the images. This data makes it possible to visually understand the user's situation.
[0699] Step 5:
[0700] Based on the analysis results, the server applies a generative AI model to generate the optimal response. This model performs inference according to the prompt text and designs a response that is appropriate to the user's state. The response is generated in text format and then converted to speech using speech synthesis technology as needed.
[0701] Step 6:
[0702] The server sends the generated response to the communication device. The communication device displays the response as text or plays it back as audio. Specifically, it provides suggestions for relaxation methods and contact information for support organizations as needed. As a result, the user is encouraged to take appropriate action and is in a position to receive support.
[0703] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0704] This invention is a novel care support system that incorporates an emotion engine to recognize the user's emotions and provide highly personalized support based on those emotions. The system includes a communication terminal, a server, and an emotion engine.
[0705] Users can input questions and information about caregiving using a communication terminal. The input data is appropriately digitized on the terminal and sent to the server. Upon receiving the input data, the server first analyzes it using natural language processing and multimodal analysis techniques. This analysis involves sentence structure analysis if the data is text, speech recognition if it is audio, and image recognition if it is images to understand the content.
[0706] After this data analysis, the server uses an emotion engine to identify emotions from the user's input. The emotion engine analyzes, for example, the tone of text, the intonation of speech, and facial expressions in images, and classifies them into emotion categories such as positive, negative, and neutral. Based on the results of the emotion engine, the user's psychological state is accurately understood.
[0707] The server then synthesizes these analysis results and proceeds to a process of generating a response that aligns with the user's intentions. This response generation process includes considering the emotions identified by the emotion engine and creating a detailed response that matches the support the user desires and their psychological state. The generated response is sent to the terminal as text and, in some cases, presented as speech using speech synthesis technology.
[0708] For example, if a user enters "My mother hasn't been feeling well lately, and I'm worried about spending time with her," the server analyzes this message and uses an emotion engine to read the user's anxiety. The server then generates emotional support and practical advice that corresponds to this emotion, providing a response such as, "I recommend you wait and see for a while, and consult your doctor if necessary. We can also think of activities to cheer her up together."
[0709] This system allows users to receive care support that is tailored to their emotional needs, thereby reducing the burden of daily caregiving.
[0710] The following describes the processing flow.
[0711] Step 1:
[0712] Users input questions or concerns in text, voice, or image format using a communication terminal. The communication terminal digitizes the input data and prepares it to be immediately transmitted to the server.
[0713] Step 2:
[0714] The terminal transmits digitized input data to the server. The communication is encrypted to ensure the security of user data.
[0715] Step 3:
[0716] The server analyzes the received data. For text, it uses natural language processing for syntactic analysis and keyword extraction; for speech, it uses speech recognition to convert it to text; and for images, it uses image recognition technology to determine visual elements.
[0717] Step 4:
[0718] The server passes the analyzed data through the emotion engine. The emotion engine analyzes the emotions contained in the user's input and assigns one of three emotion labels: positive, negative, or neutral.
[0719] Step 5:
[0720] The server comprehensively understands the user's intent and emotions based on the results of natural language processing and the emotion engine. It then moves on to a process of generating the most appropriate response for the user based on this information.
[0721] Step 6:
[0722] The server uses a generative AI model to create responses that align with the user's intentions. These responses include emotional support and practical advice based on the emotions indicated by the emotion engine.
[0723] Step 7:
[0724] The server sends the completed response to the terminal. The terminal displays the received response to the user. If necessary, the response can also be provided in voice using speech synthesis technology.
[0725] Step 8:
[0726] Under the monitoring of the emotion engine, if the server determines that user input is urgent, it will automatically notify pre-configured external support organizations and initiate procedures to request prompt assistance.
[0727] (Example 2)
[0728] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0729] In modern society, caregivers are required to respond quickly and appropriately to increasingly complex needs. However, caregivers often find it difficult to understand the underlying emotions and psychological states of users, limiting their ability to provide appropriate support. This problem highlights the need to develop systems that enable even non-professionals to provide appropriate advice and support tailored to the situation.
[0730] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0731] In this invention, the server includes means for analyzing information from the user using language processing technology and various data analysis technologies, means for identifying the user's abstract intentions and emotional state based on the analysis results, and means for creating an individualized response based on the identified information. This makes it possible to quickly provide individualized support that is in line with the user's emotions and situation, reduce the burden on caregivers, and improve the quality of care.
[0732] A "communication device" is a device that acquires information from a user and transmits it to a data processing device.
[0733] A "data processing device" is a device that analyzes received information to identify the user's intentions and emotional state.
[0734] "Language processing technology" refers to techniques for analyzing text data and understanding its structure and meaning.
[0735] "Diverse data analysis technologies" refer to technologies for analyzing data in various formats, such as text, audio, and images, and understanding their content.
[0736] A "personalized response" is a response that is generated in a way that is tailored to the user's specific situation and emotions.
[0737] "Abstract user intent" refers to the goals or needs that users have implicitly but do not explicitly express.
[0738] "Emotional state" refers to the psychological state a user exhibits, such as positive, negative, or neutral emotions.
[0739] One embodiment of the present invention is configured as a care support system incorporating emotion recognition technology. This system includes a communication terminal, a data processing device, and an emotion engine.
[0740] Users can input questions and information about caregiving using a communication terminal. The communication terminal digitizes this input and transmits it to a data processing device. Encryption and secure communication protocols (e.g., HTTPS) are used here.
[0741] The data processing device analyzes the received data using language processing techniques (e.g., natural language processing libraries) and various data analysis techniques (speech recognition software, image recognition algorithms, etc.). If it is text, structural analysis is performed; if it is audio, it is converted to text; and if it is an image, important features are extracted to understand the context.
[0742] After data analysis, the data processing unit utilizes an emotion engine to identify the user's emotional state. The emotion engine analyzes text tone, voice intonation, and facial expressions in images to categorize the user's emotions as positive, negative, neutral, etc.
[0743] For example, if a user enters "My mother hasn't been feeling well lately, and I'm worried about spending time with her," the device sends this message to the data processing unit. The data processing unit uses an emotion engine to identify the user's anxiety and generates a reassuring response. For example, it might provide a response such as, "I recommend you wait and see for a while, and consult your doctor if necessary. We can also think of activities to cheer her up together."
[0744] An example of a prompt message for a generative AI model would be: "Identify the emotions the user is feeling and suggest care support based on those emotions."
[0745] In this way, the system can provide personalized care support tailored to the user's emotions, reduce the burden on caregivers, and improve the quality of care.
[0746] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0747] Step 1:
[0748] Users input questions and information about caregiving into a communication terminal. This input can be text, voice, or images. The input information is converted into a digital format. For example, voice input is saved as an audio file.
[0749] Step 2:
[0750] The terminal transmits digitized information to a data processing unit. The data is protected using encryption technology and transmitted via a secure protocol (e.g., HTTPS). This ensures the data is transferred safely.
[0751] Step 3:
[0752] The server analyzes the received data using language processing techniques and various data analysis techniques. For text data, it analyzes the grammatical structure and extracts meaning. Speech data is converted to text using speech recognition technology, and image data has its main features extracted using image processing technology. The output of the analysis provides unified information in the form of text data.
[0753] Step 4:
[0754] The server uses an emotion engine to identify emotional states from the analyzed data. The emotion engine classifies the user's emotions as positive, negative, or neutral based on text tone, voice intonation, and image facial expressions. This allows the user's psychological state to be identified.
[0755] Step 5:
[0756] The server generates a response to the user based on the results of the emotion engine. Response generation creates a personalized message based on the identified emotion and the user's intent. For example, it might create a message containing support and advice for a worried user. This generated response is in text format and can be converted to speech format using speech synthesis technology as needed.
[0757] Step 6:
[0758] The terminal displays the response received from the server to the user. Text responses are displayed on the screen, and audio responses are played through the speaker. This allows the user to receive support based on their own emotions.
[0759] (Application Example 2)
[0760] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0761] In modern brick-and-mortar stores, it is difficult to instantly grasp customer emotions and provide appropriate service. Furthermore, while implementing systems that can sense customer needs and emotions in real time and provide services accordingly would greatly improve a store's competitiveness, such systems are not yet sufficiently available.
[0762] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0763] In this invention, the server includes means for acquiring input information from a user via a communication device, means for transmitting the acquired input information to a processing device, and means for analyzing the received input information in the processing device using natural language processing and complex modal analysis techniques. This enables real-time sensing of the customer's emotional state in a physical store and the rapid provision of services based on that understanding.
[0764] A "communication device" is a device used to acquire information from a user and transmit it to another device.
[0765] A "user" is a person or entity that uses the system and provides information.
[0766] "Input information" refers to data transmitted by the user, including text, audio, images, and other forms.
[0767] A "processing device" is a device that analyzes received information and generates instructions based on that analysis.
[0768] "Natural language processing" is a technology that uses computers to process and understand human language.
[0769] "Complex modal analysis technology" is a technology that integrates and analyzes multiple data formats such as text, audio, and images.
[0770] "Emotional state" refers to an individual's temporary emotional state and is classified into categories such as positive, negative, and neutral.
[0771] "Service improvement information" refers to advice and instructions provided in real time based on the customer's emotional state, aimed at improving the quality of service.
[0772] The system that realizes this invention mainly uses a communication device, a processing device, and related analysis software. The communication device functions as a user device such as a smartphone or tablet and acquires input information from the user. This information is acquired as text, audio, or images, converted into digital data, and then transmitted to the processing device.
[0773] The processing unit functions as a server, analyzing the received input information. This analysis involves processing text using a natural language processing library (e.g., SpaCy) and clarifying the user's emotional state using a sentiment analysis engine (e.g., Google Cloud Natural Language). Compound modal analysis techniques are used to comprehensively analyze audio and image information.
[0774] Based on the results of sentiment analysis, the server provides real-time advice for service improvement. This advice is transmitted to a communication device and notified to staff. This enables rapid and accurate improvement of customer service within the store.
[0775] For example, in a restaurant, if a customer says they've "waited too long," the system can sense their stress and prompt staff for a quicker response. As an example of real-time service based on customer emotional state, an example of a prompt message is as follows: "Convert the customer's statement to text, determine their emotional state, and generate a service suggestion based on that."
[0776] This invention makes it possible to provide meticulous service that takes customer emotions into consideration, thereby improving the quality of service at stores.
[0777] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0778] Step 1:
[0779] Users input either voice or text using a smartphone, which is a communication device. This input is processed in real time as digital data on the device. If the input data is text, it is acquired as text data; if it is voice, it is acquired as voice data.
[0780] Step 2:
[0781] The terminal sends the acquired input data to the server, which acts as a processing unit. In the case of audio data, it is first converted to text, and the input data is then formatted into an appropriate data format such as JSON before being transferred to the server.
[0782] Step 3:
[0783] The server analyzes the received input data using natural language processing techniques. If text is input, it analyzes its structure and extracts important keywords and phrases. This process utilizes natural language processing libraries such as SpaCy.
[0784] Step 4:
[0785] The server uses complex modal analysis techniques to gain a deep understanding of the input data. In this step, it analyzes the user's emotional state through the analysis of speech intonation and frequently occurring phrases.
[0786] Step 5:
[0787] The server uses a generative AI model to determine the user's emotions based on the results obtained from the emotion analysis engine. This process utilizes emotion analysis tools such as Google Cloud Natural Language, resulting in output classified as positive, negative, neutral, etc.
[0788] Step 6:
[0789] The server generates real-time service improvement information based on the user's emotional state. This information is generated as a concrete action plan using prompts to encourage the delivery of the best possible service to the customer.
[0790] Step 7:
[0791] The terminal notifies store staff of service improvement information received from the server. The notification is provided via voice or text to the staff's dedicated terminal to prompt the next action.
[0792] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0793] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0794] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0795] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0796] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0797] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0798] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0799] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0800] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0801] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0802] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0803] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0804] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0805] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0806] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0807] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0808] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0809] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0810] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0811] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0812] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0813] The following is further disclosed regarding the embodiments described above.
[0814] (Claim 1)
[0815] A means of obtaining input data from a user via a communication terminal,
[0816] A means of sending the acquired input data to the server,
[0817] The server includes means for analyzing received input data using natural language processing and multimodal analysis techniques,
[0818] A means for estimating the user's intentions and psychological state based on the analysis results,
[0819] A system that includes means for generating an appropriate response based on estimated information and transmitting it to a communication terminal.
[0820] (Claim 2)
[0821] The system according to claim 1, wherein the generated response is presented to the user using speech synthesis technology.
[0822] (Claim 3)
[0823] The system according to claim 1, wherein the server automatically notifies a registered external support organization when it determines that the user's situation is urgent.
[0824] "Example 1"
[0825] (Claim 1)
[0826] A means of collecting data from users via information terminals,
[0827] A means for transmitting the collected data to a data processing device,
[0828] A data processing device includes means for processing received data using natural language processing and various analytical techniques,
[0829] A means for evaluating the user's purpose and mental state based on the processing results,
[0830] A system that includes means for constructing an appropriate response based on evaluated information and transmitting it to an information terminal.
[0831] (Claim 2)
[0832] The system according to claim 1, wherein the constructed response is presented to the user using speech synthesis technology.
[0833] (Claim 3)
[0834] The system according to claim 1, wherein the data processing device automatically notifies a registered external support organization when it determines that the user's situation is urgent.
[0835] "Application Example 1"
[0836] (Claim 1)
[0837] A means of acquiring input information from a human via a communication device,
[0838] Means for transmitting acquired input information to a central processing unit,
[0839] In the central processing unit, means for analyzing received input information using overall language processing and diverse analysis techniques,
[0840] A means for estimating human intentions and psychological states based on the analysis results,
[0841] A means for generating an appropriate response based on estimated information and transmitting it to a communication device,
[0842] A system that includes means of monitoring changes in the user's living environment and mood in real time using display devices and providing appropriate support.
[0843] (Claim 2)
[0844] The system according to claim 1, wherein the generated response is presented to a human using speech synthesis technology.
[0845] (Claim 3)
[0846] The system according to claim 1, wherein the central processing unit automatically notifies a registered external support organization when it determines that a person's situation is urgent.
[0847] "Example 2 of combining an emotion engine"
[0848] (Claim 1)
[0849] A means of obtaining information from users through a communication device,
[0850] A means for transmitting acquired information to a data processing device,
[0851] A data processing device includes means for analyzing received information using language processing techniques and various data analysis techniques,
[0852] Based on the analysis results, a means for identifying the user's abstract intentions and emotional state,
[0853] A system including means for creating an individualized response based on identified information and transmitting it to a communication device.
[0854] (Claim 2)
[0855] The system according to claim 1, wherein the generated response is provided to the user using speech generation technology.
[0856] (Claim 3)
[0857] The system according to claim 1, wherein when the data processing device determines that the user's situation is critical, it automatically notifies a registered external support organization.
[0858] "Application example 2 when combining with an emotional engine"
[0859] (Claim 1)
[0860] A means of obtaining input information from the user via a communication device,
[0861] Means for transmitting acquired input information to a processing device,
[0862] The processing device includes means for analyzing received input information using natural language processing and composite modal analysis techniques,
[0863] A means for estimating the user's intentions and psychological state based on the analysis results,
[0864] A means for generating an appropriate response based on estimated information and transmitting it to a communication device,
[0865] A system that includes a means of providing staff with information in real time to improve services based on the emotional state of users.
[0866] (Claim 2)
[0867] The system according to claim 1, wherein the generated response is presented to the user using speech conversion technology.
[0868] (Claim 3)
[0869] The system according to claim 1, wherein the processing device automatically notifies a registered external support organization when it determines that the user's situation is urgent. [Explanation of Symbols]
[0870] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining input data from a user via a communication terminal, A means of sending the acquired input data to the server, The server includes means for analyzing received input data using natural language processing and multimodal analysis techniques, A means for estimating the user's intentions and psychological state based on the analysis results, A system that includes means for generating an appropriate response based on estimated information and transmitting it to a communication terminal.
2. The system according to claim 1, wherein the generated response is presented to the user using speech synthesis technology.
3. The system according to claim 1, wherein the server automatically notifies a registered external support organization when it determines that the user's situation is urgent.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A