System
The system addresses the challenge of continuous emotion monitoring and social reintegration by analyzing user inputs to provide real-time psychological support and interactive practice, ensuring timely professional intervention and enhancing social reintegration support.
Patent Information
- Application Number
- JP2024126238
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2026-02-13
AI Technical Summary
Current methods struggle to continuously monitor individuals' emotions and provide immediate mental health support, lack systematized collaboration with experts, and insufficiently support social reintegration, making it difficult to maintain mental health and facilitate smooth reintegration into daily life and social activities.
A system that analyzes users' voice or text input to determine emotions, provides real-time psychological support, requests specialist intervention if necessary, and supports interactive practice sessions, communication, and user matching to enhance social reintegration.
Enables continuous monitoring and timely professional support, facilitates effective mental health maintenance and social reintegration by providing appropriate counseling and interactive practice scenarios, and strengthens support systems for users.
Smart Images

Figure 2026023917000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, maintaining mental health and supporting social reintegration are important issues. Early detection of mental illness, appropriate treatment, and smooth reintegration into daily life and social activities are particularly important. However, current methods make it difficult to continuously monitor the emotions and state of individual users and respond immediately. Furthermore, collaboration with experts and support for social reintegration are not yet systematized. The present invention aims to solve these problems. [Means for solving the problem]
[0005] The system of the present invention includes a means for receiving and analyzing a user's voice or text input to determine their emotions. This allows for continuous monitoring of the user's mental state. Then, by generating and providing appropriate counseling messages based on the emotions, it is possible to provide real-time psychological support to the user. It also includes a means for determining the level of urgency based on the emotions and, if high, requesting a specialist, thereby enabling timely professional support. It also includes a means for receiving practice requests for everyday situations, generating scenarios based on the requests, and conducting interactive practice sessions. This allows the user to prepare for real-world situations. It also includes a means for matching with other users and supporting communication, thereby strengthening support systems after the user's return to society. Providing this system to businesses makes it possible to expand mental health support to a larger number of users.
[0006] "User" refers to an individual who uses this system and is a person who receives mental health support and social reintegration support.
[0007] "Voice or text input" refers to voice data or character data that the user provides to the system, and is information that the system uses to analyze the user's emotions and state.
[0008] The term "server" refers to a computer system that receives user input data and performs processes such as emotion analysis, generating counseling messages, and determining the level of urgency.
[0009] A "terminal" is a device that allows a user to input voice or text and receive messages from a server, and includes smartphones, tablets, PCs, etc.
[0010] "Sentiment analysis" is the process of determining a user's emotions based on their voice or text input using natural language processing and machine learning.
[0011] A "counseling message" is a message that the system generates based on the user's emotions and state and provides to the user, including advice and support information.
[0012] "Urgency assessment" is the process by which the system evaluates the urgency of a situation based on the user's emotional state and, if necessary, requests a specialist to respond.
[0013] An "expert" is a person who has knowledge and skills related to mental health, such as a psychologist, psychiatrist, or counselor, and who provides professional advice or treatment to users.
[0014] A "response request" is an action in which the system instructs or requests an expert to respond to a user's emergency situation.
[0015] A "practice scenario" is an interactive situation generated by the system that allows users to practice preparing for everyday situations or specific situations.
[0016] "Feedback" refers to evaluation and advice information that the system generates based on the results of practice and the user's condition and provides to the user.
[0017] "Matching" is the process by which the system searches and selects other users with the same position or situation from a database and pairs them appropriately.
[0018] "Communication" refers to the dialogue and message exchanges in which matched users exchange information and feelings.
[0019] "Monitoring" is the process by which the system periodically checks the content of user communications and takes action when it detects an abnormality or problem.
[0020] "Intervention" is the act of the system instructing an expert to intervene in a particular problem or emergency situation. [Brief explanation of the drawings]
[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0023] First, the terms used in the following description will be explained.
[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0029] [First embodiment]
[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0042] The present invention is a system that analyzes a user's voice or text input, determines their emotions, and provides appropriate counseling messages. This system is made up of the following components:
[0043] Components
[0044] 1. User Device:
[0045] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[0046] 2. Server:
[0047] It receives input data, analyzes emotions, generates counseling messages, and determines the level of urgency. It also generates practice scenarios, matches users with other users, and monitors communication content.
[0048] 3. Experts:
[0049] Based on requests from the server, appropriate advice and treatment will be provided to users, including psychologists, psychiatrists, and counselors.
[0050] Program processing
[0051] Handling User Input
[0052] The user inputs data by voice or text, and the device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server receives this data, performs emotional analysis, and determines that the user is "feeling stressed."
[0053] Counseling message generation
[0054] The server generates a counseling message based on the results of emotion analysis. For example, if it determines that the user is feeling stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[0055] Urgency assessment and response request
[0056] The server continuously monitors the emotion data and determines the level of urgency. If the level is deemed high, it requests a specialist to respond. For example, if a user continuously expresses depressed feelings, the server sends an alert to the specialist. The specialist receives the alert, contacts the user, and provides appropriate counseling or treatment.
[0057] Providing practice scenarios
[0058] When a user wishes to practice an everyday situation, for example, they make a request such as "I would like to practice an interview." The server receives this request and generates an appropriate interview scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on the scenario. The server analyzes the practice results, generates feedback, sends it to the device, and provides it to the user.
[0059] Matching and communication support
[0060] When a user wishes to communicate with other users, they can enter, for example, "I want to talk to someone who is also trying to reintegrate into society." The server searches and selects users in the same position from its database and performs matching. The matching results are sent to the user's device and notified. The user then begins a conversation with the matched person, and the system monitors the conversation. If necessary, the server requests expert intervention.
[0061] Specific examples
[0062] 1. Sentiment analysis and counseling message provision:
[0063] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this speech into text and sends it to the server. The server determines that the user is "feeling stressed," and generates a message such as "Take a deep breath to relax," which is sent to the device. The device then displays this message to the user.
[0064] 2. Urgency assessment and expert intervention:
[0065] The user repeatedly types "I feel heavy." The server analyzes this and determines that the situation is urgent. It sends an alert to an expert, who then contacts the user and provides counseling.
[0066] 3. Providing practice scenarios and feedback:
[0067] The user requests, "I want to practice an interview." The server generates an appropriate scenario and sends it to the device. The user practices interactively, and the server analyzes the results and sends feedback.
[0068] 4. Matching and communication support:
[0069] The user types, "I want to talk to people who are also trying to reintegrate into society." The server matches suitable users and notifies the device. The users begin a conversation, and the system monitors the content. If necessary, it requests intervention from an expert.
[0070] As described above, the system of the present invention can monitor the mental health state of the user and provide appropriate counseling and support for rehabilitation into society.
[0071] The processing flow will be explained below.
[0072] AI that listens to your heart
[0073] User Input and Sentiment Analysis
[0074] Step 1:
[0075] The user speaks or texts into the device, saying, "I've been feeling stressed lately because things haven't been going well at work."
[0076] Step 2:
[0077] The terminal converts the voice input into text data and transmits the text data to the server.
[0078] Step 3:
[0079] A server receives the text data and runs a sentiment analysis model (e.g., using natural language processing).
[0080] Step 4:
[0081] The server determines from the user's text that he or she is "feeling stressed."
[0082] Counseling message generation
[0083] Step 5:
[0084] The server generates a counseling message for stress reduction (e.g., "Take a deep breath and you'll feel more relaxed").
[0085] Step 6:
[0086] The server sends the generated counseling message to the terminal.
[0087] Step 7:
[0088] The terminal displays a counseling message to the user.
[0089] AI in harmony with specialists
[0090] Urgency assessment and response request
[0091] Step 1:
[0092] The server monitors the user's daily emotional data.
[0093] Step 2:
[0094] The server analyzes the emotional data and determines whether the person is experiencing a continuous state of depression.
[0095] Step 3:
[0096] The server determines the level of urgency and, if deemed high, sends an alert to an expert.
[0097] Step 4:
[0098] A specialist receives the alert from the server and contacts the user to provide counseling or treatment.
[0099] AI provides a platform for social reintegration
[0100] Providing practice scenarios
[0101] Step 1:
[0102] The user inputs a request to the terminal saying, "I want to practice for an interview."
[0103] Step 2:
[0104] The server receives the user's request and generates an interview scenario.
[0105] Step 3:
[0106] The server sends the generated scenario to the terminal.
[0107] Step 4:
[0108] The terminal displays a scenario, and the user practices the interview in an interactive format.
[0109] Step 5:
[0110] When the user completes the exercise, the terminal transmits the exercise results to the server.
[0111] Providing Feedback
[0112] Step 6:
[0113] The server analyzes the practice results and generates a feedback message.
[0114] Step 7:
[0115] The server sends a feedback message to the terminal, which displays it to the user.
[0116] AI that provides an empathetic community
[0117] Matching with user requests
[0118] Step 1:
[0119] The user inputs their request into the terminal, saying, "I would like to talk to someone who is also aiming to return to society."
[0120] Step 2:
[0121] The server analyzes users in the same position from the database and matches them with the appropriate users.
[0122] Step 3:
[0123] The server sends the matching results to the device.
[0124] Step 4:
[0125] The user begins communicating with the matched person (e.g., "I see you're in the same situation, ○○. Let's both do our best.").
[0126] Communications monitoring and support
[0127] Step 5:
[0128] The server monitors interactions within the community.
[0129] Step 6:
[0130] If necessary, the server calls for expert intervention.
[0131] Through the above steps, this system can comprehensively support the user's mental health maintenance and social reintegration.
[0132] Example 1
[0133] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0134] Conventional counseling systems struggle to quickly and accurately analyze users' emotions and provide appropriate counseling messages. Furthermore, delays in assessing the level of urgency and providing appropriate expert intervention can lead to a deterioration in the user's mental health. Furthermore, they lack sufficient practice in everyday situations and support for communication between users, preventing them from effectively supporting users' reintegration into society.
[0135] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0136] In this invention, the server includes means for receiving a user's voice or text input, means for converting the input into text data, means for transmitting the text data to the server, means for analyzing the text data to determine the user's emotion, means for generating a counseling message using a generative AI model based on the result of the emotion analysis, means for transmitting the counseling message to a user terminal, and means for displaying the counseling message to the user on the user terminal. This makes it possible to quickly and accurately analyze the user's emotion and effectively support appropriate counseling and social reintegration.
[0137] "Voice or text input" refers to the form of speech or written information that a user uses to communicate their feelings or requests to a system.
[0138] "Text data" refers to data obtained by converting voice input into text format, or text information entered directly by the user.
[0139] "Server" refers to the central computing system that receives and analyzes voice or text data, generates counseling messages, determines urgency, etc.
[0140] "Sentiment analysis" refers to the process of analyzing a user's text data to determine the user's emotional state.
[0141] A "generative AI model" refers to an artificial intelligence algorithm, primarily a natural language generation model, that generates appropriate messages based on the results of user sentiment analysis.
[0142] "Counseling message" refers to a message of advice or encouragement provided to a user that is generated based on the results of an analysis of the user's emotions.
[0143] "Urgency" refers to an index that evaluates the urgency of the user's emotional state and indicates whether a prompt response is required.
[0144] "Professional" refers to a psychologist, psychiatrist, counselor, or other person qualified to assist users with their mental health.
[0145] A "practice scenario" refers to a simulation scenario generated by the server when a user wishes to practice a particular situation (such as an interview).
[0146] "Feedback" refers to evaluations and advice provided to users based on the results of their practice.
[0147] "Matching" refers to the process by which a server selects and connects a suitable partner when a user wishes to communicate with another user.
[0148] "User terminal" refers to a device (smartphone, tablet, PC, etc.) used to input voice or text and receive and display counseling messages and feedback from the server.
[0149] The "HTTPS protocol" refers to a communication protocol for secure data communication over the Internet.
[0150] "Natural language processing tool" refers to a software tool (e.g., spaCy, NLTK, etc.) used to analyze a user's text data.
[0151] The present invention is a system that analyzes a user's voice or text input, determines their emotions, and provides appropriate counseling messages. This can support the user's mental health and effectively assist them in their social reintegration. A specific embodiment of this system is described below.
[0152] System Configuration
[0153] This system consists of user devices, servers, and expert components. Specific hardware components for user devices include smartphones, tablets, and PCs. Servers are high-performance servers (e.g., cloud virtual machines) with various software installed.
[0154] User terminal: A device that receives voice or text input, converts it into text data, and sends it to a server.
[0155] Server: This is a critical computing system that performs sentiment analysis, generates counseling messages, and determines urgency. Specifically, it uses natural language processing tools (e.g., spaCy, NLTK) and machine learning libraries (e.g., TensorFlow, PyTorch).
[0156] Experts: Psychologists, psychiatrists, counselors, or other qualified individuals available to support users' mental health.
[0157] Example of operation
[0158] Processing voice or text input
[0159] The user inputs data by voice or text. The user's device receives this data, and in the case of voice input, it converts it into text data using the Google Speech-to-Text API. For example, if the user says, "Recently, things haven't been going well at work and I'm feeling stressed," the content is converted into text. This text data is sent to the server using the HTTPS protocol.
[0160] Emotion analysis
[0161] The server analyzes the received text data and determines the user's emotions. Using a natural language processing tool (e.g., spaCy), it determines that the user is "feeling stressed." The emotion analysis algorithm is trained using a machine learning library (e.g., TensorFlow).
[0162] Counseling message generation
[0163] Based on the results of the sentiment analysis, a generative AI model (e.g., GPT-3) is used to generate an appropriate counseling message. For example, a message such as "Take a deep breath and you'll feel more relaxed. Try it out." This message is sent to the user's device and immediately displayed to the user.
[0164] Urgency assessment and expert intervention
[0165] The server continuously monitors the user's emotional data and determines the level of urgency. If the user repeatedly inputs "I feel heavy," the level of urgency is determined to be high. If the level of urgency is high, the server sends an alert to an expert, who then contacts the user and provides counseling.
[0166] Providing practice scenarios
[0167] When a user requests "I want to practice for an interview," the server generates an appropriate practice scenario. The scenario contains specific questions, and the contents of the scenario are sent to the user's terminal. The user practices interactively, and the server analyzes the results and generates feedback.
[0168] Prompt Sentence Examples
[0169] Emotion Analysis Prompt: Enter "I've been feeling stressed lately because things haven't been going well at work."
[0170] Urgency determination prompt: Enter "I feel heavy" repeatedly.
[0171] Practice Scenario Prompt: Type "I want to practice interviewing."
[0172] Matching prompt: Enter "I'd like to talk to someone who is also trying to reintegrate into society."
[0173] As described above, the system of the present invention is realized using advanced technology to support the user's mental health and effectively assist in their reintegration into society.
[0174] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0175] Step 1:
[0176] The user inputs data by voice or text. For example, they might say, "I've been feeling stressed lately because things haven't been going well at work." This input becomes the starting point for the system to process the data.
[0177] Input: User voice or text
[0178] Output: Raw audio or text data
[0179] Step 2:
[0180] The device converts the voice input into text data. Specifically, it uses voice recognition software (e.g., Google Speech-to-Text API) to convert the voice to text. If the input is text, it proceeds to the next step.
[0181] Input: Audio data
[0182] Output: Text data
[0183] Step 3:
[0184] The device sends the text data to the server, specifically using the HTTPS protocol to transfer the data securely.
[0185] Input: Text data
[0186] Output: Text data is sent to the server
[0187] Step 4:
[0188] The server analyzes the received text data and determines the user's emotions. Specifically, it analyzes the text using natural language processing tools (e.g., spaCy, NLTK) and classifies the emotions using machine learning models (e.g., TensorFlow).
[0189] Input: Text data
[0190] Output: Sentiment analysis result (e.g., user's emotion is "stress")
[0191] Step 5:
[0192] The server generates a counseling message using a generative AI model (e.g., GPT-3) based on the results of emotion analysis. For example, it automatically generates a message such as, "Try taking deep breaths to relax. Give it a try."
[0193] Input: Sentiment analysis results
[0194] Output: Counseling message
[0195] Step 6:
[0196] The server sends the generated counseling message to the terminal. The server sends the message using the HTTPS protocol, and the terminal receives it.
[0197] Input: Counseling message
[0198] Output: Counseling message forwarded to terminal
[0199] Step 7:
[0200] The terminal displays the counseling message to the user. Specifically, the notification function is used to make the message immediately visible to the user.
[0201] Input: Counseling message
[0202] Output: A message displayed to the user
[0203] Step 8:
[0204] The server continuously monitors the emotion data and determines the level of urgency. If the user repeatedly inputs "I feel heavy," the level of urgency is determined to be high.
[0205] Input: Continuous emotion data
[0206] Output: Urgency judgment result (e.g., high urgency)
[0207] Step 9:
[0208] If the emergency is deemed high, the server will request a response from an expert. For example, an alert email will be sent to a psychologist requesting immediate action.
[0209] Input: Urgency assessment result
[0210] Output: An alert email is sent to the expert
[0211] Step 10:
[0212] When a user requests "I want to practice for an interview," the server generates an appropriate practice scenario, which includes specific questions and sends the contents of the scenario to the terminal.
[0213] Input: Practice Request
[0214] Output: Practice scenario
[0215] Step 11:
[0216] The user practices interactively, and the server analyzes the results and generates feedback, such as "Your speech is clear and good."
[0217] Input: Practice results
[0218] Output: Feedback
[0219] (Application example 1)
[0220] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0221] In recent years, there has been an increasing demand for counseling systems to maintain users' mental health. However, existing systems lack sufficient analysis of users' emotions and assessment of the level of urgency, making it difficult to respond quickly and effectively. Furthermore, as security risks increase, systems that can provide immediate security responses based on users' emotional states are needed. Furthermore, more advanced responses using generative AI models are needed to improve the quality of practice and feedback for everyday situations requested by users. A system that can solve these issues and comprehensively support users' mental health and safety is needed.
[0222] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0223] In this invention, the server includes: means for receiving and analyzing voice or text input and determining a user's emotion; means for generating a counseling message using a generative AI model based on the determined emotion; means for providing the generated counseling message to the user; means for determining the urgency based on the emotion and requesting a response from an expert; means for issuing a security alert based on the result of the emotion analysis; means for cooperating with a security device when an alert is issued; means for receiving a practice request for an everyday situation from the user, generating a practice scenario based on the request, practicing interactively with the user, and generating feedback based on the practice results; and means for generating prompt sentences based on emotion data and generating feedback using a generative AI model. This enables the provision of counseling messages that are responsive to the user's emotional state and rapid response in high-urgency situations, and enables the immediate issuance of alerts even in high-security risk situations and the provision of comprehensive safety support in cooperation with other devices.
[0224] "Voice or text input" refers to the means by which a user provides voice or text data to a system.
[0225] "Means for determining emotion" refers to a method or device for analyzing received voice or text input and classifying the user's current emotional state as "positive," "negative," "neutral," or the like.
[0226] The "means for generating a counseling message" refers to a method or device for generating a message including appropriate advice or advice based on the result of the user's emotion determination.
[0227] A "generative AI model" is an artificial intelligence model trained based on large amounts of data, and is a means used for emotion analysis and generating counseling messages.
[0228] The "means for determining the degree of urgency" refers to a method or device for analyzing the emotional state of the user and evaluating and determining the degree of urgency of that state.
[0229] "Means for requesting a response from an expert" refers to a method or device for requesting intervention from an expert such as a psychologist or psychiatrist when the emergency is determined to be high.
[0230] "Means for issuing a security alert" refers to a method or device for issuing an alert when it is determined that the user's safety may be threatened based on the results of emotion analysis.
[0231] "Means for coordinating with security devices" refers to a method or device for coordinating with other security devices (such as surveillance cameras or intrusion detection systems) when issuing an alarm.
[0232] The "means for receiving a practice request" refers to a method or device for receiving a request for practicing a daily situation desired by a user.
[0233] "Means for generating a practice scenario" refers to a method or device for automatically generating a scenario according to the practice desired by the user.
[0234] The "means for conducting interactive practice" refers to a method or device for allowing a user and a system to have a dialogue based on a generated practice scenario.
[0235] The "means for generating feedback" refers to a method or device for analyzing the results of a user's practice and providing advice or suggestions for improvement based on the results.
[0236] A "prompt" is text data that is input into a generative AI model, providing information for the AI's response or generated message.
[0237] Components
[0238] 1. User Device:
[0239] A device that allows a user to input voice or text. It includes smartphones, tablets, personal computers, head-mounted displays, etc. The user terminal is responsible for sending input data to a server and receiving messages from the server.
[0240] 2. Server:
[0241] It receives input data, analyzes emotions, generates counseling messages, and determines the level of urgency. It also uses generative AI models to generate practice scenarios and collaborate with other security devices.
[0242] Program processing
[0243] Handling User Input
[0244] When a user inputs data by voice or text, the user device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The user device converts this data into text data and sends it to the server. The server receives this data and performs emotion analysis.
[0245] Emotion analysis
[0246] The server uses an emotion analysis model to classify the user's emotional state as either "positive," "negative," or "neutral." For example, if a user enters "I'm feeling stressed," the server will determine this as "negative."
[0247] Counseling message generation
[0248] The server generates appropriate counseling messages based on the analysis results, such as "Take a deep breath to relax," using a generative AI model, and sends the messages to the user's device.
[0249] Urgency assessment and response request
[0250] The server continuously monitors the emotion data and determines the level of urgency. If the level of urgency is deemed high, it will request a response from an expert. For example, if the user continuously expresses sadness, the server will send an alert to an expert. It will also issue a security alarm and connect to security devices.
[0251] Providing practice scenarios
[0252] When a user wants to practice an everyday situation, for example, they request, "I would like to practice an interview." The server receives this request and uses a generative AI model to generate an appropriate interview scenario. The generated scenario is sent to the user's device, where the user practices interactively. The server analyzes the practice results, generates feedback based on the prompt sentence, and sends it to the user's device.
[0253] Specific examples
[0254] 1. Sentiment analysis and counseling message provision:
[0255] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this speech into text and sends it to the server. The server determines that the user is "feeling stressed," and generates a message such as "Take a deep breath to relax," and sends it to the device.
[0256] 2. Urgency assessment and expert intervention:
[0257] The user repeatedly types "I feel heavy." The server analyzes this and determines that the situation is urgent. It sends an alert to an expert, who then contacts the user and provides counseling.
[0258] 3. Providing practice scenarios and feedback:
[0259] The user requests, "I want to practice for an interview." The server generates an appropriate scenario and sends it to the terminal. The user practices interactively, and the server analyzes the results and sends feedback. An example of a prompt sentence is, "How do you relieve nervousness during an interview?"
[0260] In this way, a system is realized in which the user terminal and server work together to support the user's mental health while also responding immediately to security risks.
[0261] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0262] Step 1:
[0263] The user provides input via voice or text.
[0264] Input: User voice or text
[0265] Output: Audio data or text data
[0266] How it works: The user uses a device such as a smartphone or head-mounted display (HMD) to input voice commands, for example, "I'm feeling stressed because things haven't been going well at work lately."
[0267] Step 2:
[0268] The device converts the voice input into text.
[0269] Input: Audio data
[0270] Output: Text data
[0271] How it works: The device's microphone records the user's voice and converts it into text using speech recognition software (e.g., Google Speech Recognition).
[0272] Step 3:
[0273] The terminal transmits the text data to the server.
[0274] Input: Text data
[0275] Output: Data sent to the server
[0276] How it works: Your device sends the converted text data to a server over your internet connection.
[0277] Step 4:
[0278] The server receives the text data and performs sentiment analysis.
[0279] Input: Text data
[0280] Output: Emotion judgment result (positive, negative, neutral, etc.)
[0281] How it works: The server inputs the received text data into a sentiment analysis model (e.g., BERT), which then classifies the user's sentiment. For example, "I'm feeling stressed" is classified as "negative."
[0282] Step 5:
[0283] The server generates a counseling message based on the emotion determination result.
[0284] Input: Emotion determination result
[0285] Output: Counseling message
[0286] How it works: The server uses the emotion determination results to input prompts into the generative AI model, generating an appropriate counseling message. For example, a message like "Take a deep breath and you'll be able to relax" is generated.
[0287] Step 6:
[0288] The server sends a counseling message to the user terminal.
[0289] Input: Counseling message
[0290] Output: Transmitted data
[0291] Operation: The server sends the generated counseling message to the user terminal.
[0292] Step 7:
[0293] The user terminal displays a counseling message to the user.
[0294] Input: Counseling message
[0295] Output: Display message
[0296] Operation: The user terminal displays the counseling message received on the screen. For example, the user sees a message on the terminal screen saying, "Take a deep breath and you'll be able to relax."
[0297] Step 8:
[0298] The server continuously monitors the emotional data and determines the level of urgency.
[0299] Input: Continuous emotion data
[0300] Output: Urgency (low, medium, high)
[0301] How it works: The server monitors the emotional data continuously sent by the user and determines the urgency based on certain criteria.
[0302] Step 9:
[0303] If the emergency is deemed high, the server will request a specialist to respond.
[0304] Input: Urgency (High)
[0305] Output: Alert to experts
[0306] How it works: When the server detects a high level of urgency, it sends an alert to registered experts and requests them to take action.
[0307] Step 10:
[0308] Additionally, the server will issue security alerts as needed and work in conjunction with security devices.
[0309] Input: High-urgency emotion data
[0310] Output: Security alerts and data sent to linked devices
[0311] How it works: The server issues security alarms and works in conjunction with security devices such as surveillance cameras and intrusion detection systems to provide comprehensive safety responses.
[0312] Step 11:
[0313] When a user submits a practice request, the server generates a practice scenario for an everyday situation.
[0314] Input: User's practice request
[0315] Output: Practice scenario
[0316] How it works: When a user requests, "I want to practice for an interview," the server uses the generative AI model to generate an appropriate practice scenario and sends it to the user's device.
[0317] Step 12:
[0318] The server analyzes the practice results, generates feedback, and sends it to the user's terminal.
[0319] Input: Practice result data
[0320] Output: Feedback message
[0321] How it works: The user performs the exercise and sends the results to the server. The server then uses the generative AI model to generate a prompt and sends feedback to the user's device. For example, the server might provide feedback such as, "It would be better if you relaxed your gaze more."
[0322] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0323] This system analyzes a user's voice or text input, determines the user's emotions, and provides appropriate counseling messages. Furthermore, by combining it with an emotion engine, it can recognize subtle changes in the user's emotional state and provide more accurate support.
[0324] Components
[0325] 1. User Device:
[0326] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[0327] 2. Server:
[0328] It receives input data and performs processes such as emotion analysis, generating counseling messages, and determining the level of urgency. It also uses an emotion engine to learn the user's emotional patterns and reflect changes in emotions in real time.
[0329] 3. Experts:
[0330] Based on requests from the server, appropriate advice and treatment will be provided to users, including psychologists, psychiatrists, and counselors.
[0331] 4. Emotion Engine:
[0332] It recognizes emotions by analyzing the user's voice, text input, and non-verbal elements (facial expressions, tone of voice, etc.). It learns emotional patterns and updates them in real time to detect subtle changes in the user's emotional state.
[0333] Program processing
[0334] Handling User Input
[0335] The user inputs data by voice or text, and the device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this data into text data and sends it to the server. The server receives this data and analyzes it using an emotion engine.
[0336] Emotion engine processing
[0337] The emotion engine in the server recognizes the user's emotions based on the input data. For example, it not only determines that the user is "feeling stressed," but also detects that the user is "very stressed" based on facial expressions and tone of voice.
[0338] Counseling message generation
[0339] The server generates a counseling message based on the analysis results of the emotion engine. For example, if it determines that the user is feeling extremely stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[0340] Urgency assessment and response request
[0341] The server determines the level of urgency based on the data obtained from the emotion engine. For example, if a user repeatedly expresses very strong feelings of depression, the server will request a response from a specialist. The specialist will receive the alert and contact the user to provide counseling or appropriate treatment.
[0342] Providing practice scenarios
[0343] When a user wants to practice an everyday situation, for example, they request, "I want to practice an interview." The server receives this request and generates an appropriate interview scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on the scenario. The emotion engine monitors and analyzes the user's reactions and provides feedback.
[0344] Matching and communication support
[0345] When a user wishes to communicate with other users, they can enter, for example, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users and sends the results to the device. The users then begin a conversation, and the server monitors the exchange, requesting expert intervention if necessary.
[0346] Specific examples
[0347] 1. Emotion analysis using an emotion engine and provision of counseling messages:
[0348] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server uses its emotion engine to determine that the user is under "very high stress," and generates a message such as, "Take a deep breath and you'll be able to relax. Try it out," which is sent to the device. The device then displays this message to the user.
[0349] 2. Urgency assessment and expert intervention:
[0350] A user repeatedly types, "I'm feeling very depressed." The server analyzes the situation using an emotion engine, and if it determines that the situation is urgent, it sends an alert to a specialist. The specialist receives the alert and contacts the user to provide counseling or treatment.
[0351] 3. Providing practice scenarios and feedback:
[0352] The user requests, "I want to practice for an interview." The server uses an emotion engine to analyze the user's emotional state and generate an appropriate scenario. The scenario is sent to the device, and the user practices. The server analyzes the practice results, generates feedback, and sends it to the device.
[0353] 4. Matching and communication support:
[0354] The user types, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users then begin to converse, and the system monitors the content and requests intervention from experts if necessary.
[0355] As described above, by combining the emotion engine, the system of the present invention can monitor the user's mental health state with higher accuracy and provide appropriate counseling and support for rehabilitation into society.
[0356] The processing flow will be explained below.
[0357] AI that listens to your heart
[0358] User Input and Sentiment Analysis
[0359] Step 1:
[0360] The user speaks or texts into the device, saying, "I've been feeling stressed lately because things haven't been going well at work."
[0361] Step 2:
[0362] The terminal converts the voice input into text data and transmits the text data to the server.
[0363] Step 3:
[0364] The server receives the text data and requests the emotion engine to analyze it.
[0365] Step 4:
[0366] The emotion engine determines from the user's text that they are "feeling stressed."
[0367] Step 5:
[0368] The emotion engine also analyzes non-verbal data such as facial expressions and tone of voice to assess the intensity of stress (if determined to be "very high stress").
[0369] Counseling message generation
[0370] Step 6:
[0371] The server generates a counseling message based on the analysis results of the emotion engine. For example, it generates a message such as, "Take a deep breath and you'll feel more relaxed. Try it out."
[0372] Step 7:
[0373] The server sends the generated counseling message to the terminal.
[0374] Step 8:
[0375] The terminal displays a counseling message to the user.
[0376] AI in harmony with specialists
[0377] Urgency assessment and response request
[0378] Step 1:
[0379] The server continuously monitors the user's daily emotional data.
[0380] Step 2:
[0381] The server determines the level of urgency based on the data obtained from the emotion engine. For example, if the user repeatedly expresses "feeling very depressed," it will determine that the level of urgency is high.
[0382] Step 3:
[0383] If the server determines that the situation is urgent, it will send an alert to an expert.
[0384] Step 4:
[0385] A specialist receives the alert from the server and contacts the user to provide counseling or treatment.
[0386] AI provides a platform for social reintegration
[0387] Providing practice scenarios
[0388] Step 1:
[0389] The user inputs a request to the terminal saying, "I want to practice for an interview."
[0390] Step 2:
[0391] The server receives the user's request and generates an interview scenario.
[0392] Step 3:
[0393] The emotion engine analyzes the user's emotional state and reflects the emotional data in the scenario.
[0394] Step 4:
[0395] The server sends the generated scenario to the terminal.
[0396] Step 5:
[0397] The terminal displays a scenario, and the user practices the interview in an interactive format.
[0398] Providing Feedback
[0399] Step 6:
[0400] When the user completes the exercise, the terminal transmits the exercise results to the server.
[0401] Step 7:
[0402] The server uses an emotion engine to analyze the practice results and generate a feedback message, such as "You did very well. You could do better if you answered with more confidence."
[0403] Step 8:
[0404] The server sends a feedback message to the terminal, which displays it to the user.
[0405] AI that provides an empathetic community
[0406] Matching with user requests
[0407] Step 1:
[0408] The user inputs their request into the terminal, saying, "I would like to talk to someone who is also aiming to return to society."
[0409] Step 2:
[0410] The server receives the user's request and analyzes it using the emotion engine.
[0411] Step 3:
[0412] The server searches and selects users in the same position from the database and works with the emotion engine to make appropriate matches.
[0413] Step 4:
[0414] The server sends the matching results to the device.
[0415] Step 5:
[0416] The terminal notifies the user of the matching result, and the user starts communication with the matched person.
[0417] Communications monitoring and support
[0418] Step 6:
[0419] The server monitors interactions within the community using an emotion engine.
[0420] Step 7:
[0421] If necessary, the server calls for expert intervention.
[0422] Through these steps, this system can comprehensively support users in maintaining their mental health and reintegrating into society. By combining it with an emotion engine, it is possible to detect even subtle changes in the user's emotional state, enabling more accurate counseling and support to be provided.
[0423] Example 2
[0424] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0425] In modern society, increasing stress and anxiety have created a demand for psychological support. However, many users have limited opportunities to receive appropriate counseling immediately. Furthermore, conventional counseling methods often fail to capture subtle changes in a user's emotions in real time, making it difficult to provide immediate, appropriate support. While there are counseling support systems that use voice or text input, few of them offer sufficient emotional analysis accuracy, real-time performance, urgency assessment, or expert intervention. This has created a demand for systems that can effectively support users' mental health.
[0426] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0427] In this invention, the server includes means for receiving a user's voice or text input, means for converting the input into text data and transmitting it to the server, means for analyzing the input to determine the user's emotion, means for generating a counseling message based on the emotion, and means for providing the counseling message to the user. This makes it possible to analyze the user's voice or text input in real time, quickly capture subtle changes in the user's emotional state, and generate and provide an appropriate counseling message. Furthermore, in cases of high urgency, it is possible to request a response from an expert, thereby enabling rapid expert intervention.
[0428] "User" refers to an individual or organization that uses the system.
[0429] "Voice input" refers to data that is converted into digital form from what a user says.
[0430] "Text input" refers to data in which a user inputs textual information using a keyboard or other input device.
[0431] "Device" means a device used by a user to input voice or text, including, for example, a smartphone, tablet, or computer.
[0432] A "server" refers to a high-performance computer system that receives and analyzes data sent by users.
[0433] "Emotion engine" refers to a system that analyzes a user's emotions from voice, text, and non-verbal data.
[0434] "Counseling message" refers to a message containing advice or support for the user, generated based on the analysis results of the emotion engine.
[0435] "Urgency" refers to an index that evaluates the level of urgency of a user's emotional state or health condition.
[0436] "Expert" refers to a person or organization qualified to provide appropriate advice or treatment for a user's mental health, such as a psychologist, psychiatrist, or counselor.
[0437] "Response request" refers to an action in which the system requests an expert to respond to the user's condition.
[0438] "Practice scenario" refers to an interactive scenario provided for a user to simulate a particular situation.
[0439] "Feedback" refers to evaluations and advice provided based on the results of the user's scenario practice.
[0440] "Matching" refers to the process of appropriately pairing users based on their purpose and status.
[0441] The present invention provides a system for supporting the mental health of a user by analyzing the user's voice or text input, determining the user's emotions, and providing appropriate counseling messages. Specific implementation methods are described in detail below.
[0442] Components
[0443] User terminal
[0444] A user terminal is a device for inputting voice or text. Terminals include smartphones, tablets, and PCs. The user terminal transmits the input data to a server and receives messages from the server.
[0445] server
[0446] The server has the central function of receiving and analyzing data sent by users. The received data is input into the emotion engine, which analyzes the user's emotions. Based on the analysis results, it generates an appropriate counseling message and sends it to the user's device. It also determines the level of urgency, requests expert assistance, provides practice scenarios and feedback, and matches users and supports communication.
[0447] Emotion Engine
[0448] The emotion engine is a system that analyzes the user's emotions using non-verbal data such as voice, text, facial expression analysis, and tone of voice. Based on the analysis results, the emotion engine detects subtle changes in the user's emotional state and reflects them in real time.
[0449] Hardware and software used
[0450] Speech Recognition Software: Uses the Google Speech-to-Text API to convert user voice input into text data.
[0451] Sentiment analysis model: We use a sentiment analysis model using the Hugging Face Transformers library.
[0452] Generative AI model: OpenAI's GPT-3 is used to generate counseling messages and practice scenarios.
[0453] Specific examples
[0454] Emotion analysis using an emotion engine and provision of counseling messages
[0455] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server uses its emotion engine to determine that the user is feeling "very stressed," and generates a message such as, "Take a deep breath and you'll be able to relax. Try it out," and sends it to the device. The device then displays this message to the user.
[0456] Urgency assessment and expert intervention
[0457] A user repeatedly types, "I'm feeling very depressed." The server analyzes the situation using an emotion engine, and if it determines that the situation is urgent, it sends an alert to a specialist. The specialist receives the alert and contacts the user to provide counseling or treatment.
[0458] Providing practice scenarios and feedback
[0459] The user requests, "I want to practice for an interview." The server uses an emotion engine to analyze the user's emotional state and generate an appropriate scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on that scenario. The server analyzes the practice results, generates feedback, and sends it to the device.
[0460] Matching and communication support
[0461] The user types, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users then begin to converse, and the server monitors the exchange and, if necessary, requests intervention from experts.
[0462] Prompt Sentence Examples
[0463] 1. If you want to use sentiment analysis with the sentiment engine:
[0464] "If a user says, 'I'm feeling stressed because things haven't been going well at work lately,' how does the emotion engine respond?"
[0465] 2. To request a priority assessment and expert intervention:
[0466] "If a user expresses very depressed feelings consecutively, how does the server determine the urgency and notify an expert?"
[0467] 3. If you would like to provide a practice scenario:
[0468] "If a user wants to practice an interview, how does the system generate scenarios and provide feedback?"
[0469] 4. If you wish to match and communicate with other users:
[0470] “If a user inputs that they want to talk to other users with similar goals of reintegration, how does the system match them with the right users and support their interactions?”
[0471] As described above, the present invention can support users' mental health with high accuracy through real-time emotion analysis using an emotion engine and the provision of counseling messages and practice scenarios using a generative AI model.
[0472] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0473] Step 1:
[0474] The user inputs data by voice or text. An example of input data is the text "I'm feeling stressed lately because things aren't going well at work." In the case of voice input, the device uses speech recognition software (e.g., Google Speech-to-Text API) to convert the voice into text data. The input for this step is the user's voice or text, and the output is text data.
[0475] Step 2:
[0476] The device sends text data to the server. Specifically, the device sends the text data entered to the server via an HTTP request. For example, text data is sent in the form of {"user_input": "Recently, things haven't been going well at work and I'm feeling stressed"}. The input of this step is text data, and the output is the data sent to the server.
[0477] Step 3:
[0478] The server passes the received input data to the emotion engine. The emotion engine uses an emotion analysis model using Hugging Face's Transformers library to analyze the text data and determine the user's emotion. For example, it determines "very strong stress." The input of this step is the received text data, and the output is the emotion analysis result.
[0479] Step 4:
[0480] The server generates a counseling message using a generative AI model (for example, OpenAI's GPT-3) based on the analysis results of the emotion engine. The prompt text is "The user's stress level is very high. Please provide some advice on how to relax." The AI model generates a message such as "Taking deep breaths can help you relax. Let's try it." The input for this step is the emotion analysis result, and the output is a counseling message.
[0481] Step 5:
[0482] The server sends the generated counseling message to the user's device. Specifically, it sends the following message using an HTTP response: {"counseling_message": "Try taking a deep breath to relax."} The device then displays the received message to the user. The input to this step is the counseling message, and the output is the message displayed on the user's device.
[0483] Step 6:
[0484] The server determines the urgency level based on the analysis results of the emotion engine. For example, if the user repeatedly expresses "very depressed feelings," the server determines the urgency level as "high." The input of this step is the emotion analysis result, and the output is the urgency level determination result.
[0485] Step 7:
[0486] If the urgency is determined to be "high," the server requests a response from an expert. An alert is sent to the expert, for example, a notification saying, "User A's urgency is high. Please respond immediately." The expert receives the alert and contacts the user to provide counseling or appropriate treatment. The input to this step is the urgency determination result, and the output is an alert notification to the expert.
[0487] Step 8:
[0488] When a user wishes to practice an everyday situation, they input a request, for example, "I would like to practice an interview." The server receives the request, analyzes the user's emotional state using an emotion engine, and generates an appropriate practice scenario. The generated scenario is sent to the device, and the user practices interactively based on that scenario. The server analyzes the practice results, generates feedback, and sends it to the device. The input to this step is the user's practice request, and the output is the practice scenario and feedback.
[0489] Step 9:
[0490] When a user wishes to communicate with other users, they input, for example, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users begin a conversation, and the server monitors the exchange, requesting intervention from experts if necessary. The input to this step is the user's communication desire, and the output is the matching results and conversation monitoring.
[0491] (Application example 2)
[0492] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0493] In recent years, with the spread of electronic payment services, users have been experiencing increasing stress and anxiety during transactions and payments. This mental burden not only worsens the user experience, but may also lead to transaction failures and increased security risks. Therefore, there is a need for a system that can monitor users' emotions in real time during electronic payment services and provide appropriate support to improve the user experience.
[0494] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a user's voice or text input, means for analyzing the input to determine the user's emotion, means for generating a counseling message based on the emotion, means for providing the counseling message to the user, and means for monitoring the user's emotion when conducting a transaction or payment and taking measures to reduce stress. In this way, by understanding the user's emotional state in real time and providing a counseling message as needed, stress during a transaction or payment can be reduced, allowing the user to use the service with peace of mind.
[0495] The "means for receiving user voice or text input" is an interface for acquiring voice or text data uttered by the user and processing it within the system.
[0496] The "means for analyzing the input and determining the user's emotions" refers to an algorithm or engine that analyzes the received voice or text data and determines the user's emotional state (e.g., stress, joy, surprise, etc.) based on that analysis.
[0497] The "means for generating a counseling message based on the emotion" is a processing unit for automatically generating an appropriate message of advice or comfort according to the determined emotional state of the user.
[0498] The "means for providing the counseling message to the user" is a module for displaying, playing, or transmitting the generated counseling message to the user's device.
[0499] "Means to monitor users' emotions when making transactions or payments and take measures to reduce stress" refers to a function that monitors users' emotional state in real time during electronic payments or transactions and provides guidelines and support messages to reduce stress and anxiety.
[0500] The "means for determining the degree of urgency based on the emotion" is a system for assessing the seriousness of the user's emotional state and determining whether an emergency response is required, if necessary.
[0501] The "means for requesting a response from an expert when the urgency is high" is a module that sends an alert to an expert such as a counselor or a specialist doctor when a high urgency is determined, urging them to intervene quickly.
[0502] The "means for receiving a user's practice request for an everyday situation" is an interface for receiving a request when a user wishes to practice a specific scenario (for example, an interview, a presentation, etc.).
[0503] The "means for generating a practice scenario based on the request" is a module for automatically creating a specific practice scenario in response to a request from a user.
[0504] The "means for interactively practicing with the user based on the practice scenario" is an interface that allows the user to practice while interacting with the system based on the generated scenario.
[0505] The "means for generating feedback based on the practice results" is a system that analyzes the user's practice results and automatically creates and provides feedback such as areas for improvement and results.
[0506] This invention is a system that analyzes a user's voice or text input, determines the user's emotions, and provides appropriate counseling messages. Furthermore, by using an emotion engine, it is possible to monitor the user's emotions in real time when conducting transactions or payments, and take necessary measures. Such a system aims to reduce the user's mental burden and provide a better user experience.
[0507] System Configuration
[0508] The system includes the following components:
[0509] 1. User Device
[0510] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[0511] 2. Server
[0512] It receives input data and performs processes such as emotion analysis, generating counseling messages, monitoring emotions during transactions and payments, and determining urgency. It also uses an emotion engine to learn the user's emotional patterns and reflect emotional changes in real time.
[0513] 3. Emotion Engine
[0514] It recognizes emotions by analyzing the user's voice, text input, and non-verbal elements (facial expressions, tone of voice, etc.). It learns emotional patterns and updates them in real time to detect subtle changes in the user's emotional state.
[0515] System Operation
[0516] In the present invention, the system operates using the following means.
[0517] 1. Handling User Input
[0518] The user inputs data by voice or text, and the device sends the data to the server. For example, the user says, "This transaction is very stressful." The device converts this data into text and sends it to the server.
[0519] 2. Emotion Engine Processing
[0520] The emotion engine in the server recognizes the user's emotions based on the input data. For example, it not only determines that the user is "feeling stressed," but also detects that the user is "very stressed" based on facial expressions and tone of voice.
[0521] 3. Generating counseling messages
[0522] The server generates a counseling message based on the analysis results of the emotion engine. For example, if it determines that the user is feeling extremely stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[0523] 4. Emotion monitoring and countermeasures during transactions and payments
[0524] The server monitors the user's emotions in real time when trading or making payments, and if it determines that the user is feeling stressed, it takes measures to reduce stress. For example, if a user enters "This transaction is very stressful" while trading, the server analyzes this and generates an advice message such as "Try taking deep breaths to relax."
[0525] Specific examples and examples of prompts for generative AI models
[0526] Example 1:
[0527] If the user enters "This transaction is very stressful," the server processes it as follows:
[0528] Emotion analysis result: NEGATIVE
[0529] Score: 0.95
[0530] Counseling message: "You seem to be stressed. Take a deep breath and try to relax a bit."
[0531] Example prompt sentence:
[0532] Create a program that analyzes a user's voice or text input, checks whether the user is feeling anxious or stressed, and provides an appropriate counseling message. Use Transformers as the emotion analysis model, and display a stress relief message if the input indicates a negative emotion.
[0533] The above is a specific embodiment for carrying out this invention. By combining this system with an emotion engine, it is possible to monitor the user's mental health with greater precision and provide appropriate counseling and transaction / payment support.
[0534] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0535] Step 1:
[0536] The user provides voice or text input. For example, the user might say, "This transaction is very stressful." This input is captured by the terminal and converted into text data.
[0537] Step 2:
[0538] The device sends the converted text data to the server, which receives the data and prepares it for sentiment analysis. The input is the user's voice or text data, and the output is the text data sent to the server.
[0539] Step 3:
[0540] The server uses an emotion engine to analyze the incoming data. Specifically, it applies a generative AI model such as Transformers to determine the user's emotion. The input is text data, and the output is an emotion (e.g., NEGATIVE) and its score (e.g., 0.95).
[0541] Step 4:
[0542] The server runs an algorithm that generates a counseling message based on the analysis results. Specifically, if the emotion is "NEGATIVE" and the score is high, a counseling message such as "You seem to be feeling stressed. Take a deep breath and try to relax a bit" is created. The input is the emotion analysis result and score, and the output is the generated counseling message.
[0543] Step 5:
[0544] The server sends the generated counseling message to the terminal. The terminal receives this message and displays or plays it aloud to the user. The input is the counseling message sent from the server, and the output is the message provided in a format that the user can view.
[0545] Step 6:
[0546] The server monitors users' emotions in real time and continuously provides stress-reducing measures during transactions and settlements. For example, if a transaction is prolonged or if a particular frustration persists, it generates additional messages such as "Take a short break to refresh yourself." The input is continuous emotion analysis data, and the output is support messages provided at any time.
[0547] Step 7:
[0548] The system determines the urgency of the emotion, and if the server deems it necessary, sends an alert to an expert. For example, if a user repeatedly expresses very strong, depressed emotions, the system will contact an expert and ask for their response. The input is the emotion analysis result and urgency assessment data, and the output is an alert to the expert and a request for their response.
[0549] Step 8:
[0550] Based on the user's request, the server generates a practice scenario for an everyday situation. If the user requests "I want to practice an interview," the server generates an appropriate scenario and sends it to the terminal. The input is the user's practice request, and the output is the generated practice scenario.
[0551] Step 9:
[0552] The user performs interactive practice based on the generated scenario, which involves the system analyzing the user's responses and providing appropriate feedback. The input is the user's interaction data, and the output is the feedback provided by the system.
[0553] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0554] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0555] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0556] [Second embodiment]
[0557] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0558] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0559] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0560] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0561] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0562] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0563] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0564] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0565] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0566] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0567] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0568] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0569] The present invention is a system that analyzes a user's voice or text input, determines their emotions, and provides appropriate counseling messages. This system is made up of the following components:
[0570] Components
[0571] 1. User Device:
[0572] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[0573] 2. Server:
[0574] It receives input data, analyzes emotions, generates counseling messages, and determines the level of urgency. It also generates practice scenarios, matches users with other users, and monitors communication content.
[0575] 3. Experts:
[0576] Based on requests from the server, appropriate advice and treatment will be provided to users, including psychologists, psychiatrists, and counselors.
[0577] Program processing
[0578] Handling User Input
[0579] The user inputs data by voice or text, and the device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server receives this data, performs emotional analysis, and determines that the user is "feeling stressed."
[0580] Counseling message generation
[0581] The server generates a counseling message based on the results of emotion analysis. For example, if it determines that the user is feeling stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[0582] Urgency assessment and response request
[0583] The server continuously monitors the emotion data and determines the level of urgency. If the level is deemed high, it requests a specialist to respond. For example, if a user continuously expresses depressed feelings, the server sends an alert to the specialist. The specialist receives the alert, contacts the user, and provides appropriate counseling or treatment.
[0584] Providing practice scenarios
[0585] When a user wishes to practice an everyday situation, for example, they make a request such as "I would like to practice an interview." The server receives this request and generates an appropriate interview scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on the scenario. The server analyzes the practice results, generates feedback, sends it to the device, and provides it to the user.
[0586] Matching and communication support
[0587] When a user wishes to communicate with other users, they can enter, for example, "I want to talk to someone who is also trying to reintegrate into society." The server searches and selects users in the same position from its database and performs matching. The matching results are sent to the user's device and notified. The user then begins a conversation with the matched person, and the system monitors the conversation. If necessary, the server requests expert intervention.
[0588] Specific examples
[0589] 1. Sentiment analysis and counseling message provision:
[0590] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this speech into text and sends it to the server. The server determines that the user is "feeling stressed," and generates a message such as "Take a deep breath to relax," which is sent to the device. The device then displays this message to the user.
[0591] 2. Urgency assessment and expert intervention:
[0592] The user repeatedly types "I feel heavy." The server analyzes this and determines that the situation is urgent. It sends an alert to an expert, who then contacts the user and provides counseling.
[0593] 3. Providing practice scenarios and feedback:
[0594] The user requests, "I want to practice an interview." The server generates an appropriate scenario and sends it to the device. The user practices interactively, and the server analyzes the results and sends feedback.
[0595] 4. Matching and communication support:
[0596] The user types, "I want to talk to people who are also trying to reintegrate into society." The server matches suitable users and notifies the device. The users begin a conversation, and the system monitors the content. If necessary, it requests intervention from an expert.
[0597] As described above, the system of the present invention can monitor the mental health state of the user and provide appropriate counseling and support for rehabilitation into society.
[0598] The processing flow will be explained below.
[0599] AI that listens to your heart
[0600] User Input and Sentiment Analysis
[0601] Step 1:
[0602] The user speaks or texts into the device, saying, "I've been feeling stressed lately because things haven't been going well at work."
[0603] Step 2:
[0604] The terminal converts the voice input into text data and transmits the text data to the server.
[0605] Step 3:
[0606] A server receives the text data and runs a sentiment analysis model (e.g., using natural language processing).
[0607] Step 4:
[0608] The server determines from the user's text that he or she is "feeling stressed."
[0609] Counseling message generation
[0610] Step 5:
[0611] The server generates a counseling message for stress reduction (e.g., "Take a deep breath and you'll feel more relaxed").
[0612] Step 6:
[0613] The server sends the generated counseling message to the terminal.
[0614] Step 7:
[0615] The terminal displays a counseling message to the user.
[0616] AI in harmony with specialists
[0617] Urgency assessment and response request
[0618] Step 1:
[0619] The server monitors the user's daily emotional data.
[0620] Step 2:
[0621] The server analyzes the emotional data and determines whether the person is experiencing a continuous state of depression.
[0622] Step 3:
[0623] The server determines the level of urgency and, if deemed high, sends an alert to an expert.
[0624] Step 4:
[0625] A specialist receives the alert from the server and contacts the user to provide counseling or treatment.
[0626] AI provides a platform for social reintegration
[0627] Providing practice scenarios
[0628] Step 1:
[0629] The user inputs a request to the terminal saying, "I want to practice for an interview."
[0630] Step 2:
[0631] The server receives the user's request and generates an interview scenario.
[0632] Step 3:
[0633] The server sends the generated scenario to the terminal.
[0634] Step 4:
[0635] The terminal displays a scenario, and the user practices the interview in an interactive format.
[0636] Step 5:
[0637] When the user completes the exercise, the terminal transmits the exercise results to the server.
[0638] Providing Feedback
[0639] Step 6:
[0640] The server analyzes the practice results and generates a feedback message.
[0641] Step 7:
[0642] The server sends a feedback message to the terminal, which displays it to the user.
[0643] AI that provides an empathetic community
[0644] Matching with user requests
[0645] Step 1:
[0646] The user inputs their request into the terminal, saying, "I would like to talk to someone who is also aiming to return to society."
[0647] Step 2:
[0648] The server analyzes users in the same position from the database and matches them with the appropriate users.
[0649] Step 3:
[0650] The server sends the matching results to the device.
[0651] Step 4:
[0652] The user begins communicating with the matched person (e.g., "I see you're in the same situation, ○○. Let's both do our best.").
[0653] Communications monitoring and support
[0654] Step 5:
[0655] The server monitors interactions within the community.
[0656] Step 6:
[0657] If necessary, the server calls for expert intervention.
[0658] Through the above steps, this system can comprehensively support the user's mental health maintenance and social reintegration.
[0659] Example 1
[0660] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0661] Conventional counseling systems struggle to quickly and accurately analyze users' emotions and provide appropriate counseling messages. Furthermore, delays in assessing the level of urgency and providing appropriate expert intervention can lead to a deterioration in the user's mental health. Furthermore, they lack sufficient practice in everyday situations and support for communication between users, preventing them from effectively supporting users' reintegration into society.
[0662] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0663] In this invention, the server includes means for receiving a user's voice or text input, means for converting the input into text data, means for transmitting the text data to the server, means for analyzing the text data to determine the user's emotion, means for generating a counseling message using a generative AI model based on the result of the emotion analysis, means for transmitting the counseling message to a user terminal, and means for displaying the counseling message to the user on the user terminal. This makes it possible to quickly and accurately analyze the user's emotion and effectively support appropriate counseling and social reintegration.
[0664] "Voice or text input" refers to the form of speech or written information that a user uses to communicate their feelings or requests to a system.
[0665] "Text data" refers to data obtained by converting voice input into text format, or text information entered directly by the user.
[0666] "Server" refers to the central computing system that receives and analyzes voice or text data, generates counseling messages, determines urgency, etc.
[0667] "Sentiment analysis" refers to the process of analyzing a user's text data to determine the user's emotional state.
[0668] A "generative AI model" refers to an artificial intelligence algorithm, primarily a natural language generation model, that generates appropriate messages based on the results of user sentiment analysis.
[0669] "Counseling message" refers to a message of advice or encouragement provided to a user that is generated based on the results of an analysis of the user's emotions.
[0670] "Urgency" refers to an index that evaluates the urgency of the user's emotional state and indicates whether a prompt response is required.
[0671] "Professional" refers to a psychologist, psychiatrist, counselor, or other person qualified to assist users with their mental health.
[0672] A "practice scenario" refers to a simulation scenario generated by the server when a user wishes to practice a particular situation (such as an interview).
[0673] "Feedback" refers to evaluations and advice provided to users based on the results of their practice.
[0674] "Matching" refers to the process by which a server selects and connects a suitable partner when a user wishes to communicate with another user.
[0675] "User terminal" refers to a device (smartphone, tablet, PC, etc.) used to input voice or text and receive and display counseling messages and feedback from the server.
[0676] The "HTTPS protocol" refers to a communication protocol for secure data communication over the Internet.
[0677] "Natural language processing tool" refers to a software tool (e.g., spaCy, NLTK, etc.) used to analyze a user's text data.
[0678] The present invention is a system that analyzes a user's voice or text input, determines their emotions, and provides appropriate counseling messages. This can support the user's mental health and effectively assist them in their social reintegration. A specific embodiment of this system is described below.
[0679] System Configuration
[0680] This system consists of user devices, servers, and expert components. Specific hardware components for user devices include smartphones, tablets, and PCs. Servers are high-performance servers (e.g., cloud virtual machines) with various software installed.
[0681] User terminal: A device that receives voice or text input, converts it into text data, and sends it to a server.
[0682] Server: This is a critical computing system that performs sentiment analysis, generates counseling messages, and determines urgency. Specifically, it uses natural language processing tools (e.g., spaCy, NLTK) and machine learning libraries (e.g., TensorFlow, PyTorch).
[0683] Experts: Psychologists, psychiatrists, counselors, or other qualified individuals available to support users' mental health.
[0684] Example of operation
[0685] Processing voice or text input
[0686] The user inputs data by voice or text. The user's device receives this data, and in the case of voice input, it converts it into text data using the Google Speech-to-Text API. For example, if the user says, "Recently, things haven't been going well at work and I'm feeling stressed," the content is converted into text. This text data is sent to the server using the HTTPS protocol.
[0687] Emotion analysis
[0688] The server analyzes the received text data and determines the user's emotions. Using a natural language processing tool (e.g., spaCy), it determines that the user is "feeling stressed." The emotion analysis algorithm is trained using a machine learning library (e.g., TensorFlow).
[0689] Counseling message generation
[0690] Based on the results of the sentiment analysis, a generative AI model (e.g., GPT-3) is used to generate an appropriate counseling message. For example, a message such as "Take a deep breath and you'll feel more relaxed. Try it out." This message is sent to the user's device and immediately displayed to the user.
[0691] Urgency assessment and expert intervention
[0692] The server continuously monitors the user's emotional data and determines the level of urgency. If the user repeatedly inputs "I feel heavy," the level of urgency is determined to be high. If the level of urgency is high, the server sends an alert to an expert, who then contacts the user and provides counseling.
[0693] Providing practice scenarios
[0694] When a user requests "I want to practice for an interview," the server generates an appropriate practice scenario. The scenario contains specific questions, and the contents of the scenario are sent to the user's terminal. The user practices interactively, and the server analyzes the results and generates feedback.
[0695] Prompt Sentence Examples
[0696] Emotion Analysis Prompt: Enter "I've been feeling stressed lately because things haven't been going well at work."
[0697] Urgency determination prompt: Enter "I feel heavy" repeatedly.
[0698] Practice Scenario Prompt: Type "I want to practice interviewing."
[0699] Matching prompt: Enter "I'd like to talk to someone who is also trying to reintegrate into society."
[0700] As described above, the system of the present invention is realized using advanced technology to support the user's mental health and effectively assist in their reintegration into society.
[0701] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0702] Step 1:
[0703] The user inputs data by voice or text. For example, they might say, "I've been feeling stressed lately because things haven't been going well at work." This input becomes the starting point for the system to process the data.
[0704] Input: User voice or text
[0705] Output: Raw audio or text data
[0706] Step 2:
[0707] The device converts the voice input into text data. Specifically, it uses voice recognition software (e.g., Google Speech-to-Text API) to convert the voice to text. If the input is text, it proceeds to the next step.
[0708] Input: Audio data
[0709] Output: Text data
[0710] Step 3:
[0711] The device sends the text data to the server, specifically using the HTTPS protocol to transfer the data securely.
[0712] Input: Text data
[0713] Output: Text data is sent to the server
[0714] Step 4:
[0715] The server analyzes the received text data and determines the user's emotions. Specifically, it analyzes the text using natural language processing tools (e.g., spaCy, NLTK) and classifies the emotions using machine learning models (e.g., TensorFlow).
[0716] Input: Text data
[0717] Output: Sentiment analysis result (e.g., user's emotion is "stress")
[0718] Step 5:
[0719] The server generates a counseling message using a generative AI model (e.g., GPT-3) based on the results of emotion analysis. For example, it automatically generates a message such as, "Try taking deep breaths to relax. Give it a try."
[0720] Input: Sentiment analysis results
[0721] Output: Counseling message
[0722] Step 6:
[0723] The server sends the generated counseling message to the terminal. The server sends the message using the HTTPS protocol, and the terminal receives it.
[0724] Input: Counseling message
[0725] Output: Counseling message forwarded to terminal
[0726] Step 7:
[0727] The terminal displays the counseling message to the user. Specifically, the notification function is used to make the message immediately visible to the user.
[0728] Input: Counseling message
[0729] Output: A message displayed to the user
[0730] Step 8:
[0731] The server continuously monitors the emotion data and determines the level of urgency. If the user repeatedly inputs "I feel heavy," the level of urgency is determined to be high.
[0732] Input: Continuous emotion data
[0733] Output: Urgency judgment result (e.g., high urgency)
[0734] Step 9:
[0735] If the emergency is deemed high, the server will request a response from an expert. For example, an alert email will be sent to a psychologist requesting immediate action.
[0736] Input: Urgency assessment result
[0737] Output: An alert email is sent to the expert
[0738] Step 10:
[0739] When a user requests "I want to practice for an interview," the server generates an appropriate practice scenario, which includes specific questions and sends the contents of the scenario to the terminal.
[0740] Input: Practice Request
[0741] Output: Practice scenario
[0742] Step 11:
[0743] The user practices interactively, and the server analyzes the results and generates feedback, such as "Your speech is clear and good."
[0744] Input: Practice results
[0745] Output: Feedback
[0746] (Application example 1)
[0747] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0748] In recent years, there has been an increasing demand for counseling systems to maintain users' mental health. However, existing systems lack sufficient analysis of users' emotions and assessment of the level of urgency, making it difficult to respond quickly and effectively. Furthermore, as security risks increase, systems that can provide immediate security responses based on users' emotional states are needed. Furthermore, more advanced responses using generative AI models are needed to improve the quality of practice and feedback for everyday situations requested by users. A system that can solve these issues and comprehensively support users' mental health and safety is needed.
[0749] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0750] In this invention, the server includes: means for receiving and analyzing voice or text input and determining a user's emotion; means for generating a counseling message using a generative AI model based on the determined emotion; means for providing the generated counseling message to the user; means for determining the urgency based on the emotion and requesting a response from an expert; means for issuing a security alert based on the result of the emotion analysis; means for cooperating with a security device when an alert is issued; means for receiving a practice request for an everyday situation from the user, generating a practice scenario based on the request, practicing interactively with the user, and generating feedback based on the practice results; and means for generating prompt sentences based on emotion data and generating feedback using a generative AI model. This enables the provision of counseling messages that are responsive to the user's emotional state and rapid response in high-urgency situations, and enables the immediate issuance of alerts even in high-security risk situations and the provision of comprehensive safety support in cooperation with other devices.
[0751] "Voice or text input" refers to the means by which a user provides voice or text data to a system.
[0752] "Means for determining emotion" refers to a method or device for analyzing received voice or text input and classifying the user's current emotional state as "positive," "negative," "neutral," or the like.
[0753] The "means for generating a counseling message" refers to a method or device for generating a message including appropriate advice or advice based on the result of the user's emotion determination.
[0754] A "generative AI model" is an artificial intelligence model trained based on large amounts of data, and is a means used for emotion analysis and generating counseling messages.
[0755] The "means for determining the degree of urgency" refers to a method or device for analyzing the emotional state of the user and evaluating and determining the degree of urgency of that state.
[0756] "Means for requesting a response from an expert" refers to a method or device for requesting intervention from an expert such as a psychologist or psychiatrist when the emergency is determined to be high.
[0757] "Means for issuing a security alert" refers to a method or device for issuing an alert when it is determined that the user's safety may be threatened based on the results of emotion analysis.
[0758] "Means for coordinating with security devices" refers to a method or device for coordinating with other security devices (such as surveillance cameras or intrusion detection systems) when issuing an alarm.
[0759] The "means for receiving a practice request" refers to a method or device for receiving a request for practicing a daily situation desired by a user.
[0760] "Means for generating a practice scenario" refers to a method or device for automatically generating a scenario according to the practice desired by the user.
[0761] The "means for conducting interactive practice" refers to a method or device for allowing a user and a system to have a dialogue based on a generated practice scenario.
[0762] The "means for generating feedback" refers to a method or device for analyzing the results of a user's practice and providing advice or suggestions for improvement based on the results.
[0763] A "prompt" is text data that is input into a generative AI model, providing information for the AI's response or generated message.
[0764] Components
[0765] 1. User Device:
[0766] A device that allows a user to input voice or text. It includes smartphones, tablets, personal computers, head-mounted displays, etc. The user terminal is responsible for sending input data to a server and receiving messages from the server.
[0767] 2. Server:
[0768] It receives input data, analyzes emotions, generates counseling messages, and determines the level of urgency. It also uses generative AI models to generate practice scenarios and collaborate with other security devices.
[0769] Program processing
[0770] Handling User Input
[0771] When a user inputs data by voice or text, the user device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The user device converts this data into text data and sends it to the server. The server receives this data and performs emotion analysis.
[0772] Emotion analysis
[0773] The server uses an emotion analysis model to classify the user's emotional state as either "positive," "negative," or "neutral." For example, if a user enters "I'm feeling stressed," the server will determine this as "negative."
[0774] Counseling message generation
[0775] The server generates appropriate counseling messages based on the analysis results, such as "Take a deep breath to relax," using a generative AI model, and sends the messages to the user's device.
[0776] Urgency assessment and response request
[0777] The server continuously monitors the emotion data and determines the level of urgency. If the level of urgency is deemed high, it will request a response from an expert. For example, if the user continuously expresses sadness, the server will send an alert to an expert. It will also issue a security alarm and connect to security devices.
[0778] Providing practice scenarios
[0779] When a user wants to practice an everyday situation, for example, they request, "I would like to practice an interview." The server receives this request and uses a generative AI model to generate an appropriate interview scenario. The generated scenario is sent to the user's device, where the user practices interactively. The server analyzes the practice results, generates feedback based on the prompt sentence, and sends it to the user's device.
[0780] Specific examples
[0781] 1. Sentiment analysis and counseling message provision:
[0782] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this speech into text and sends it to the server. The server determines that the user is "feeling stressed," and generates a message such as "Take a deep breath to relax," and sends it to the device.
[0783] 2. Urgency assessment and expert intervention:
[0784] The user repeatedly types "I feel heavy." The server analyzes this and determines that the situation is urgent. It sends an alert to an expert, who then contacts the user and provides counseling.
[0785] 3. Providing practice scenarios and feedback:
[0786] The user requests, "I want to practice for an interview." The server generates an appropriate scenario and sends it to the terminal. The user practices interactively, and the server analyzes the results and sends feedback. An example of a prompt sentence is, "How do you relieve nervousness during an interview?"
[0787] In this way, a system is realized in which the user terminal and server work together to support the user's mental health while also responding immediately to security risks.
[0788] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0789] Step 1:
[0790] The user provides input via voice or text.
[0791] Input: User voice or text
[0792] Output: Audio data or text data
[0793] How it works: The user uses a device such as a smartphone or head-mounted display (HMD) to input voice commands, for example, "I'm feeling stressed because things haven't been going well at work lately."
[0794] Step 2:
[0795] The device converts the voice input into text.
[0796] Input: Audio data
[0797] Output: Text data
[0798] How it works: The device's microphone records the user's voice and converts it into text using speech recognition software (e.g., Google Speech Recognition).
[0799] Step 3:
[0800] The terminal transmits the text data to the server.
[0801] Input: Text data
[0802] Output: Data sent to the server
[0803] How it works: Your device sends the converted text data to a server over your internet connection.
[0804] Step 4:
[0805] The server receives the text data and performs sentiment analysis.
[0806] Input: Text data
[0807] Output: Emotion judgment result (positive, negative, neutral, etc.)
[0808] How it works: The server inputs the received text data into a sentiment analysis model (e.g., BERT), which then classifies the user's sentiment. For example, "I'm feeling stressed" is classified as "negative."
[0809] Step 5:
[0810] The server generates a counseling message based on the emotion determination result.
[0811] Input: Emotion determination result
[0812] Output: Counseling message
[0813] How it works: The server uses the emotion determination results to input prompts into the generative AI model, generating an appropriate counseling message. For example, a message like "Take a deep breath and you'll be able to relax" is generated.
[0814] Step 6:
[0815] The server sends a counseling message to the user terminal.
[0816] Input: Counseling message
[0817] Output: Transmitted data
[0818] Operation: The server sends the generated counseling message to the user terminal.
[0819] Step 7:
[0820] The user terminal displays a counseling message to the user.
[0821] Input: Counseling message
[0822] Output: Display message
[0823] Operation: The user terminal displays the counseling message received on the screen. For example, the user sees a message on the terminal screen saying, "Take a deep breath and you'll be able to relax."
[0824] Step 8:
[0825] The server continuously monitors the emotional data and determines the level of urgency.
[0826] Input: Continuous emotion data
[0827] Output: Urgency (low, medium, high)
[0828] How it works: The server monitors the emotional data continuously sent by the user and determines the urgency based on certain criteria.
[0829] Step 9:
[0830] If the emergency is deemed high, the server will request a specialist to respond.
[0831] Input: Urgency (High)
[0832] Output: Alert to experts
[0833] How it works: When the server detects a high level of urgency, it sends an alert to registered experts and requests them to take action.
[0834] Step 10:
[0835] Additionally, the server will issue security alerts as needed and work in conjunction with security devices.
[0836] Input: High-urgency emotion data
[0837] Output: Security alerts and data sent to linked devices
[0838] How it works: The server issues security alarms and works in conjunction with security devices such as surveillance cameras and intrusion detection systems to provide comprehensive safety responses.
[0839] Step 11:
[0840] When a user submits a practice request, the server generates a practice scenario for an everyday situation.
[0841] Input: User's practice request
[0842] Output: Practice scenario
[0843] How it works: When a user requests, "I want to practice for an interview," the server uses the generative AI model to generate an appropriate practice scenario and sends it to the user's device.
[0844] Step 12:
[0845] The server analyzes the practice results, generates feedback, and sends it to the user's terminal.
[0846] Input: Practice result data
[0847] Output: Feedback message
[0848] How it works: The user performs the exercise and sends the results to the server. The server then uses the generative AI model to generate a prompt and sends feedback to the user's device. For example, the server might provide feedback such as, "It would be better if you relaxed your gaze more."
[0849] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0850] This system analyzes a user's voice or text input, determines the user's emotions, and provides appropriate counseling messages. Furthermore, by combining it with an emotion engine, it can recognize subtle changes in the user's emotional state and provide more accurate support.
[0851] Components
[0852] 1. User Device:
[0853] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[0854] 2. Server:
[0855] It receives input data and performs processes such as emotion analysis, generating counseling messages, and determining the level of urgency. It also uses an emotion engine to learn the user's emotional patterns and reflect changes in emotions in real time.
[0856] 3. Experts:
[0857] Based on requests from the server, appropriate advice and treatment will be provided to users, including psychologists, psychiatrists, and counselors.
[0858] 4. Emotion Engine:
[0859] It recognizes emotions by analyzing the user's voice, text input, and non-verbal elements (facial expressions, tone of voice, etc.). It learns emotional patterns and updates them in real time to detect subtle changes in the user's emotional state.
[0860] Program processing
[0861] Handling User Input
[0862] The user inputs data by voice or text, and the device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this data into text data and sends it to the server. The server receives this data and analyzes it using an emotion engine.
[0863] Emotion engine processing
[0864] The emotion engine in the server recognizes the user's emotions based on the input data. For example, it not only determines that the user is "feeling stressed," but also detects that the user is "very stressed" based on facial expressions and tone of voice.
[0865] Counseling message generation
[0866] The server generates a counseling message based on the analysis results of the emotion engine. For example, if it determines that the user is feeling extremely stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[0867] Urgency assessment and response request
[0868] The server determines the level of urgency based on the data obtained from the emotion engine. For example, if a user repeatedly expresses very strong feelings of depression, the server will request a response from a specialist. The specialist will receive the alert and contact the user to provide counseling or appropriate treatment.
[0869] Providing practice scenarios
[0870] When a user wants to practice an everyday situation, for example, they request, "I want to practice an interview." The server receives this request and generates an appropriate interview scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on the scenario. The emotion engine monitors and analyzes the user's reactions and provides feedback.
[0871] Matching and communication support
[0872] When a user wishes to communicate with other users, they can enter, for example, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users and sends the results to the device. The users then begin a conversation, and the server monitors the exchange, requesting expert intervention if necessary.
[0873] Specific examples
[0874] 1. Emotion analysis using an emotion engine and provision of counseling messages:
[0875] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server uses its emotion engine to determine that the user is under "very high stress," and generates a message such as, "Take a deep breath and you'll be able to relax. Try it out," which is sent to the device. The device then displays this message to the user.
[0876] 2. Urgency assessment and expert intervention:
[0877] A user repeatedly types, "I'm feeling very depressed." The server analyzes the situation using an emotion engine, and if it determines that the situation is urgent, it sends an alert to a specialist. The specialist receives the alert and contacts the user to provide counseling or treatment.
[0878] 3. Providing practice scenarios and feedback:
[0879] The user requests, "I want to practice for an interview." The server uses an emotion engine to analyze the user's emotional state and generate an appropriate scenario. The scenario is sent to the device, and the user practices. The server analyzes the practice results, generates feedback, and sends it to the device.
[0880] 4. Matching and communication support:
[0881] The user types, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users then begin to converse, and the system monitors the content and requests intervention from experts if necessary.
[0882] As described above, by combining the emotion engine, the system of the present invention can monitor the user's mental health state with higher accuracy and provide appropriate counseling and support for rehabilitation into society.
[0883] The processing flow will be explained below.
[0884] AI that listens to your heart
[0885] User Input and Sentiment Analysis
[0886] Step 1:
[0887] The user speaks or texts into the device, saying, "I've been feeling stressed lately because things haven't been going well at work."
[0888] Step 2:
[0889] The terminal converts the voice input into text data and transmits the text data to the server.
[0890] Step 3:
[0891] The server receives the text data and requests the emotion engine to analyze it.
[0892] Step 4:
[0893] The emotion engine determines from the user's text that they are "feeling stressed."
[0894] Step 5:
[0895] The emotion engine also analyzes non-verbal data such as facial expressions and tone of voice to assess the intensity of stress (if determined to be "very high stress").
[0896] Counseling message generation
[0897] Step 6:
[0898] The server generates a counseling message based on the analysis results of the emotion engine. For example, it generates a message such as, "Take a deep breath and you'll feel more relaxed. Try it out."
[0899] Step 7:
[0900] The server sends the generated counseling message to the terminal.
[0901] Step 8:
[0902] The terminal displays a counseling message to the user.
[0903] AI in harmony with specialists
[0904] Urgency assessment and response request
[0905] Step 1:
[0906] The server continuously monitors the user's daily emotional data.
[0907] Step 2:
[0908] The server determines the level of urgency based on the data obtained from the emotion engine. For example, if the user repeatedly expresses "feeling very depressed," it will determine that the level of urgency is high.
[0909] Step 3:
[0910] If the server determines that the situation is urgent, it will send an alert to an expert.
[0911] Step 4:
[0912] A specialist receives the alert from the server and contacts the user to provide counseling or treatment.
[0913] AI provides a platform for social reintegration
[0914] Providing practice scenarios
[0915] Step 1:
[0916] The user inputs a request to the terminal saying, "I want to practice for an interview."
[0917] Step 2:
[0918] The server receives the user's request and generates an interview scenario.
[0919] Step 3:
[0920] The emotion engine analyzes the user's emotional state and reflects the emotional data in the scenario.
[0921] Step 4:
[0922] The server sends the generated scenario to the terminal.
[0923] Step 5:
[0924] The terminal displays a scenario, and the user practices the interview in an interactive format.
[0925] Providing Feedback
[0926] Step 6:
[0927] When the user completes the exercise, the terminal transmits the exercise results to the server.
[0928] Step 7:
[0929] The server uses an emotion engine to analyze the practice results and generate a feedback message, such as "You did very well. You could do better if you answered with more confidence."
[0930] Step 8:
[0931] The server sends a feedback message to the terminal, which displays it to the user.
[0932] AI that provides an empathetic community
[0933] Matching with user requests
[0934] Step 1:
[0935] The user inputs their request into the terminal, saying, "I would like to talk to someone who is also aiming to return to society."
[0936] Step 2:
[0937] The server receives the user's request and analyzes it using the emotion engine.
[0938] Step 3:
[0939] The server searches and selects users in the same position from the database and works with the emotion engine to make appropriate matches.
[0940] Step 4:
[0941] The server sends the matching results to the device.
[0942] Step 5:
[0943] The terminal notifies the user of the matching result, and the user starts communication with the matched person.
[0944] Communications monitoring and support
[0945] Step 6:
[0946] The server monitors interactions within the community using an emotion engine.
[0947] Step 7:
[0948] If necessary, the server calls for expert intervention.
[0949] Through these steps, this system can comprehensively support users in maintaining their mental health and reintegrating into society. By combining it with an emotion engine, it is possible to detect even subtle changes in the user's emotional state, enabling more accurate counseling and support to be provided.
[0950] Example 2
[0951] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0952] In modern society, increasing stress and anxiety have created a demand for psychological support. However, many users have limited opportunities to receive appropriate counseling immediately. Furthermore, conventional counseling methods often fail to capture subtle changes in a user's emotions in real time, making it difficult to provide immediate, appropriate support. While there are counseling support systems that use voice or text input, few of them offer sufficient emotional analysis accuracy, real-time performance, urgency assessment, or expert intervention. This has created a demand for systems that can effectively support users' mental health.
[0953] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0954] In this invention, the server includes means for receiving a user's voice or text input, means for converting the input into text data and transmitting it to the server, means for analyzing the input to determine the user's emotion, means for generating a counseling message based on the emotion, and means for providing the counseling message to the user. This makes it possible to analyze the user's voice or text input in real time, quickly capture subtle changes in the user's emotional state, and generate and provide an appropriate counseling message. Furthermore, in cases of high urgency, it is possible to request a response from an expert, thereby enabling rapid expert intervention.
[0955] "User" refers to an individual or organization that uses the system.
[0956] "Voice input" refers to data that is converted into digital form from what a user says.
[0957] "Text input" refers to data in which a user inputs textual information using a keyboard or other input device.
[0958] "Device" means a device used by a user to input voice or text, including, for example, a smartphone, tablet, or computer.
[0959] A "server" refers to a high-performance computer system that receives and analyzes data sent by users.
[0960] "Emotion engine" refers to a system that analyzes a user's emotions from voice, text, and non-verbal data.
[0961] "Counseling message" refers to a message containing advice or support for the user, generated based on the analysis results of the emotion engine.
[0962] "Urgency" refers to an index that evaluates the level of urgency of a user's emotional state or health condition.
[0963] "Expert" refers to a person or organization qualified to provide appropriate advice or treatment for a user's mental health, such as a psychologist, psychiatrist, or counselor.
[0964] "Response request" refers to an action in which the system requests an expert to respond to the user's condition.
[0965] "Practice scenario" refers to an interactive scenario provided for a user to simulate a particular situation.
[0966] "Feedback" refers to evaluations and advice provided based on the results of the user's scenario practice.
[0967] "Matching" refers to the process of appropriately pairing users based on their purpose and status.
[0968] The present invention provides a system for supporting the mental health of a user by analyzing the user's voice or text input, determining the user's emotions, and providing appropriate counseling messages. Specific implementation methods are described in detail below.
[0969] Components
[0970] User terminal
[0971] A user terminal is a device for inputting voice or text. Terminals include smartphones, tablets, and PCs. The user terminal transmits the input data to a server and receives messages from the server.
[0972] server
[0973] The server has the central function of receiving and analyzing data sent by users. The received data is input into the emotion engine, which analyzes the user's emotions. Based on the analysis results, it generates an appropriate counseling message and sends it to the user's device. It also determines the level of urgency, requests expert assistance, provides practice scenarios and feedback, and matches users and supports communication.
[0974] Emotion Engine
[0975] The emotion engine is a system that analyzes the user's emotions using non-verbal data such as voice, text, facial expression analysis, and tone of voice. Based on the analysis results, the emotion engine detects subtle changes in the user's emotional state and reflects them in real time.
[0976] Hardware and software used
[0977] Speech Recognition Software: Uses the Google Speech-to-Text API to convert user voice input into text data.
[0978] Sentiment analysis model: We use a sentiment analysis model using the Hugging Face Transformers library.
[0979] Generative AI model: OpenAI's GPT-3 is used to generate counseling messages and practice scenarios.
[0980] Specific examples
[0981] Emotion analysis using an emotion engine and provision of counseling messages
[0982] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server uses its emotion engine to determine that the user is feeling "very stressed," and generates a message such as, "Take a deep breath and you'll be able to relax. Try it out," and sends it to the device. The device then displays this message to the user.
[0983] Urgency assessment and expert intervention
[0984] A user repeatedly types, "I'm feeling very depressed." The server analyzes the situation using an emotion engine, and if it determines that the situation is urgent, it sends an alert to a specialist. The specialist receives the alert and contacts the user to provide counseling or treatment.
[0985] Providing practice scenarios and feedback
[0986] The user requests, "I want to practice for an interview." The server uses an emotion engine to analyze the user's emotional state and generate an appropriate scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on that scenario. The server analyzes the practice results, generates feedback, and sends it to the device.
[0987] Matching and communication support
[0988] The user types, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users then begin to converse, and the server monitors the exchange and, if necessary, requests intervention from experts.
[0989] Prompt Sentence Examples
[0990] 1. If you want to use sentiment analysis with the sentiment engine:
[0991] "If a user says, 'I'm feeling stressed because things haven't been going well at work lately,' how does the emotion engine respond?"
[0992] 2. To request a priority assessment and expert intervention:
[0993] "If a user expresses very depressed feelings consecutively, how does the server determine the urgency and notify an expert?"
[0994] 3. If you would like to provide a practice scenario:
[0995] "If a user wants to practice an interview, how does the system generate scenarios and provide feedback?"
[0996] 4. If you wish to match and communicate with other users:
[0997] “If a user inputs that they want to talk to other users with similar goals of reintegration, how does the system match them with the right users and support their interactions?”
[0998] As described above, the present invention can support users' mental health with high accuracy through real-time emotion analysis using an emotion engine and the provision of counseling messages and practice scenarios using a generative AI model.
[0999] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1000] Step 1:
[1001] The user inputs data by voice or text. An example of input data is the text "I'm feeling stressed lately because things aren't going well at work." In the case of voice input, the device uses speech recognition software (e.g., Google Speech-to-Text API) to convert the voice into text data. The input for this step is the user's voice or text, and the output is text data.
[1002] Step 2:
[1003] The device sends text data to the server. Specifically, the device sends the text data entered to the server via an HTTP request. For example, text data is sent in the form of {"user_input": "Recently, things haven't been going well at work and I'm feeling stressed"}. The input of this step is text data, and the output is the data sent to the server.
[1004] Step 3:
[1005] The server passes the received input data to the emotion engine. The emotion engine uses an emotion analysis model using Hugging Face's Transformers library to analyze the text data and determine the user's emotion. For example, it determines "very strong stress." The input of this step is the received text data, and the output is the emotion analysis result.
[1006] Step 4:
[1007] The server generates a counseling message using a generative AI model (for example, OpenAI's GPT-3) based on the analysis results of the emotion engine. The prompt text is "The user's stress level is very high. Please provide some advice on how to relax." The AI model generates a message such as "Taking deep breaths can help you relax. Let's try it." The input for this step is the emotion analysis result, and the output is a counseling message.
[1008] Step 5:
[1009] The server sends the generated counseling message to the user's device. Specifically, it sends the following message using an HTTP response: {"counseling_message": "Try taking a deep breath to relax."} The device then displays the received message to the user. The input to this step is the counseling message, and the output is the message displayed on the user's device.
[1010] Step 6:
[1011] The server determines the urgency level based on the analysis results of the emotion engine. For example, if the user repeatedly expresses "very depressed feelings," the server determines the urgency level as "high." The input of this step is the emotion analysis result, and the output is the urgency level determination result.
[1012] Step 7:
[1013] If the urgency is determined to be "high," the server requests a response from an expert. An alert is sent to the expert, for example, a notification saying, "User A's urgency is high. Please respond immediately." The expert receives the alert and contacts the user to provide counseling or appropriate treatment. The input to this step is the urgency determination result, and the output is an alert notification to the expert.
[1014] Step 8:
[1015] When a user wishes to practice an everyday situation, they input a request, for example, "I would like to practice an interview." The server receives the request, analyzes the user's emotional state using an emotion engine, and generates an appropriate practice scenario. The generated scenario is sent to the device, and the user practices interactively based on that scenario. The server analyzes the practice results, generates feedback, and sends it to the device. The input to this step is the user's practice request, and the output is the practice scenario and feedback.
[1016] Step 9:
[1017] When a user wishes to communicate with other users, they input, for example, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users begin a conversation, and the server monitors the exchange, requesting intervention from experts if necessary. The input to this step is the user's communication desire, and the output is the matching results and conversation monitoring.
[1018] (Application example 2)
[1019] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1020] In recent years, with the spread of electronic payment services, users have been experiencing increasing stress and anxiety during transactions and payments. This mental burden not only worsens the user experience, but may also lead to transaction failures and increased security risks. Therefore, there is a need for a system that can monitor users' emotions in real time during electronic payment services and provide appropriate support to improve the user experience.
[1021] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a user's voice or text input, means for analyzing the input to determine the user's emotion, means for generating a counseling message based on the emotion, means for providing the counseling message to the user, and means for monitoring the user's emotion when conducting a transaction or payment and taking measures to reduce stress. In this way, by understanding the user's emotional state in real time and providing a counseling message as needed, stress during a transaction or payment can be reduced, allowing the user to use the service with peace of mind.
[1022] The "means for receiving user voice or text input" is an interface for acquiring voice or text data uttered by the user and processing it within the system.
[1023] The "means for analyzing the input and determining the user's emotions" refers to an algorithm or engine that analyzes the received voice or text data and determines the user's emotional state (e.g., stress, joy, surprise, etc.) based on that analysis.
[1024] The "means for generating a counseling message based on the emotion" is a processing unit for automatically generating an appropriate message of advice or comfort according to the determined emotional state of the user.
[1025] The "means for providing the counseling message to the user" is a module for displaying, playing, or transmitting the generated counseling message to the user's device.
[1026] "Means to monitor users' emotions when making transactions or payments and take measures to reduce stress" refers to a function that monitors users' emotional state in real time during electronic payments or transactions and provides guidelines and support messages to reduce stress and anxiety.
[1027] The "means for determining the degree of urgency based on the emotion" is a system for assessing the seriousness of the user's emotional state and determining whether an emergency response is required, if necessary.
[1028] The "means for requesting a response from an expert when the urgency is high" is a module that sends an alert to an expert such as a counselor or a specialist doctor when a high urgency is determined, urging them to intervene quickly.
[1029] The "means for receiving a user's practice request for an everyday situation" is an interface for receiving a request when a user wishes to practice a specific scenario (for example, an interview, a presentation, etc.).
[1030] The "means for generating a practice scenario based on the request" is a module for automatically creating a specific practice scenario in response to a request from a user.
[1031] The "means for interactively practicing with the user based on the practice scenario" is an interface that allows the user to practice while interacting with the system based on the generated scenario.
[1032] The "means for generating feedback based on the practice results" is a system that analyzes the user's practice results and automatically creates and provides feedback such as areas for improvement and results.
[1033] This invention is a system that analyzes a user's voice or text input, determines the user's emotions, and provides appropriate counseling messages. Furthermore, by using an emotion engine, it is possible to monitor the user's emotions in real time when conducting transactions or payments, and take necessary measures. Such a system aims to reduce the user's mental burden and provide a better user experience.
[1034] System Configuration
[1035] The system includes the following components:
[1036] 1. User Device
[1037] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[1038] 2. Server
[1039] It receives input data and performs processes such as emotion analysis, generating counseling messages, monitoring emotions during transactions and payments, and determining urgency. It also uses an emotion engine to learn the user's emotional patterns and reflect emotional changes in real time.
[1040] 3. Emotion Engine
[1041] It recognizes emotions by analyzing the user's voice, text input, and non-verbal elements (facial expressions, tone of voice, etc.). It learns emotional patterns and updates them in real time to detect subtle changes in the user's emotional state.
[1042] System Operation
[1043] In the present invention, the system operates using the following means.
[1044] 1. Handling User Input
[1045] The user inputs data by voice or text, and the device sends the data to the server. For example, the user says, "This transaction is very stressful." The device converts this data into text and sends it to the server.
[1046] 2. Emotion Engine Processing
[1047] The emotion engine in the server recognizes the user's emotions based on the input data. For example, it not only determines that the user is "feeling stressed," but also detects that the user is "very stressed" based on facial expressions and tone of voice.
[1048] 3. Generating counseling messages
[1049] The server generates a counseling message based on the analysis results of the emotion engine. For example, if it determines that the user is feeling extremely stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[1050] 4. Emotion monitoring and countermeasures during transactions and payments
[1051] The server monitors the user's emotions in real time when trading or making payments, and if it determines that the user is feeling stressed, it takes measures to reduce stress. For example, if a user enters "This transaction is very stressful" while trading, the server analyzes this and generates an advice message such as "Try taking deep breaths to relax."
[1052] Specific examples and examples of prompts for generative AI models
[1053] Example 1:
[1054] If the user enters "This transaction is very stressful," the server processes it as follows:
[1055] Emotion analysis result: NEGATIVE
[1056] Score: 0.95
[1057] Counseling message: "You seem to be stressed. Take a deep breath and try to relax a bit."
[1058] Example prompt sentence:
[1059] Create a program that analyzes a user's voice or text input, checks whether the user is feeling anxious or stressed, and provides an appropriate counseling message. Use Transformers as the emotion analysis model, and display a stress relief message if the input indicates a negative emotion.
[1060] The above is a specific embodiment for carrying out this invention. By combining this system with an emotion engine, it is possible to monitor the user's mental health with greater precision and provide appropriate counseling and transaction / payment support.
[1061] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1062] Step 1:
[1063] The user provides voice or text input. For example, the user might say, "This transaction is very stressful." This input is captured by the terminal and converted into text data.
[1064] Step 2:
[1065] The device sends the converted text data to the server, which receives the data and prepares it for sentiment analysis. The input is the user's voice or text data, and the output is the text data sent to the server.
[1066] Step 3:
[1067] The server uses an emotion engine to analyze the incoming data. Specifically, it applies a generative AI model such as Transformers to determine the user's emotion. The input is text data, and the output is an emotion (e.g., NEGATIVE) and its score (e.g., 0.95).
[1068] Step 4:
[1069] The server runs an algorithm that generates a counseling message based on the analysis results. Specifically, if the emotion is "NEGATIVE" and the score is high, a counseling message such as "You seem to be feeling stressed. Take a deep breath and try to relax a bit" is created. The input is the emotion analysis result and score, and the output is the generated counseling message.
[1070] Step 5:
[1071] The server sends the generated counseling message to the terminal. The terminal receives this message and displays or plays it aloud to the user. The input is the counseling message sent from the server, and the output is the message provided in a format that the user can view.
[1072] Step 6:
[1073] The server monitors users' emotions in real time and continuously provides stress-reducing measures during transactions and settlements. For example, if a transaction is prolonged or if a particular frustration persists, it generates additional messages such as "Take a short break to refresh yourself." The input is continuous emotion analysis data, and the output is support messages provided at any time.
[1074] Step 7:
[1075] The system determines the urgency of the emotion, and if the server deems it necessary, sends an alert to an expert. For example, if a user repeatedly expresses very strong, depressed emotions, the system will contact an expert and ask for their response. The input is the emotion analysis result and urgency assessment data, and the output is an alert to the expert and a request for their response.
[1076] Step 8:
[1077] Based on the user's request, the server generates a practice scenario for an everyday situation. If the user requests "I want to practice an interview," the server generates an appropriate scenario and sends it to the terminal. The input is the user's practice request, and the output is the generated practice scenario.
[1078] Step 9:
[1079] The user performs interactive practice based on the generated scenario, which involves the system analyzing the user's responses and providing appropriate feedback. The input is the user's interaction data, and the output is the feedback provided by the system.
[1080] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1081] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1082] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1083] [Third embodiment]
[1084] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1085] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1086] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1087] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1088] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1089] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1090] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1091] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1092] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1093] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1094] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1095] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1096] The present invention is a system that analyzes a user's voice or text input, determines their emotions, and provides appropriate counseling messages. This system is made up of the following components:
[1097] Components
[1098] 1. User Device:
[1099] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[1100] 2. Server:
[1101] It receives input data, analyzes emotions, generates counseling messages, and determines the level of urgency. It also generates practice scenarios, matches users with other users, and monitors communication content.
[1102] 3. Experts:
[1103] Based on requests from the server, appropriate advice and treatment will be provided to users, including psychologists, psychiatrists, and counselors.
[1104] Program processing
[1105] Handling User Input
[1106] The user inputs data by voice or text, and the device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server receives this data, performs emotional analysis, and determines that the user is "feeling stressed."
[1107] Counseling message generation
[1108] The server generates a counseling message based on the results of emotion analysis. For example, if it determines that the user is feeling stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[1109] Urgency assessment and response request
[1110] The server continuously monitors the emotion data and determines the level of urgency. If the level is deemed high, it requests a specialist to respond. For example, if a user continuously expresses depressed feelings, the server sends an alert to the specialist. The specialist receives the alert, contacts the user, and provides appropriate counseling or treatment.
[1111] Providing practice scenarios
[1112] When a user wishes to practice an everyday situation, for example, they make a request such as "I would like to practice an interview." The server receives this request and generates an appropriate interview scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on the scenario. The server analyzes the practice results, generates feedback, sends it to the device, and provides it to the user.
[1113] Matching and communication support
[1114] When a user wishes to communicate with other users, they can enter, for example, "I want to talk to someone who is also trying to reintegrate into society." The server searches and selects users in the same position from its database and performs matching. The matching results are sent to the user's device and notified. The user then begins a conversation with the matched person, and the system monitors the conversation. If necessary, the server requests expert intervention.
[1115] Specific examples
[1116] 1. Sentiment analysis and counseling message provision:
[1117] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this speech into text and sends it to the server. The server determines that the user is "feeling stressed," and generates a message such as "Take a deep breath to relax," which is sent to the device. The device then displays this message to the user.
[1118] 2. Urgency assessment and expert intervention:
[1119] The user repeatedly types "I feel heavy." The server analyzes this and determines that the situation is urgent. It sends an alert to an expert, who then contacts the user and provides counseling.
[1120] 3. Providing practice scenarios and feedback:
[1121] The user requests, "I want to practice an interview." The server generates an appropriate scenario and sends it to the device. The user practices interactively, and the server analyzes the results and sends feedback.
[1122] 4. Matching and communication support:
[1123] The user types, "I want to talk to people who are also trying to reintegrate into society." The server matches suitable users and notifies the device. The users begin a conversation, and the system monitors the content. If necessary, it requests intervention from an expert.
[1124] As described above, the system of the present invention can monitor the mental health state of the user and provide appropriate counseling and support for rehabilitation into society.
[1125] The processing flow will be explained below.
[1126] AI that listens to your heart
[1127] User Input and Sentiment Analysis
[1128] Step 1:
[1129] The user speaks or texts into the device, saying, "I've been feeling stressed lately because things haven't been going well at work."
[1130] Step 2:
[1131] The terminal converts the voice input into text data and transmits the text data to the server.
[1132] Step 3:
[1133] A server receives the text data and runs a sentiment analysis model (e.g., using natural language processing).
[1134] Step 4:
[1135] The server determines from the user's text that he or she is "feeling stressed."
[1136] Counseling message generation
[1137] Step 5:
[1138] The server generates a counseling message for stress reduction (e.g., "Take a deep breath and you'll feel more relaxed").
[1139] Step 6:
[1140] The server sends the generated counseling message to the terminal.
[1141] Step 7:
[1142] The terminal displays a counseling message to the user.
[1143] AI in harmony with specialists
[1144] Urgency assessment and response request
[1145] Step 1:
[1146] The server monitors the user's daily emotional data.
[1147] Step 2:
[1148] The server analyzes the emotional data and determines whether the person is experiencing a continuous state of depression.
[1149] Step 3:
[1150] The server determines the level of urgency and, if deemed high, sends an alert to an expert.
[1151] Step 4:
[1152] A specialist receives the alert from the server and contacts the user to provide counseling or treatment.
[1153] AI provides a platform for social reintegration
[1154] Providing practice scenarios
[1155] Step 1:
[1156] The user inputs a request to the terminal saying, "I want to practice for an interview."
[1157] Step 2:
[1158] The server receives the user's request and generates an interview scenario.
[1159] Step 3:
[1160] The server sends the generated scenario to the terminal.
[1161] Step 4:
[1162] The terminal displays a scenario, and the user practices the interview in an interactive format.
[1163] Step 5:
[1164] When the user completes the exercise, the terminal transmits the exercise results to the server.
[1165] Providing Feedback
[1166] Step 6:
[1167] The server analyzes the practice results and generates a feedback message.
[1168] Step 7:
[1169] The server sends a feedback message to the terminal, which displays it to the user.
[1170] AI that provides an empathetic community
[1171] Matching with user requests
[1172] Step 1:
[1173] The user inputs their request into the terminal, saying, "I would like to talk to someone who is also aiming to return to society."
[1174] Step 2:
[1175] The server analyzes users in the same position from the database and matches them with the appropriate users.
[1176] Step 3:
[1177] The server sends the matching results to the device.
[1178] Step 4:
[1179] The user begins communicating with the matched person (e.g., "I see you're in the same situation, ○○. Let's both do our best.").
[1180] Communications monitoring and support
[1181] Step 5:
[1182] The server monitors interactions within the community.
[1183] Step 6:
[1184] If necessary, the server calls for expert intervention.
[1185] Through the above steps, this system can comprehensively support the user's mental health maintenance and social reintegration.
[1186] Example 1
[1187] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1188] Conventional counseling systems struggle to quickly and accurately analyze users' emotions and provide appropriate counseling messages. Furthermore, delays in assessing the level of urgency and providing appropriate expert intervention can lead to a deterioration in the user's mental health. Furthermore, they lack sufficient practice in everyday situations and support for communication between users, preventing them from effectively supporting users' reintegration into society.
[1189] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1190] In this invention, the server includes means for receiving a user's voice or text input, means for converting the input into text data, means for transmitting the text data to the server, means for analyzing the text data to determine the user's emotion, means for generating a counseling message using a generative AI model based on the result of the emotion analysis, means for transmitting the counseling message to a user terminal, and means for displaying the counseling message to the user on the user terminal. This makes it possible to quickly and accurately analyze the user's emotion and effectively support appropriate counseling and social reintegration.
[1191] "Voice or text input" refers to the form of speech or written information that a user uses to communicate their feelings or requests to a system.
[1192] "Text data" refers to data obtained by converting voice input into text format, or text information entered directly by the user.
[1193] "Server" refers to the central computing system that receives and analyzes voice or text data, generates counseling messages, determines urgency, etc.
[1194] "Sentiment analysis" refers to the process of analyzing a user's text data to determine the user's emotional state.
[1195] A "generative AI model" refers to an artificial intelligence algorithm, primarily a natural language generation model, that generates appropriate messages based on the results of user sentiment analysis.
[1196] "Counseling message" refers to a message of advice or encouragement provided to a user that is generated based on the results of an analysis of the user's emotions.
[1197] "Urgency" refers to an index that evaluates the urgency of the user's emotional state and indicates whether a prompt response is required.
[1198] "Professional" refers to a psychologist, psychiatrist, counselor, or other person qualified to assist users with their mental health.
[1199] A "practice scenario" refers to a simulation scenario generated by the server when a user wishes to practice a particular situation (such as an interview).
[1200] "Feedback" refers to evaluations and advice provided to users based on the results of their practice.
[1201] "Matching" refers to the process by which a server selects and connects a suitable partner when a user wishes to communicate with another user.
[1202] "User terminal" refers to a device (smartphone, tablet, PC, etc.) used to input voice or text and receive and display counseling messages and feedback from the server.
[1203] The "HTTPS protocol" refers to a communication protocol for secure data communication over the Internet.
[1204] "Natural language processing tool" refers to a software tool (e.g., spaCy, NLTK, etc.) used to analyze a user's text data.
[1205] The present invention is a system that analyzes a user's voice or text input, determines their emotions, and provides appropriate counseling messages. This can support the user's mental health and effectively assist them in their social reintegration. A specific embodiment of this system is described below.
[1206] System Configuration
[1207] This system consists of user devices, servers, and expert components. Specific hardware components for user devices include smartphones, tablets, and PCs. Servers are high-performance servers (e.g., cloud virtual machines) with various software installed.
[1208] User terminal: A device that receives voice or text input, converts it into text data, and sends it to a server.
[1209] Server: This is a critical computing system that performs sentiment analysis, generates counseling messages, and determines urgency. Specifically, it uses natural language processing tools (e.g., spaCy, NLTK) and machine learning libraries (e.g., TensorFlow, PyTorch).
[1210] Experts: Psychologists, psychiatrists, counselors, or other qualified individuals available to support users' mental health.
[1211] Example of operation
[1212] Processing voice or text input
[1213] The user inputs data by voice or text. The user's device receives this data, and in the case of voice input, it converts it into text data using the Google Speech-to-Text API. For example, if the user says, "Recently, things haven't been going well at work and I'm feeling stressed," the content is converted into text. This text data is sent to the server using the HTTPS protocol.
[1214] Emotion analysis
[1215] The server analyzes the received text data and determines the user's emotions. Using a natural language processing tool (e.g., spaCy), it determines that the user is "feeling stressed." The emotion analysis algorithm is trained using a machine learning library (e.g., TensorFlow).
[1216] Counseling message generation
[1217] Based on the results of the sentiment analysis, a generative AI model (e.g., GPT-3) is used to generate an appropriate counseling message. For example, a message such as "Take a deep breath and you'll feel more relaxed. Try it out." This message is sent to the user's device and immediately displayed to the user.
[1218] Urgency assessment and expert intervention
[1219] The server continuously monitors the user's emotional data and determines the level of urgency. If the user repeatedly inputs "I feel heavy," the level of urgency is determined to be high. If the level of urgency is high, the server sends an alert to an expert, who then contacts the user and provides counseling.
[1220] Providing practice scenarios
[1221] When a user requests "I want to practice for an interview," the server generates an appropriate practice scenario. The scenario contains specific questions, and the contents of the scenario are sent to the user's terminal. The user practices interactively, and the server analyzes the results and generates feedback.
[1222] Prompt Sentence Examples
[1223] Emotion Analysis Prompt: Enter "I've been feeling stressed lately because things haven't been going well at work."
[1224] Urgency determination prompt: Enter "I feel heavy" repeatedly.
[1225] Practice Scenario Prompt: Type "I want to practice interviewing."
[1226] Matching prompt: Enter "I'd like to talk to someone who is also trying to reintegrate into society."
[1227] As described above, the system of the present invention is realized using advanced technology to support the user's mental health and effectively assist in their reintegration into society.
[1228] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1229] Step 1:
[1230] The user inputs data by voice or text. For example, they might say, "I've been feeling stressed lately because things haven't been going well at work." This input becomes the starting point for the system to process the data.
[1231] Input: User voice or text
[1232] Output: Raw audio or text data
[1233] Step 2:
[1234] The device converts the voice input into text data. Specifically, it uses voice recognition software (e.g., Google Speech-to-Text API) to convert the voice to text. If the input is text, it proceeds to the next step.
[1235] Input: Audio data
[1236] Output: Text data
[1237] Step 3:
[1238] The device sends the text data to the server, specifically using the HTTPS protocol to transfer the data securely.
[1239] Input: Text data
[1240] Output: Text data is sent to the server
[1241] Step 4:
[1242] The server analyzes the received text data and determines the user's emotions. Specifically, it analyzes the text using natural language processing tools (e.g., spaCy, NLTK) and classifies the emotions using machine learning models (e.g., TensorFlow).
[1243] Input: Text data
[1244] Output: Sentiment analysis result (e.g., user's emotion is "stress")
[1245] Step 5:
[1246] The server generates a counseling message using a generative AI model (e.g., GPT-3) based on the results of emotion analysis. For example, it automatically generates a message such as, "Try taking deep breaths to relax. Give it a try."
[1247] Input: Sentiment analysis results
[1248] Output: Counseling message
[1249] Step 6:
[1250] The server sends the generated counseling message to the terminal. The server sends the message using the HTTPS protocol, and the terminal receives it.
[1251] Input: Counseling message
[1252] Output: Counseling message forwarded to terminal
[1253] Step 7:
[1254] The terminal displays the counseling message to the user. Specifically, the notification function is used to make the message immediately visible to the user.
[1255] Input: Counseling message
[1256] Output: A message displayed to the user
[1257] Step 8:
[1258] The server continuously monitors the emotion data and determines the level of urgency. If the user repeatedly inputs "I feel heavy," the level of urgency is determined to be high.
[1259] Input: Continuous emotion data
[1260] Output: Urgency judgment result (e.g., high urgency)
[1261] Step 9:
[1262] If the emergency is deemed high, the server will request a response from an expert. For example, an alert email will be sent to a psychologist requesting immediate action.
[1263] Input: Urgency assessment result
[1264] Output: An alert email is sent to the expert
[1265] Step 10:
[1266] When a user requests "I want to practice for an interview," the server generates an appropriate practice scenario, which includes specific questions and sends the contents of the scenario to the terminal.
[1267] Input: Practice Request
[1268] Output: Practice scenario
[1269] Step 11:
[1270] The user practices interactively, and the server analyzes the results and generates feedback, such as "Your speech is clear and good."
[1271] Input: Practice results
[1272] Output: Feedback
[1273] (Application example 1)
[1274] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1275] In recent years, there has been an increasing demand for counseling systems to maintain users' mental health. However, existing systems lack sufficient analysis of users' emotions and assessment of the level of urgency, making it difficult to respond quickly and effectively. Furthermore, as security risks increase, systems that can provide immediate security responses based on users' emotional states are needed. Furthermore, more advanced responses using generative AI models are needed to improve the quality of practice and feedback for everyday situations requested by users. A system that can solve these issues and comprehensively support users' mental health and safety is needed.
[1276] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1277] In this invention, the server includes: means for receiving and analyzing voice or text input and determining a user's emotion; means for generating a counseling message using a generative AI model based on the determined emotion; means for providing the generated counseling message to the user; means for determining the urgency based on the emotion and requesting a response from an expert; means for issuing a security alert based on the result of the emotion analysis; means for cooperating with a security device when an alert is issued; means for receiving a practice request for an everyday situation from the user, generating a practice scenario based on the request, practicing interactively with the user, and generating feedback based on the practice results; and means for generating prompt sentences based on emotion data and generating feedback using a generative AI model. This enables the provision of counseling messages that are responsive to the user's emotional state and rapid response in high-urgency situations, and enables the immediate issuance of alerts even in high-security risk situations and the provision of comprehensive safety support in cooperation with other devices.
[1278] "Voice or text input" refers to the means by which a user provides voice or text data to a system.
[1279] "Means for determining emotion" refers to a method or device for analyzing received voice or text input and classifying the user's current emotional state as "positive," "negative," "neutral," or the like.
[1280] The "means for generating a counseling message" refers to a method or device for generating a message including appropriate advice or advice based on the result of the user's emotion determination.
[1281] A "generative AI model" is an artificial intelligence model trained based on large amounts of data, and is a means used for emotion analysis and generating counseling messages.
[1282] The "means for determining the degree of urgency" refers to a method or device for analyzing the emotional state of the user and evaluating and determining the degree of urgency of that state.
[1283] "Means for requesting a response from an expert" refers to a method or device for requesting intervention from an expert such as a psychologist or psychiatrist when the emergency is determined to be high.
[1284] "Means for issuing a security alert" refers to a method or device for issuing an alert when it is determined that the user's safety may be threatened based on the results of emotion analysis.
[1285] "Means for coordinating with security devices" refers to a method or device for coordinating with other security devices (such as surveillance cameras or intrusion detection systems) when issuing an alarm.
[1286] The "means for receiving a practice request" refers to a method or device for receiving a request for practicing a daily situation desired by a user.
[1287] "Means for generating a practice scenario" refers to a method or device for automatically generating a scenario according to the practice desired by the user.
[1288] The "means for conducting interactive practice" refers to a method or device for allowing a user and a system to have a dialogue based on a generated practice scenario.
[1289] The "means for generating feedback" refers to a method or device for analyzing the results of a user's practice and providing advice or suggestions for improvement based on the results.
[1290] A "prompt" is text data that is input into a generative AI model, providing information for the AI's response or generated message.
[1291] Components
[1292] 1. User Device:
[1293] A device that allows a user to input voice or text. It includes smartphones, tablets, personal computers, head-mounted displays, etc. The user terminal is responsible for sending input data to a server and receiving messages from the server.
[1294] 2. Server:
[1295] It receives input data, analyzes emotions, generates counseling messages, and determines the level of urgency. It also uses generative AI models to generate practice scenarios and collaborate with other security devices.
[1296] Program processing
[1297] Handling User Input
[1298] When a user inputs data by voice or text, the user device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The user device converts this data into text data and sends it to the server. The server receives this data and performs emotion analysis.
[1299] Emotion analysis
[1300] The server uses an emotion analysis model to classify the user's emotional state as either "positive," "negative," or "neutral." For example, if a user enters "I'm feeling stressed," the server will determine this as "negative."
[1301] Counseling message generation
[1302] The server generates appropriate counseling messages based on the analysis results, such as "Take a deep breath to relax," using a generative AI model, and sends the messages to the user's device.
[1303] Urgency assessment and response request
[1304] The server continuously monitors the emotion data and determines the level of urgency. If the level of urgency is deemed high, it will request a response from an expert. For example, if the user continuously expresses sadness, the server will send an alert to an expert. It will also issue a security alarm and connect to security devices.
[1305] Providing practice scenarios
[1306] When a user wants to practice an everyday situation, for example, they request, "I would like to practice an interview." The server receives this request and uses a generative AI model to generate an appropriate interview scenario. The generated scenario is sent to the user's device, where the user practices interactively. The server analyzes the practice results, generates feedback based on the prompt sentence, and sends it to the user's device.
[1307] Specific examples
[1308] 1. Sentiment analysis and counseling message provision:
[1309] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this speech into text and sends it to the server. The server determines that the user is "feeling stressed," and generates a message such as "Take a deep breath to relax," and sends it to the device.
[1310] 2. Urgency assessment and expert intervention:
[1311] The user repeatedly types "I feel heavy." The server analyzes this and determines that the situation is urgent. It sends an alert to an expert, who then contacts the user and provides counseling.
[1312] 3. Providing practice scenarios and feedback:
[1313] The user requests, "I want to practice for an interview." The server generates an appropriate scenario and sends it to the terminal. The user practices interactively, and the server analyzes the results and sends feedback. An example of a prompt sentence is, "How do you relieve nervousness during an interview?"
[1314] In this way, a system is realized in which the user terminal and server work together to support the user's mental health while also responding immediately to security risks.
[1315] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1316] Step 1:
[1317] The user provides input via voice or text.
[1318] Input: User voice or text
[1319] Output: Audio data or text data
[1320] How it works: The user uses a device such as a smartphone or head-mounted display (HMD) to input voice commands, for example, "I'm feeling stressed because things haven't been going well at work lately."
[1321] Step 2:
[1322] The device converts the voice input into text.
[1323] Input: Audio data
[1324] Output: Text data
[1325] How it works: The device's microphone records the user's voice and converts it into text using speech recognition software (e.g., Google Speech Recognition).
[1326] Step 3:
[1327] The terminal transmits the text data to the server.
[1328] Input: Text data
[1329] Output: Data sent to the server
[1330] How it works: Your device sends the converted text data to a server over your internet connection.
[1331] Step 4:
[1332] The server receives the text data and performs sentiment analysis.
[1333] Input: Text data
[1334] Output: Emotion judgment result (positive, negative, neutral, etc.)
[1335] How it works: The server inputs the received text data into a sentiment analysis model (e.g., BERT), which then classifies the user's sentiment. For example, "I'm feeling stressed" is classified as "negative."
[1336] Step 5:
[1337] The server generates a counseling message based on the emotion determination result.
[1338] Input: Emotion determination result
[1339] Output: Counseling message
[1340] How it works: The server uses the emotion determination results to input prompts into the generative AI model, generating an appropriate counseling message. For example, a message like "Take a deep breath and you'll be able to relax" is generated.
[1341] Step 6:
[1342] The server sends a counseling message to the user terminal.
[1343] Input: Counseling message
[1344] Output: Transmitted data
[1345] Operation: The server sends the generated counseling message to the user terminal.
[1346] Step 7:
[1347] The user terminal displays a counseling message to the user.
[1348] Input: Counseling message
[1349] Output: Display message
[1350] Operation: The user terminal displays the counseling message received on the screen. For example, the user sees a message on the terminal screen saying, "Take a deep breath and you'll be able to relax."
[1351] Step 8:
[1352] The server continuously monitors the emotional data and determines the level of urgency.
[1353] Input: Continuous emotion data
[1354] Output: Urgency (low, medium, high)
[1355] How it works: The server monitors the emotional data continuously sent by the user and determines the urgency based on certain criteria.
[1356] Step 9:
[1357] If the emergency is deemed high, the server will request a specialist to respond.
[1358] Input: Urgency (High)
[1359] Output: Alert to experts
[1360] How it works: When the server detects a high level of urgency, it sends an alert to registered experts and requests them to take action.
[1361] Step 10:
[1362] Additionally, the server will issue security alerts as needed and work in conjunction with security devices.
[1363] Input: High-urgency emotion data
[1364] Output: Security alerts and data sent to linked devices
[1365] How it works: The server issues security alarms and works in conjunction with security devices such as surveillance cameras and intrusion detection systems to provide comprehensive safety responses.
[1366] Step 11:
[1367] When a user submits a practice request, the server generates a practice scenario for an everyday situation.
[1368] Input: User's practice request
[1369] Output: Practice scenario
[1370] How it works: When a user requests, "I want to practice for an interview," the server uses the generative AI model to generate an appropriate practice scenario and sends it to the user's device.
[1371] Step 12:
[1372] The server analyzes the practice results, generates feedback, and sends it to the user's terminal.
[1373] Input: Practice result data
[1374] Output: Feedback message
[1375] How it works: The user performs the exercise and sends the results to the server. The server then uses the generative AI model to generate a prompt and sends feedback to the user's device. For example, the server might provide feedback such as, "It would be better if you relaxed your gaze more."
[1376] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1377] This system analyzes a user's voice or text input, determines the user's emotions, and provides appropriate counseling messages. Furthermore, by combining it with an emotion engine, it can recognize subtle changes in the user's emotional state and provide more accurate support.
[1378] Components
[1379] 1. User Device:
[1380] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[1381] 2. Server:
[1382] It receives input data and performs processes such as emotion analysis, generating counseling messages, and determining the level of urgency. It also uses an emotion engine to learn the user's emotional patterns and reflect changes in emotions in real time.
[1383] 3. Experts:
[1384] Based on requests from the server, appropriate advice and treatment will be provided to users, including psychologists, psychiatrists, and counselors.
[1385] 4. Emotion Engine:
[1386] It recognizes emotions by analyzing the user's voice, text input, and non-verbal elements (facial expressions, tone of voice, etc.). It learns emotional patterns and updates them in real time to detect subtle changes in the user's emotional state.
[1387] Program processing
[1388] Handling User Input
[1389] The user inputs data by voice or text, and the device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this data into text data and sends it to the server. The server receives this data and analyzes it using an emotion engine.
[1390] Emotion engine processing
[1391] The emotion engine in the server recognizes the user's emotions based on the input data. For example, it not only determines that the user is "feeling stressed," but also detects that the user is "very stressed" based on facial expressions and tone of voice.
[1392] Counseling message generation
[1393] The server generates a counseling message based on the analysis results of the emotion engine. For example, if it determines that the user is feeling extremely stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[1394] Urgency assessment and response request
[1395] The server determines the level of urgency based on the data obtained from the emotion engine. For example, if a user repeatedly expresses very strong feelings of depression, the server will request a response from a specialist. The specialist will receive the alert and contact the user to provide counseling or appropriate treatment.
[1396] Providing practice scenarios
[1397] When a user wants to practice an everyday situation, for example, they request, "I want to practice an interview." The server receives this request and generates an appropriate interview scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on the scenario. The emotion engine monitors and analyzes the user's reactions and provides feedback.
[1398] Matching and communication support
[1399] When a user wishes to communicate with other users, they can enter, for example, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users and sends the results to the device. The users then begin a conversation, and the server monitors the exchange, requesting expert intervention if necessary.
[1400] Specific examples
[1401] 1. Emotion analysis using an emotion engine and provision of counseling messages:
[1402] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server uses its emotion engine to determine that the user is under "very high stress," and generates a message such as, "Take a deep breath and you'll be able to relax. Try it out," which is sent to the device. The device then displays this message to the user.
[1403] 2. Urgency assessment and expert intervention:
[1404] A user repeatedly types, "I'm feeling very depressed." The server analyzes the situation using an emotion engine, and if it determines that the situation is urgent, it sends an alert to a specialist. The specialist receives the alert and contacts the user to provide counseling or treatment.
[1405] 3. Providing practice scenarios and feedback:
[1406] The user requests, "I want to practice for an interview." The server uses an emotion engine to analyze the user's emotional state and generate an appropriate scenario. The scenario is sent to the device, and the user practices. The server analyzes the practice results, generates feedback, and sends it to the device.
[1407] 4. Matching and communication support:
[1408] The user types, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users then begin to converse, and the system monitors the content and requests intervention from experts if necessary.
[1409] As described above, by combining the emotion engine, the system of the present invention can monitor the user's mental health state with higher accuracy and provide appropriate counseling and support for rehabilitation into society.
[1410] The processing flow will be explained below.
[1411] AI that listens to your heart
[1412] User Input and Sentiment Analysis
[1413] Step 1:
[1414] The user speaks or texts into the device, saying, "I've been feeling stressed lately because things haven't been going well at work."
[1415] Step 2:
[1416] The terminal converts the voice input into text data and transmits the text data to the server.
[1417] Step 3:
[1418] The server receives the text data and requests the emotion engine to analyze it.
[1419] Step 4:
[1420] The emotion engine determines from the user's text that they are "feeling stressed."
[1421] Step 5:
[1422] The emotion engine also analyzes non-verbal data such as facial expressions and tone of voice to assess the intensity of stress (if determined to be "very high stress").
[1423] Counseling message generation
[1424] Step 6:
[1425] The server generates a counseling message based on the analysis results of the emotion engine. For example, it generates a message such as, "Take a deep breath and you'll feel more relaxed. Try it out."
[1426] Step 7:
[1427] The server sends the generated counseling message to the terminal.
[1428] Step 8:
[1429] The terminal displays a counseling message to the user.
[1430] AI in harmony with specialists
[1431] Urgency assessment and response request
[1432] Step 1:
[1433] The server continuously monitors the user's daily emotional data.
[1434] Step 2:
[1435] The server determines the level of urgency based on the data obtained from the emotion engine. For example, if the user repeatedly expresses "feeling very depressed," it will determine that the level of urgency is high.
[1436] Step 3:
[1437] If the server determines that the situation is urgent, it will send an alert to an expert.
[1438] Step 4:
[1439] A specialist receives the alert from the server and contacts the user to provide counseling or treatment.
[1440] AI provides a platform for social reintegration
[1441] Providing practice scenarios
[1442] Step 1:
[1443] The user inputs a request to the terminal saying, "I want to practice for an interview."
[1444] Step 2:
[1445] The server receives the user's request and generates an interview scenario.
[1446] Step 3:
[1447] The emotion engine analyzes the user's emotional state and reflects the emotional data in the scenario.
[1448] Step 4:
[1449] The server sends the generated scenario to the terminal.
[1450] Step 5:
[1451] The terminal displays a scenario, and the user practices the interview in an interactive format.
[1452] Providing Feedback
[1453] Step 6:
[1454] When the user completes the exercise, the terminal transmits the exercise results to the server.
[1455] Step 7:
[1456] The server uses an emotion engine to analyze the practice results and generate a feedback message, such as "You did very well. You could do better if you answered with more confidence."
[1457] Step 8:
[1458] The server sends a feedback message to the terminal, which displays it to the user.
[1459] AI that provides an empathetic community
[1460] Matching with user requests
[1461] Step 1:
[1462] The user inputs their request into the terminal, saying, "I would like to talk to someone who is also aiming to return to society."
[1463] Step 2:
[1464] The server receives the user's request and analyzes it using the emotion engine.
[1465] Step 3:
[1466] The server searches and selects users in the same position from the database and works with the emotion engine to make appropriate matches.
[1467] Step 4:
[1468] The server sends the matching results to the device.
[1469] Step 5:
[1470] The terminal notifies the user of the matching result, and the user starts communication with the matched person.
[1471] Communications monitoring and support
[1472] Step 6:
[1473] The server monitors interactions within the community using an emotion engine.
[1474] Step 7:
[1475] If necessary, the server calls for expert intervention.
[1476] Through these steps, this system can comprehensively support users in maintaining their mental health and reintegrating into society. By combining it with an emotion engine, it is possible to detect even subtle changes in the user's emotional state, enabling more accurate counseling and support to be provided.
[1477] Example 2
[1478] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1479] In modern society, increasing stress and anxiety have created a demand for psychological support. However, many users have limited opportunities to receive appropriate counseling immediately. Furthermore, conventional counseling methods often fail to capture subtle changes in a user's emotions in real time, making it difficult to provide immediate, appropriate support. While there are counseling support systems that use voice or text input, few of them offer sufficient emotional analysis accuracy, real-time performance, urgency assessment, or expert intervention. This has created a demand for systems that can effectively support users' mental health.
[1480] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1481] In this invention, the server includes means for receiving a user's voice or text input, means for converting the input into text data and transmitting it to the server, means for analyzing the input to determine the user's emotion, means for generating a counseling message based on the emotion, and means for providing the counseling message to the user. This makes it possible to analyze the user's voice or text input in real time, quickly capture subtle changes in the user's emotional state, and generate and provide an appropriate counseling message. Furthermore, in cases of high urgency, it is possible to request a response from an expert, thereby enabling rapid expert intervention.
[1482] "User" refers to an individual or organization that uses the system.
[1483] "Voice input" refers to data that is converted into digital form from what a user says.
[1484] "Text input" refers to data in which a user inputs textual information using a keyboard or other input device.
[1485] "Device" means a device used by a user to input voice or text, including, for example, a smartphone, tablet, or computer.
[1486] A "server" refers to a high-performance computer system that receives and analyzes data sent by users.
[1487] "Emotion engine" refers to a system that analyzes a user's emotions from voice, text, and non-verbal data.
[1488] "Counseling message" refers to a message containing advice or support for the user, generated based on the analysis results of the emotion engine.
[1489] "Urgency" refers to an index that evaluates the level of urgency of a user's emotional state or health condition.
[1490] "Expert" refers to a person or organization qualified to provide appropriate advice or treatment for a user's mental health, such as a psychologist, psychiatrist, or counselor.
[1491] "Response request" refers to an action in which the system requests an expert to respond to the user's condition.
[1492] "Practice scenario" refers to an interactive scenario provided for a user to simulate a particular situation.
[1493] "Feedback" refers to evaluations and advice provided based on the results of the user's scenario practice.
[1494] "Matching" refers to the process of appropriately pairing users based on their purpose and status.
[1495] The present invention provides a system for supporting the mental health of a user by analyzing the user's voice or text input, determining the user's emotions, and providing appropriate counseling messages. Specific implementation methods are described in detail below.
[1496] Components
[1497] User terminal
[1498] A user terminal is a device for inputting voice or text. Terminals include smartphones, tablets, and PCs. The user terminal transmits the input data to a server and receives messages from the server.
[1499] server
[1500] The server has the central function of receiving and analyzing data sent by users. The received data is input into the emotion engine, which analyzes the user's emotions. Based on the analysis results, it generates an appropriate counseling message and sends it to the user's device. It also determines the level of urgency, requests expert assistance, provides practice scenarios and feedback, and matches users and supports communication.
[1501] Emotion Engine
[1502] The emotion engine is a system that analyzes the user's emotions using non-verbal data such as voice, text, facial expression analysis, and tone of voice. Based on the analysis results, the emotion engine detects subtle changes in the user's emotional state and reflects them in real time.
[1503] Hardware and software used
[1504] Speech Recognition Software: Uses the Google Speech-to-Text API to convert user voice input into text data.
[1505] Sentiment analysis model: We use a sentiment analysis model using the Hugging Face Transformers library.
[1506] Generative AI model: OpenAI's GPT-3 is used to generate counseling messages and practice scenarios.
[1507] Specific examples
[1508] Emotion analysis using an emotion engine and provision of counseling messages
[1509] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server uses its emotion engine to determine that the user is feeling "very stressed," and generates a message such as, "Take a deep breath and you'll be able to relax. Try it out," and sends it to the device. The device then displays this message to the user.
[1510] Urgency assessment and expert intervention
[1511] A user repeatedly types, "I'm feeling very depressed." The server analyzes the situation using an emotion engine, and if it determines that the situation is urgent, it sends an alert to a specialist. The specialist receives the alert and contacts the user to provide counseling or treatment.
[1512] Providing practice scenarios and feedback
[1513] The user requests, "I want to practice for an interview." The server uses an emotion engine to analyze the user's emotional state and generate an appropriate scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on that scenario. The server analyzes the practice results, generates feedback, and sends it to the device.
[1514] Matching and communication support
[1515] The user types, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users then begin to converse, and the server monitors the exchange and, if necessary, requests intervention from experts.
[1516] Prompt Sentence Examples
[1517] 1. If you want to use sentiment analysis with the sentiment engine:
[1518] "If a user says, 'I'm feeling stressed because things haven't been going well at work lately,' how does the emotion engine respond?"
[1519] 2. To request a priority assessment and expert intervention:
[1520] "If a user expresses very depressed feelings consecutively, how does the server determine the urgency and notify an expert?"
[1521] 3. If you would like to provide a practice scenario:
[1522] "If a user wants to practice an interview, how does the system generate scenarios and provide feedback?"
[1523] 4. If you wish to match and communicate with other users:
[1524] “If a user inputs that they want to talk to other users with similar goals of reintegration, how does the system match them with the right users and support their interactions?”
[1525] As described above, the present invention can support users' mental health with high accuracy through real-time emotion analysis using an emotion engine and the provision of counseling messages and practice scenarios using a generative AI model.
[1526] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1527] Step 1:
[1528] The user inputs data by voice or text. An example of input data is the text "I'm feeling stressed lately because things aren't going well at work." In the case of voice input, the device uses speech recognition software (e.g., Google Speech-to-Text API) to convert the voice into text data. The input for this step is the user's voice or text, and the output is text data.
[1529] Step 2:
[1530] The device sends text data to the server. Specifically, the device sends the text data entered to the server via an HTTP request. For example, text data is sent in the form of {"user_input": "Recently, things haven't been going well at work and I'm feeling stressed"}. The input of this step is text data, and the output is the data sent to the server.
[1531] Step 3:
[1532] The server passes the received input data to the emotion engine. The emotion engine uses an emotion analysis model using Hugging Face's Transformers library to analyze the text data and determine the user's emotion. For example, it determines "very strong stress." The input of this step is the received text data, and the output is the emotion analysis result.
[1533] Step 4:
[1534] The server generates a counseling message using a generative AI model (for example, OpenAI's GPT-3) based on the analysis results of the emotion engine. The prompt text is "The user's stress level is very high. Please provide some advice on how to relax." The AI model generates a message such as "Taking deep breaths can help you relax. Let's try it." The input for this step is the emotion analysis result, and the output is a counseling message.
[1535] Step 5:
[1536] The server sends the generated counseling message to the user's device. Specifically, it sends the following message using an HTTP response: {"counseling_message": "Try taking a deep breath to relax."} The device then displays the received message to the user. The input to this step is the counseling message, and the output is the message displayed on the user's device.
[1537] Step 6:
[1538] The server determines the urgency level based on the analysis results of the emotion engine. For example, if the user repeatedly expresses "very depressed feelings," the server determines the urgency level as "high." The input of this step is the emotion analysis result, and the output is the urgency level determination result.
[1539] Step 7:
[1540] If the urgency is determined to be "high," the server requests a response from an expert. An alert is sent to the expert, for example, a notification saying, "User A's urgency is high. Please respond immediately." The expert receives the alert and contacts the user to provide counseling or appropriate treatment. The input to this step is the urgency determination result, and the output is an alert notification to the expert.
[1541] Step 8:
[1542] When a user wishes to practice an everyday situation, they input a request, for example, "I would like to practice an interview." The server receives the request, analyzes the user's emotional state using an emotion engine, and generates an appropriate practice scenario. The generated scenario is sent to the device, and the user practices interactively based on that scenario. The server analyzes the practice results, generates feedback, and sends it to the device. The input to this step is the user's practice request, and the output is the practice scenario and feedback.
[1543] Step 9:
[1544] When a user wishes to communicate with other users, they input, for example, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users begin a conversation, and the server monitors the exchange, requesting intervention from experts if necessary. The input to this step is the user's communication desire, and the output is the matching results and conversation monitoring.
[1545] (Application example 2)
[1546] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1547] In recent years, with the spread of electronic payment services, users have been experiencing increasing stress and anxiety during transactions and payments. This mental burden not only worsens the user experience, but may also lead to transaction failures and increased security risks. Therefore, there is a need for a system that can monitor users' emotions in real time during electronic payment services and provide appropriate support to improve the user experience.
[1548] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a user's voice or text input, means for analyzing the input to determine the user's emotion, means for generating a counseling message based on the emotion, means for providing the counseling message to the user, and means for monitoring the user's emotion when conducting a transaction or payment and taking measures to reduce stress. In this way, by understanding the user's emotional state in real time and providing a counseling message as needed, stress during a transaction or payment can be reduced, allowing the user to use the service with peace of mind.
[1549] The "means for receiving user voice or text input" is an interface for acquiring voice or text data uttered by the user and processing it within the system.
[1550] The "means for analyzing the input and determining the user's emotions" refers to an algorithm or engine that analyzes the received voice or text data and determines the user's emotional state (e.g., stress, joy, surprise, etc.) based on that analysis.
[1551] The "means for generating a counseling message based on the emotion" is a processing unit for automatically generating an appropriate message of advice or comfort according to the determined emotional state of the user.
[1552] The "means for providing the counseling message to the user" is a module for displaying, playing, or transmitting the generated counseling message to the user's device.
[1553] "Means to monitor users' emotions when making transactions or payments and take measures to reduce stress" refers to a function that monitors users' emotional state in real time during electronic payments or transactions and provides guidelines and support messages to reduce stress and anxiety.
[1554] The "means for determining the degree of urgency based on the emotion" is a system for assessing the seriousness of the user's emotional state and determining whether an emergency response is required, if necessary.
[1555] The "means for requesting a response from an expert when the urgency is high" is a module that sends an alert to an expert such as a counselor or a specialist doctor when a high urgency is determined, urging them to intervene quickly.
[1556] The "means for receiving a user's practice request for an everyday situation" is an interface for receiving a request when a user wishes to practice a specific scenario (for example, an interview, a presentation, etc.).
[1557] The "means for generating a practice scenario based on the request" is a module for automatically creating a specific practice scenario in response to a request from a user.
[1558] The "means for interactively practicing with the user based on the practice scenario" is an interface that allows the user to practice while interacting with the system based on the generated scenario.
[1559] The "means for generating feedback based on the practice results" is a system that analyzes the user's practice results and automatically creates and provides feedback such as areas for improvement and results.
[1560] This invention is a system that analyzes a user's voice or text input, determines the user's emotions, and provides appropriate counseling messages. Furthermore, by using an emotion engine, it is possible to monitor the user's emotions in real time when conducting transactions or payments, and take necessary measures. Such a system aims to reduce the user's mental burden and provide a better user experience.
[1561] System Configuration
[1562] The system includes the following components:
[1563] 1. User Device
[1564] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[1565] 2. Server
[1566] It receives input data and performs processes such as emotion analysis, generating counseling messages, monitoring emotions during transactions and payments, and determining urgency. It also uses an emotion engine to learn the user's emotional patterns and reflect emotional changes in real time.
[1567] 3. Emotion Engine
[1568] It recognizes emotions by analyzing the user's voice, text input, and non-verbal elements (facial expressions, tone of voice, etc.). It learns emotional patterns and updates them in real time to detect subtle changes in the user's emotional state.
[1569] System Operation
[1570] In the present invention, the system operates using the following means.
[1571] 1. Handling User Input
[1572] The user inputs data by voice or text, and the device sends the data to the server. For example, the user says, "This transaction is very stressful." The device converts this data into text and sends it to the server.
[1573] 2. Emotion Engine Processing
[1574] The emotion engine in the server recognizes the user's emotions based on the input data. For example, it not only determines that the user is "feeling stressed," but also detects that the user is "very stressed" based on facial expressions and tone of voice.
[1575] 3. Generating counseling messages
[1576] The server generates a counseling message based on the analysis results of the emotion engine. For example, if it determines that the user is feeling extremely stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[1577] 4. Emotion monitoring and countermeasures during transactions and payments
[1578] The server monitors the user's emotions in real time when trading or making payments, and if it determines that the user is feeling stressed, it takes measures to reduce stress. For example, if a user enters "This transaction is very stressful" while trading, the server analyzes this and generates an advice message such as "Try taking deep breaths to relax."
[1579] Specific examples and examples of prompts for generative AI models
[1580] Example 1:
[1581] If the user enters "This transaction is very stressful," the server processes it as follows:
[1582] Emotion analysis result: NEGATIVE
[1583] Score: 0.95
[1584] Counseling message: "You seem to be stressed. Take a deep breath and try to relax a bit."
[1585] Example prompt sentence:
[1586] Create a program that analyzes a user's voice or text input, checks whether the user is feeling anxious or stressed, and provides an appropriate counseling message. Use Transformers as the emotion analysis model, and display a stress relief message if the input indicates a negative emotion.
[1587] The above is a specific embodiment for carrying out this invention. By combining this system with an emotion engine, it is possible to monitor the user's mental health with greater precision and provide appropriate counseling and transaction / payment support.
[1588] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1589] Step 1:
[1590] The user provides voice or text input. For example, the user might say, "This transaction is very stressful." This input is captured by the terminal and converted into text data.
[1591] Step 2:
[1592] The device sends the converted text data to the server, which receives the data and prepares it for sentiment analysis. The input is the user's voice or text data, and the output is the text data sent to the server.
[1593] Step 3:
[1594] The server uses an emotion engine to analyze the incoming data. Specifically, it applies a generative AI model such as Transformers to determine the user's emotion. The input is text data, and the output is an emotion (e.g., NEGATIVE) and its score (e.g., 0.95).
[1595] Step 4:
[1596] The server runs an algorithm that generates a counseling message based on the analysis results. Specifically, if the emotion is "NEGATIVE" and the score is high, a counseling message such as "You seem to be feeling stressed. Take a deep breath and try to relax a bit" is created. The input is the emotion analysis result and score, and the output is the generated counseling message.
[1597] Step 5:
[1598] The server sends the generated counseling message to the terminal. The terminal receives this message and displays or plays it aloud to the user. The input is the counseling message sent from the server, and the output is the message provided in a format that the user can view.
[1599] Step 6:
[1600] The server monitors users' emotions in real time and continuously provides stress-reducing measures during transactions and settlements. For example, if a transaction is prolonged or if a particular frustration persists, it generates additional messages such as "Take a short break to refresh yourself." The input is continuous emotion analysis data, and the output is support messages provided at any time.
[1601] Step 7:
[1602] The system determines the urgency of the emotion, and if the server deems it necessary, sends an alert to an expert. For example, if a user repeatedly expresses very strong, depressed emotions, the system will contact an expert and ask for their response. The input is the emotion analysis result and urgency assessment data, and the output is an alert to the expert and a request for their response.
[1603] Step 8:
[1604] Based on the user's request, the server generates a practice scenario for an everyday situation. If the user requests "I want to practice an interview," the server generates an appropriate scenario and sends it to the terminal. The input is the user's practice request, and the output is the generated practice scenario.
[1605] Step 9:
[1606] The user performs interactive practice based on the generated scenario, which involves the system analyzing the user's responses and providing appropriate feedback. The input is the user's interaction data, and the output is the feedback provided by the system.
[1607] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1608] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1609] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1610] [Fourth embodiment]
[1611] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1612] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1613] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1614] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1615] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1616] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1617] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1618] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1619] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1620] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1621] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1622] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1623] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1624] The present invention is a system that analyzes a user's voice or text input, determines their emotions, and provides appropriate counseling messages. This system is made up of the following components:
[1625] Components
[1626] 1. User Device:
[1627] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[1628] 2. Server:
[1629] It receives input data, analyzes emotions, generates counseling messages, and determines the level of urgency. It also generates practice scenarios, matches users with other users, and monitors communication content.
[1630] 3. Experts:
[1631] Based on requests from the server, appropriate advice and treatment will be provided to users, including psychologists, psychiatrists, and counselors.
[1632] Program processing
[1633] Handling User Input
[1634] The user inputs data by voice or text, and the device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server receives this data, performs emotional analysis, and determines that the user is "feeling stressed."
[1635] Counseling message generation
[1636] The server generates a counseling message based on the results of emotion analysis. For example, if it determines that the user is feeling stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[1637] Urgency assessment and response request
[1638] The server continuously monitors the emotion data and determines the level of urgency. If the level is deemed high, it requests a specialist to respond. For example, if a user continuously expresses depressed feelings, the server sends an alert to the specialist. The specialist receives the alert, contacts the user, and provides appropriate counseling or treatment.
[1639] Providing practice scenarios
[1640] When a user wishes to practice an everyday situation, for example, they make a request such as "I would like to practice an interview." The server receives this request and generates an appropriate interview scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on the scenario. The server analyzes the practice results, generates feedback, sends it to the device, and provides it to the user.
[1641] Matching and communication support
[1642] When a user wishes to communicate with other users, they can enter, for example, "I want to talk to someone who is also trying to reintegrate into society." The server searches and selects users in the same position from its database and performs matching. The matching results are sent to the user's device and notified. The user then begins a conversation with the matched person, and the system monitors the conversation. If necessary, the server requests expert intervention.
[1643] Specific examples
[1644] 1. Sentiment analysis and counseling message provision:
[1645] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this speech into text and sends it to the server. The server determines that the user is "feeling stressed," and generates a message such as "Take a deep breath to relax," which is sent to the device. The device then displays this message to the user.
[1646] 2. Urgency assessment and expert intervention:
[1647] The user repeatedly types "I feel heavy." The server analyzes this and determines that the situation is urgent. It sends an alert to an expert, who then contacts the user and provides counseling.
[1648] 3. Providing practice scenarios and feedback:
[1649] The user requests, "I want to practice an interview." The server generates an appropriate scenario and sends it to the device. The user practices interactively, and the server analyzes the results and sends feedback.
[1650] 4. Matching and communication support:
[1651] The user types, "I want to talk to people who are also trying to reintegrate into society." The server matches suitable users and notifies the device. The users begin a conversation, and the system monitors the content. If necessary, it requests intervention from an expert.
[1652] As described above, the system of the present invention can monitor the mental health state of the user and provide appropriate counseling and support for rehabilitation into society.
[1653] The processing flow will be explained below.
[1654] AI that listens to your heart
[1655] User Input and Sentiment Analysis
[1656] Step 1:
[1657] The user speaks or texts into the device, saying, "I've been feeling stressed lately because things haven't been going well at work."
[1658] Step 2:
[1659] The terminal converts the voice input into text data and transmits the text data to the server.
[1660] Step 3:
[1661] A server receives the text data and runs a sentiment analysis model (e.g., using natural language processing).
[1662] Step 4:
[1663] The server determines from the user's text that he or she is "feeling stressed."
[1664] Counseling message generation
[1665] Step 5:
[1666] The server generates a counseling message for stress reduction (e.g., "Take a deep breath and you'll feel more relaxed").
[1667] Step 6:
[1668] The server sends the generated counseling message to the terminal.
[1669] Step 7:
[1670] The terminal displays a counseling message to the user.
[1671] AI in harmony with specialists
[1672] Urgency assessment and response request
[1673] Step 1:
[1674] The server monitors the user's daily emotional data.
[1675] Step 2:
[1676] The server analyzes the emotional data and determines whether the person is experiencing a continuous state of depression.
[1677] Step 3:
[1678] The server determines the level of urgency and, if deemed high, sends an alert to an expert.
[1679] Step 4:
[1680] A specialist receives the alert from the server and contacts the user to provide counseling or treatment.
[1681] AI provides a platform for social reintegration
[1682] Providing practice scenarios
[1683] Step 1:
[1684] The user inputs a request to the terminal saying, "I want to practice for an interview."
[1685] Step 2:
[1686] The server receives the user's request and generates an interview scenario.
[1687] Step 3:
[1688] The server sends the generated scenario to the terminal.
[1689] Step 4:
[1690] The terminal displays a scenario, and the user practices the interview in an interactive format.
[1691] Step 5:
[1692] When the user completes the exercise, the terminal transmits the exercise results to the server.
[1693] Providing Feedback
[1694] Step 6:
[1695] The server analyzes the practice results and generates a feedback message.
[1696] Step 7:
[1697] The server sends a feedback message to the terminal, which displays it to the user.
[1698] AI that provides an empathetic community
[1699] Matching with user requests
[1700] Step 1:
[1701] The user inputs their request into the terminal, saying, "I would like to talk to someone who is also aiming to return to society."
[1702] Step 2:
[1703] The server analyzes users in the same position from the database and matches them with the appropriate users.
[1704] Step 3:
[1705] The server sends the matching results to the device.
[1706] Step 4:
[1707] The user begins communicating with the matched person (e.g., "I see you're in the same situation, ○○. Let's both do our best.").
[1708] Communications monitoring and support
[1709] Step 5:
[1710] The server monitors interactions within the community.
[1711] Step 6:
[1712] If necessary, the server calls for expert intervention.
[1713] Through the above steps, this system can comprehensively support the user's mental health maintenance and social reintegration.
[1714] Example 1
[1715] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1716] Conventional counseling systems struggle to quickly and accurately analyze users' emotions and provide appropriate counseling messages. Furthermore, delays in assessing the level of urgency and providing appropriate expert intervention can lead to a deterioration in the user's mental health. Furthermore, they lack sufficient practice in everyday situations and support for communication between users, preventing them from effectively supporting users' reintegration into society.
[1717] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1718] In this invention, the server includes means for receiving a user's voice or text input, means for converting the input into text data, means for transmitting the text data to the server, means for analyzing the text data to determine the user's emotion, means for generating a counseling message using a generative AI model based on the result of the emotion analysis, means for transmitting the counseling message to a user terminal, and means for displaying the counseling message to the user on the user terminal. This makes it possible to quickly and accurately analyze the user's emotion and effectively support appropriate counseling and social reintegration.
[1719] "Voice or text input" refers to the form of speech or written information that a user uses to communicate their feelings or requests to a system.
[1720] "Text data" refers to data obtained by converting voice input into text format, or text information entered directly by the user.
[1721] "Server" refers to the central computing system that receives and analyzes voice or text data, generates counseling messages, determines urgency, etc.
[1722] "Sentiment analysis" refers to the process of analyzing a user's text data to determine the user's emotional state.
[1723] A "generative AI model" refers to an artificial intelligence algorithm, primarily a natural language generation model, that generates appropriate messages based on the results of user sentiment analysis.
[1724] "Counseling message" refers to a message of advice or encouragement provided to a user that is generated based on the results of an analysis of the user's emotions.
[1725] "Urgency" refers to an index that evaluates the urgency of the user's emotional state and indicates whether a prompt response is required.
[1726] "Professional" refers to a psychologist, psychiatrist, counselor, or other person qualified to assist users with their mental health.
[1727] A "practice scenario" refers to a simulation scenario generated by the server when a user wishes to practice a particular situation (such as an interview).
[1728] "Feedback" refers to evaluations and advice provided to users based on the results of their practice.
[1729] "Matching" refers to the process by which a server selects and connects a suitable partner when a user wishes to communicate with another user.
[1730] "User terminal" refers to a device (smartphone, tablet, PC, etc.) used to input voice or text and receive and display counseling messages and feedback from the server.
[1731] The "HTTPS protocol" refers to a communication protocol for secure data communication over the Internet.
[1732] "Natural language processing tool" refers to a software tool (e.g., spaCy, NLTK, etc.) used to analyze a user's text data.
[1733] The present invention is a system that analyzes a user's voice or text input, determines their emotions, and provides appropriate counseling messages. This can support the user's mental health and effectively assist them in their social reintegration. A specific embodiment of this system is described below.
[1734] System Configuration
[1735] This system consists of user devices, servers, and expert components. Specific hardware components for user devices include smartphones, tablets, and PCs. Servers are high-performance servers (e.g., cloud virtual machines) with various software installed.
[1736] User terminal: A device that receives voice or text input, converts it into text data, and sends it to a server.
[1737] Server: This is a critical computing system that performs sentiment analysis, generates counseling messages, and determines urgency. Specifically, it uses natural language processing tools (e.g., spaCy, NLTK) and machine learning libraries (e.g., TensorFlow, PyTorch).
[1738] Experts: Psychologists, psychiatrists, counselors, or other qualified individuals available to support users' mental health.
[1739] Example of operation
[1740] Processing voice or text input
[1741] The user inputs data by voice or text. The user's device receives this data, and in the case of voice input, it converts it into text data using the Google Speech-to-Text API. For example, if the user says, "Recently, things haven't been going well at work and I'm feeling stressed," the content is converted into text. This text data is sent to the server using the HTTPS protocol.
[1742] Emotion analysis
[1743] The server analyzes the received text data and determines the user's emotions. Using a natural language processing tool (e.g., spaCy), it determines that the user is "feeling stressed." The emotion analysis algorithm is trained using a machine learning library (e.g., TensorFlow).
[1744] Counseling message generation
[1745] Based on the results of the sentiment analysis, a generative AI model (e.g., GPT-3) is used to generate an appropriate counseling message. For example, a message such as "Take a deep breath and you'll feel more relaxed. Try it out." This message is sent to the user's device and immediately displayed to the user.
[1746] Urgency assessment and expert intervention
[1747] The server continuously monitors the user's emotional data and determines the level of urgency. If the user repeatedly inputs "I feel heavy," the level of urgency is determined to be high. If the level of urgency is high, the server sends an alert to an expert, who then contacts the user and provides counseling.
[1748] Providing practice scenarios
[1749] When a user requests "I want to practice for an interview," the server generates an appropriate practice scenario. The scenario contains specific questions, and the contents of the scenario are sent to the user's terminal. The user practices interactively, and the server analyzes the results and generates feedback.
[1750] Prompt Sentence Examples
[1751] Emotion Analysis Prompt: Enter "I've been feeling stressed lately because things haven't been going well at work."
[1752] Urgency determination prompt: Enter "I feel heavy" repeatedly.
[1753] Practice Scenario Prompt: Type "I want to practice interviewing."
[1754] Matching prompt: Enter "I'd like to talk to someone who is also trying to reintegrate into society."
[1755] As described above, the system of the present invention is realized using advanced technology to support the user's mental health and effectively assist in their reintegration into society.
[1756] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1757] Step 1:
[1758] The user inputs data by voice or text. For example, they might say, "I've been feeling stressed lately because things haven't been going well at work." This input becomes the starting point for the system to process the data.
[1759] Input: User voice or text
[1760] Output: Raw audio or text data
[1761] Step 2:
[1762] The device converts the voice input into text data. Specifically, it uses voice recognition software (e.g., Google Speech-to-Text API) to convert the voice to text. If the input is text, it proceeds to the next step.
[1763] Input: Audio data
[1764] Output: Text data
[1765] Step 3:
[1766] The device sends the text data to the server, specifically using the HTTPS protocol to transfer the data securely.
[1767] Input: Text data
[1768] Output: Text data is sent to the server
[1769] Step 4:
[1770] The server analyzes the received text data and determines the user's emotions. Specifically, it analyzes the text using natural language processing tools (e.g., spaCy, NLTK) and classifies the emotions using machine learning models (e.g., TensorFlow).
[1771] Input: Text data
[1772] Output: Sentiment analysis result (e.g., user's emotion is "stress")
[1773] Step 5:
[1774] The server generates a counseling message using a generative AI model (e.g., GPT-3) based on the results of emotion analysis. For example, it automatically generates a message such as, "Try taking deep breaths to relax. Give it a try."
[1775] Input: Sentiment analysis results
[1776] Output: Counseling message
[1777] Step 6:
[1778] The server sends the generated counseling message to the terminal. The server sends the message using the HTTPS protocol, and the terminal receives it.
[1779] Input: Counseling message
[1780] Output: Counseling message forwarded to terminal
[1781] Step 7:
[1782] The terminal displays the counseling message to the user. Specifically, the notification function is used to make the message immediately visible to the user.
[1783] Input: Counseling message
[1784] Output: A message displayed to the user
[1785] Step 8:
[1786] The server continuously monitors the emotion data and determines the level of urgency. If the user repeatedly inputs "I feel heavy," the level of urgency is determined to be high.
[1787] Input: Continuous emotion data
[1788] Output: Urgency judgment result (e.g., high urgency)
[1789] Step 9:
[1790] If the emergency is deemed high, the server will request a response from an expert. For example, an alert email will be sent to a psychologist requesting immediate action.
[1791] Input: Urgency assessment result
[1792] Output: An alert email is sent to the expert
[1793] Step 10:
[1794] When a user requests "I want to practice for an interview," the server generates an appropriate practice scenario, which includes specific questions and sends the contents of the scenario to the terminal.
[1795] Input: Practice Request
[1796] Output: Practice scenario
[1797] Step 11:
[1798] The user practices interactively, and the server analyzes the results and generates feedback, such as "Your speech is clear and good."
[1799] Input: Practice results
[1800] Output: Feedback
[1801] (Application example 1)
[1802] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1803] In recent years, there has been an increasing demand for counseling systems to maintain users' mental health. However, existing systems lack sufficient analysis of users' emotions and assessment of the level of urgency, making it difficult to respond quickly and effectively. Furthermore, as security risks increase, systems that can provide immediate security responses based on users' emotional states are needed. Furthermore, more advanced responses using generative AI models are needed to improve the quality of practice and feedback for everyday situations requested by users. A system that can solve these issues and comprehensively support users' mental health and safety is needed.
[1804] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1805] In this invention, the server includes: means for receiving and analyzing voice or text input and determining a user's emotion; means for generating a counseling message using a generative AI model based on the determined emotion; means for providing the generated counseling message to the user; means for determining the urgency based on the emotion and requesting a response from an expert; means for issuing a security alert based on the result of the emotion analysis; means for cooperating with a security device when an alert is issued; means for receiving a practice request for an everyday situation from the user, generating a practice scenario based on the request, practicing interactively with the user, and generating feedback based on the practice results; and means for generating prompt sentences based on emotion data and generating feedback using a generative AI model. This enables the provision of counseling messages that are responsive to the user's emotional state and rapid response in high-urgency situations, and enables the immediate issuance of alerts even in high-security risk situations and the provision of comprehensive safety support in cooperation with other devices.
[1806] "Voice or text input" refers to the means by which a user provides voice or text data to a system.
[1807] "Means for determining emotion" refers to a method or device for analyzing received voice or text input and classifying the user's current emotional state as "positive," "negative," "neutral," or the like.
[1808] The "means for generating a counseling message" refers to a method or device for generating a message including appropriate advice or advice based on the result of the user's emotion determination.
[1809] A "generative AI model" is an artificial intelligence model trained based on large amounts of data, and is a means used for emotion analysis and generating counseling messages.
[1810] The "means for determining the degree of urgency" refers to a method or device for analyzing the emotional state of the user and evaluating and determining the degree of urgency of that state.
[1811] "Means for requesting a response from an expert" refers to a method or device for requesting intervention from an expert such as a psychologist or psychiatrist when the emergency is determined to be high.
[1812] "Means for issuing a security alert" refers to a method or device for issuing an alert when it is determined that the user's safety may be threatened based on the results of emotion analysis.
[1813] "Means for coordinating with security devices" refers to a method or device for coordinating with other security devices (such as surveillance cameras or intrusion detection systems) when issuing an alarm.
[1814] The "means for receiving a practice request" refers to a method or device for receiving a request for practicing a daily situation desired by a user.
[1815] "Means for generating a practice scenario" refers to a method or device for automatically generating a scenario according to the practice desired by the user.
[1816] The "means for conducting interactive practice" refers to a method or device for allowing a user and a system to have a dialogue based on a generated practice scenario.
[1817] The "means for generating feedback" refers to a method or device for analyzing the results of a user's practice and providing advice or suggestions for improvement based on the results.
[1818] A "prompt" is text data that is input into a generative AI model, providing information for the AI's response or generated message.
[1819] Components
[1820] 1. User Device:
[1821] A device that allows a user to input voice or text. It includes smartphones, tablets, personal computers, head-mounted displays, etc. The user terminal is responsible for sending input data to a server and receiving messages from the server.
[1822] 2. Server:
[1823] It receives input data, analyzes emotions, generates counseling messages, and determines the level of urgency. It also uses generative AI models to generate practice scenarios and collaborate with other security devices.
[1824] Program processing
[1825] Handling User Input
[1826] When a user inputs data by voice or text, the user device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The user device converts this data into text data and sends it to the server. The server receives this data and performs emotion analysis.
[1827] Emotion analysis
[1828] The server uses an emotion analysis model to classify the user's emotional state as either "positive," "negative," or "neutral." For example, if a user enters "I'm feeling stressed," the server will determine this as "negative."
[1829] Counseling message generation
[1830] The server generates appropriate counseling messages based on the analysis results, such as "Take a deep breath to relax," using a generative AI model, and sends the messages to the user's device.
[1831] Urgency assessment and response request
[1832] The server continuously monitors the emotion data and determines the level of urgency. If the level of urgency is deemed high, it will request a response from an expert. For example, if the user continuously expresses sadness, the server will send an alert to an expert. It will also issue a security alarm and connect to security devices.
[1833] Providing practice scenarios
[1834] When a user wants to practice an everyday situation, for example, they request, "I would like to practice an interview." The server receives this request and uses a generative AI model to generate an appropriate interview scenario. The generated scenario is sent to the user's device, where the user practices interactively. The server analyzes the practice results, generates feedback based on the prompt sentence, and sends it to the user's device.
[1835] Specific examples
[1836] 1. Sentiment analysis and counseling message provision:
[1837] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this speech into text and sends it to the server. The server determines that the user is "feeling stressed," and generates a message such as "Take a deep breath to relax," and sends it to the device.
[1838] 2. Urgency assessment and expert intervention:
[1839] The user repeatedly types "I feel heavy." The server analyzes this and determines that the situation is urgent. It sends an alert to an expert, who then contacts the user and provides counseling.
[1840] 3. Providing practice scenarios and feedback:
[1841] The user requests, "I want to practice for an interview." The server generates an appropriate scenario and sends it to the terminal. The user practices interactively, and the server analyzes the results and sends feedback. An example of a prompt sentence is, "How do you relieve nervousness during an interview?"
[1842] In this way, a system is realized in which the user terminal and server work together to support the user's mental health while also responding immediately to security risks.
[1843] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1844] Step 1:
[1845] The user provides input via voice or text.
[1846] Input: User voice or text
[1847] Output: Audio data or text data
[1848] How it works: The user uses a device such as a smartphone or head-mounted display (HMD) to input voice commands, for example, "I'm feeling stressed because things haven't been going well at work lately."
[1849] Step 2:
[1850] The device converts the voice input into text.
[1851] Input: Audio data
[1852] Output: Text data
[1853] How it works: The device's microphone records the user's voice and converts it into text using speech recognition software (e.g., Google Speech Recognition).
[1854] Step 3:
[1855] The terminal transmits the text data to the server.
[1856] Input: Text data
[1857] Output: Data sent to the server
[1858] How it works: Your device sends the converted text data to a server over your internet connection.
[1859] Step 4:
[1860] The server receives the text data and performs sentiment analysis.
[1861] Input: Text data
[1862] Output: Emotion judgment result (positive, negative, neutral, etc.)
[1863] How it works: The server inputs the received text data into a sentiment analysis model (e.g., BERT), which then classifies the user's sentiment. For example, "I'm feeling stressed" is classified as "negative."
[1864] Step 5:
[1865] The server generates a counseling message based on the emotion determination result.
[1866] Input: Emotion determination result
[1867] Output: Counseling message
[1868] How it works: The server uses the emotion determination results to input prompts into the generative AI model, generating an appropriate counseling message. For example, a message like "Take a deep breath and you'll be able to relax" is generated.
[1869] Step 6:
[1870] The server sends a counseling message to the user terminal.
[1871] Input: Counseling message
[1872] Output: Transmitted data
[1873] Operation: The server sends the generated counseling message to the user terminal.
[1874] Step 7:
[1875] The user terminal displays a counseling message to the user.
[1876] Input: Counseling message
[1877] Output: Display message
[1878] Operation: The user terminal displays the counseling message received on the screen. For example, the user sees a message on the terminal screen saying, "Take a deep breath and you'll be able to relax."
[1879] Step 8:
[1880] The server continuously monitors the emotional data and determines the level of urgency.
[1881] Input: Continuous emotion data
[1882] Output: Urgency (low, medium, high)
[1883] How it works: The server monitors the emotional data continuously sent by the user and determines the urgency based on certain criteria.
[1884] Step 9:
[1885] If the emergency is deemed high, the server will request a specialist to respond.
[1886] Input: Urgency (High)
[1887] Output: Alert to experts
[1888] How it works: When the server detects a high level of urgency, it sends an alert to registered experts and requests them to take action.
[1889] Step 10:
[1890] Additionally, the server will issue security alerts as needed and work in conjunction with security devices.
[1891] Input: High-urgency emotion data
[1892] Output: Security alerts and data sent to linked devices
[1893] How it works: The server issues security alarms and works in conjunction with security devices such as surveillance cameras and intrusion detection systems to provide comprehensive safety responses.
[1894] Step 11:
[1895] When a user submits a practice request, the server generates a practice scenario for an everyday situation.
[1896] Input: User's practice request
[1897] Output: Practice scenario
[1898] How it works: When a user requests, "I want to practice for an interview," the server uses the generative AI model to generate an appropriate practice scenario and sends it to the user's device.
[1899] Step 12:
[1900] The server analyzes the practice results, generates feedback, and sends it to the user's terminal.
[1901] Input: Practice result data
[1902] Output: Feedback message
[1903] How it works: The user performs the exercise and sends the results to the server. The server then uses the generative AI model to generate a prompt and sends feedback to the user's device. For example, the server might provide feedback such as, "It would be better if you relaxed your gaze more."
[1904] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1905] This system analyzes a user's voice or text input, determines the user's emotions, and provides appropriate counseling messages. Furthermore, by combining it with an emotion engine, it can recognize subtle changes in the user's emotional state and provide more accurate support.
[1906] Components
[1907] 1. User Device:
[1908] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[1909] 2. Server:
[1910] It receives input data and performs processes such as emotion analysis, generating counseling messages, and determining the level of urgency. It also uses an emotion engine to learn the user's emotional patterns and reflect changes in emotions in real time.
[1911] 3. Experts:
[1912] Based on requests from the server, appropriate advice and treatment will be provided to users, including psychologists, psychiatrists, and counselors.
[1913] 4. Emotion Engine:
[1914] It recognizes emotions by analyzing the user's voice, text input, and non-verbal elements (facial expressions, tone of voice, etc.). It learns emotional patterns and updates them in real time to detect subtle changes in the user's emotional state.
[1915] Program processing
[1916] Handling User Input
[1917] The user inputs data by voice or text, and the device sends the data to the server. For example, the user might say, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this data into text data and sends it to the server. The server receives this data and analyzes it using an emotion engine.
[1918] Emotion engine processing
[1919] The emotion engine in the server recognizes the user's emotions based on the input data. For example, it not only determines that the user is "feeling stressed," but also detects that the user is "very stressed" based on facial expressions and tone of voice.
[1920] Counseling message generation
[1921] The server generates a counseling message based on the analysis results of the emotion engine. For example, if it determines that the user is feeling extremely stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[1922] Urgency assessment and response request
[1923] The server determines the level of urgency based on the data obtained from the emotion engine. For example, if a user repeatedly expresses very strong feelings of depression, the server will request a response from a specialist. The specialist will receive the alert and contact the user to provide counseling or appropriate treatment.
[1924] Providing practice scenarios
[1925] When a user wants to practice an everyday situation, for example, they request, "I want to practice an interview." The server receives this request and generates an appropriate interview scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on the scenario. The emotion engine monitors and analyzes the user's reactions and provides feedback.
[1926] Matching and communication support
[1927] When a user wishes to communicate with other users, they can enter, for example, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users and sends the results to the device. The users then begin a conversation, and the server monitors the exchange, requesting expert intervention if necessary.
[1928] Specific examples
[1929] 1. Emotion analysis using an emotion engine and provision of counseling messages:
[1930] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server uses its emotion engine to determine that the user is under "very high stress," and generates a message such as, "Take a deep breath and you'll be able to relax. Try it out," which is sent to the device. The device then displays this message to the user.
[1931] 2. Urgency assessment and expert intervention:
[1932] A user repeatedly types, "I'm feeling very depressed." The server analyzes the situation using an emotion engine, and if it determines that the situation is urgent, it sends an alert to a specialist. The specialist receives the alert and contacts the user to provide counseling or treatment.
[1933] 3. Providing practice scenarios and feedback:
[1934] The user requests, "I want to practice for an interview." The server uses an emotion engine to analyze the user's emotional state and generate an appropriate scenario. The scenario is sent to the device, and the user practices. The server analyzes the practice results, generates feedback, and sends it to the device.
[1935] 4. Matching and communication support:
[1936] The user types, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users then begin to converse, and the system monitors the content and requests intervention from experts if necessary.
[1937] As described above, by combining the emotion engine, the system of the present invention can monitor the user's mental health state with higher accuracy and provide appropriate counseling and support for rehabilitation into society.
[1938] The processing flow will be explained below.
[1939] AI that listens to your heart
[1940] User Input and Sentiment Analysis
[1941] Step 1:
[1942] The user speaks or texts into the device, saying, "I've been feeling stressed lately because things haven't been going well at work."
[1943] Step 2:
[1944] The terminal converts the voice input into text data and transmits the text data to the server.
[1945] Step 3:
[1946] The server receives the text data and requests the emotion engine to analyze it.
[1947] Step 4:
[1948] The emotion engine determines from the user's text that they are "feeling stressed."
[1949] Step 5:
[1950] The emotion engine also analyzes non-verbal data such as facial expressions and tone of voice to assess the intensity of stress (if determined to be "very high stress").
[1951] Counseling message generation
[1952] Step 6:
[1953] The server generates a counseling message based on the analysis results of the emotion engine. For example, it generates a message such as, "Take a deep breath and you'll feel more relaxed. Try it out."
[1954] Step 7:
[1955] The server sends the generated counseling message to the terminal.
[1956] Step 8:
[1957] The terminal displays a counseling message to the user.
[1958] AI in harmony with specialists
[1959] Urgency assessment and response request
[1960] Step 1:
[1961] The server continuously monitors the user's daily emotional data.
[1962] Step 2:
[1963] The server determines the level of urgency based on the data obtained from the emotion engine. For example, if the user repeatedly expresses "feeling very depressed," it will determine that the level of urgency is high.
[1964] Step 3:
[1965] If the server determines that the situation is urgent, it will send an alert to an expert.
[1966] Step 4:
[1967] A specialist receives the alert from the server and contacts the user to provide counseling or treatment.
[1968] AI provides a platform for social reintegration
[1969] Providing practice scenarios
[1970] Step 1:
[1971] The user inputs a request to the terminal saying, "I want to practice for an interview."
[1972] Step 2:
[1973] The server receives the user's request and generates an interview scenario.
[1974] Step 3:
[1975] The emotion engine analyzes the user's emotional state and reflects the emotional data in the scenario.
[1976] Step 4:
[1977] The server sends the generated scenario to the terminal.
[1978] Step 5:
[1979] The terminal displays a scenario, and the user practices the interview in an interactive format.
[1980] Providing Feedback
[1981] Step 6:
[1982] When the user completes the exercise, the terminal transmits the exercise results to the server.
[1983] Step 7:
[1984] The server uses an emotion engine to analyze the practice results and generate a feedback message, such as "You did very well. You could do better if you answered with more confidence."
[1985] Step 8:
[1986] The server sends a feedback message to the terminal, which displays it to the user.
[1987] AI that provides an empathetic community
[1988] Matching with user requests
[1989] Step 1:
[1990] The user inputs their request into the terminal, saying, "I would like to talk to someone who is also aiming to return to society."
[1991] Step 2:
[1992] The server receives the user's request and analyzes it using the emotion engine.
[1993] Step 3:
[1994] The server searches and selects users in the same position from the database and works with the emotion engine to make appropriate matches.
[1995] Step 4:
[1996] The server sends the matching results to the device.
[1997] Step 5:
[1998] The terminal notifies the user of the matching result, and the user starts communication with the matched person.
[1999] Communications monitoring and support
[2000] Step 6:
[2001] The server monitors interactions within the community using an emotion engine.
[2002] Step 7:
[2003] If necessary, the server calls for expert intervention.
[2004] Through these steps, this system can comprehensively support users in maintaining their mental health and reintegrating into society. By combining it with an emotion engine, it is possible to detect even subtle changes in the user's emotional state, enabling more accurate counseling and support to be provided.
[2005] Example 2
[2006] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2007] In modern society, increasing stress and anxiety have created a demand for psychological support. However, many users have limited opportunities to receive appropriate counseling immediately. Furthermore, conventional counseling methods often fail to capture subtle changes in a user's emotions in real time, making it difficult to provide immediate, appropriate support. While there are counseling support systems that use voice or text input, few of them offer sufficient emotional analysis accuracy, real-time performance, urgency assessment, or expert intervention. This has created a demand for systems that can effectively support users' mental health.
[2008] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2009] In this invention, the server includes means for receiving a user's voice or text input, means for converting the input into text data and transmitting it to the server, means for analyzing the input to determine the user's emotion, means for generating a counseling message based on the emotion, and means for providing the counseling message to the user. This makes it possible to analyze the user's voice or text input in real time, quickly capture subtle changes in the user's emotional state, and generate and provide an appropriate counseling message. Furthermore, in cases of high urgency, it is possible to request a response from an expert, thereby enabling rapid expert intervention.
[2010] "User" refers to an individual or organization that uses the system.
[2011] "Voice input" refers to data that is converted into digital form from what a user says.
[2012] "Text input" refers to data in which a user inputs textual information using a keyboard or other input device.
[2013] "Device" means a device used by a user to input voice or text, including, for example, a smartphone, tablet, or computer.
[2014] A "server" refers to a high-performance computer system that receives and analyzes data sent by users.
[2015] "Emotion engine" refers to a system that analyzes a user's emotions from voice, text, and non-verbal data.
[2016] "Counseling message" refers to a message containing advice or support for the user, generated based on the analysis results of the emotion engine.
[2017] "Urgency" refers to an index that evaluates the level of urgency of a user's emotional state or health condition.
[2018] "Expert" refers to a person or organization qualified to provide appropriate advice or treatment for a user's mental health, such as a psychologist, psychiatrist, or counselor.
[2019] "Response request" refers to an action in which the system requests an expert to respond to the user's condition.
[2020] "Practice scenario" refers to an interactive scenario provided for a user to simulate a particular situation.
[2021] "Feedback" refers to evaluations and advice provided based on the results of the user's scenario practice.
[2022] "Matching" refers to the process of appropriately pairing users based on their purpose and status.
[2023] The present invention provides a system for supporting the mental health of a user by analyzing the user's voice or text input, determining the user's emotions, and providing appropriate counseling messages. Specific implementation methods are described in detail below.
[2024] Components
[2025] User terminal
[2026] A user terminal is a device for inputting voice or text. Terminals include smartphones, tablets, and PCs. The user terminal transmits the input data to a server and receives messages from the server.
[2027] server
[2028] The server has the central function of receiving and analyzing data sent by users. The received data is input into the emotion engine, which analyzes the user's emotions. Based on the analysis results, it generates an appropriate counseling message and sends it to the user's device. It also determines the level of urgency, requests expert assistance, provides practice scenarios and feedback, and matches users and supports communication.
[2029] Emotion Engine
[2030] The emotion engine is a system that analyzes the user's emotions using non-verbal data such as voice, text, facial expression analysis, and tone of voice. Based on the analysis results, the emotion engine detects subtle changes in the user's emotional state and reflects them in real time.
[2031] Hardware and software used
[2032] Speech Recognition Software: Uses the Google Speech-to-Text API to convert user voice input into text data.
[2033] Sentiment analysis model: We use a sentiment analysis model using the Hugging Face Transformers library.
[2034] Generative AI model: OpenAI's GPT-3 is used to generate counseling messages and practice scenarios.
[2035] Specific examples
[2036] Emotion analysis using an emotion engine and provision of counseling messages
[2037] The user says to the device, "Recently, things haven't been going well at work and I'm feeling stressed." The device converts this into text data and sends it to the server. The server uses its emotion engine to determine that the user is feeling "very stressed," and generates a message such as, "Take a deep breath and you'll be able to relax. Try it out," and sends it to the device. The device then displays this message to the user.
[2038] Urgency assessment and expert intervention
[2039] A user repeatedly types, "I'm feeling very depressed." The server analyzes the situation using an emotion engine, and if it determines that the situation is urgent, it sends an alert to a specialist. The specialist receives the alert and contacts the user to provide counseling or treatment.
[2040] Providing practice scenarios and feedback
[2041] The user requests, "I want to practice for an interview." The server uses an emotion engine to analyze the user's emotional state and generate an appropriate scenario. The generated scenario is sent to the device, and the user engages in interactive practice based on that scenario. The server analyzes the practice results, generates feedback, and sends it to the device.
[2042] Matching and communication support
[2043] The user types, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users then begin to converse, and the server monitors the exchange and, if necessary, requests intervention from experts.
[2044] Prompt Sentence Examples
[2045] 1. If you want to use sentiment analysis with the sentiment engine:
[2046] "If a user says, 'I'm feeling stressed because things haven't been going well at work lately,' how does the emotion engine respond?"
[2047] 2. To request a priority assessment and expert intervention:
[2048] "If a user expresses very depressed feelings consecutively, how does the server determine the urgency and notify an expert?"
[2049] 3. If you would like to provide a practice scenario:
[2050] "If a user wants to practice an interview, how does the system generate scenarios and provide feedback?"
[2051] 4. If you wish to match and communicate with other users:
[2052] “If a user inputs that they want to talk to other users with similar goals of reintegration, how does the system match them with the right users and support their interactions?”
[2053] As described above, the present invention can support users' mental health with high accuracy through real-time emotion analysis using an emotion engine and the provision of counseling messages and practice scenarios using a generative AI model.
[2054] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2055] Step 1:
[2056] The user inputs data by voice or text. An example of input data is the text "I'm feeling stressed lately because things aren't going well at work." In the case of voice input, the device uses speech recognition software (e.g., Google Speech-to-Text API) to convert the voice into text data. The input for this step is the user's voice or text, and the output is text data.
[2057] Step 2:
[2058] The device sends text data to the server. Specifically, the device sends the text data entered to the server via an HTTP request. For example, text data is sent in the form of {"user_input": "Recently, things haven't been going well at work and I'm feeling stressed"}. The input of this step is text data, and the output is the data sent to the server.
[2059] Step 3:
[2060] The server passes the received input data to the emotion engine. The emotion engine uses an emotion analysis model using Hugging Face's Transformers library to analyze the text data and determine the user's emotion. For example, it determines "very strong stress." The input of this step is the received text data, and the output is the emotion analysis result.
[2061] Step 4:
[2062] The server generates a counseling message using a generative AI model (for example, OpenAI's GPT-3) based on the analysis results of the emotion engine. The prompt text is "The user's stress level is very high. Please provide some advice on how to relax." The AI model generates a message such as "Taking deep breaths can help you relax. Let's try it." The input for this step is the emotion analysis result, and the output is a counseling message.
[2063] Step 5:
[2064] The server sends the generated counseling message to the user's device. Specifically, it sends the following message using an HTTP response: {"counseling_message": "Try taking a deep breath to relax."} The device then displays the received message to the user. The input to this step is the counseling message, and the output is the message displayed on the user's device.
[2065] Step 6:
[2066] The server determines the urgency level based on the analysis results of the emotion engine. For example, if the user repeatedly expresses "very depressed feelings," the server determines the urgency level as "high." The input of this step is the emotion analysis result, and the output is the urgency level determination result.
[2067] Step 7:
[2068] If the urgency is determined to be "high," the server requests a response from an expert. An alert is sent to the expert, for example, a notification saying, "User A's urgency is high. Please respond immediately." The expert receives the alert and contacts the user to provide counseling or appropriate treatment. The input to this step is the urgency determination result, and the output is an alert notification to the expert.
[2069] Step 8:
[2070] When a user wishes to practice an everyday situation, they input a request, for example, "I would like to practice an interview." The server receives the request, analyzes the user's emotional state using an emotion engine, and generates an appropriate practice scenario. The generated scenario is sent to the device, and the user practices interactively based on that scenario. The server analyzes the practice results, generates feedback, and sends it to the device. The input to this step is the user's practice request, and the output is the practice scenario and feedback.
[2071] Step 9:
[2072] When a user wishes to communicate with other users, they input, for example, "I want to talk to people who are also trying to reintegrate into society." The server uses an emotion engine to match suitable users from the database and sends the results to the device. The users begin a conversation, and the server monitors the exchange, requesting intervention from experts if necessary. The input to this step is the user's communication desire, and the output is the matching results and conversation monitoring.
[2073] (Application example 2)
[2074] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2075] In recent years, with the spread of electronic payment services, users have been experiencing increasing stress and anxiety during transactions and payments. This mental burden not only worsens the user experience, but may also lead to transaction failures and increased security risks. Therefore, there is a need for a system that can monitor users' emotions in real time during electronic payment services and provide appropriate support to improve the user experience.
[2076] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving a user's voice or text input, means for analyzing the input to determine the user's emotion, means for generating a counseling message based on the emotion, means for providing the counseling message to the user, and means for monitoring the user's emotion when conducting a transaction or payment and taking measures to reduce stress. In this way, by understanding the user's emotional state in real time and providing a counseling message as needed, stress during a transaction or payment can be reduced, allowing the user to use the service with peace of mind.
[2077] The "means for receiving user voice or text input" is an interface for acquiring voice or text data uttered by the user and processing it within the system.
[2078] The "means for analyzing the input and determining the user's emotions" refers to an algorithm or engine that analyzes the received voice or text data and determines the user's emotional state (e.g., stress, joy, surprise, etc.) based on that analysis.
[2079] The "means for generating a counseling message based on the emotion" is a processing unit for automatically generating an appropriate message of advice or comfort according to the determined emotional state of the user.
[2080] The "means for providing the counseling message to the user" is a module for displaying, playing, or transmitting the generated counseling message to the user's device.
[2081] "Means to monitor users' emotions when making transactions or payments and take measures to reduce stress" refers to a function that monitors users' emotional state in real time during electronic payments or transactions and provides guidelines and support messages to reduce stress and anxiety.
[2082] The "means for determining the degree of urgency based on the emotion" is a system for assessing the seriousness of the user's emotional state and determining whether an emergency response is required, if necessary.
[2083] The "means for requesting a response from an expert when the urgency is high" is a module that sends an alert to an expert such as a counselor or a specialist doctor when a high urgency is determined, urging them to intervene quickly.
[2084] The "means for receiving a user's practice request for an everyday situation" is an interface for receiving a request when a user wishes to practice a specific scenario (for example, an interview, a presentation, etc.).
[2085] The "means for generating a practice scenario based on the request" is a module for automatically creating a specific practice scenario in response to a request from a user.
[2086] The "means for interactively practicing with the user based on the practice scenario" is an interface that allows the user to practice while interacting with the system based on the generated scenario.
[2087] The "means for generating feedback based on the practice results" is a system that analyzes the user's practice results and automatically creates and provides feedback such as areas for improvement and results.
[2088] This invention is a system that analyzes a user's voice or text input, determines the user's emotions, and provides appropriate counseling messages. Furthermore, by using an emotion engine, it is possible to monitor the user's emotions in real time when conducting transactions or payments, and take necessary measures. Such a system aims to reduce the user's mental burden and provide a better user experience.
[2089] System Configuration
[2090] The system includes the following components:
[2091] 1. User Device
[2092] A device that allows a user to input voice or text. It includes smartphones, tablets, and PCs. The user device sends input data to a server and receives messages from the server.
[2093] 2. Server
[2094] It receives input data and performs processes such as emotion analysis, generating counseling messages, monitoring emotions during transactions and payments, and determining urgency. It also uses an emotion engine to learn the user's emotional patterns and reflect emotional changes in real time.
[2095] 3. Emotion Engine
[2096] It recognizes emotions by analyzing the user's voice, text input, and non-verbal elements (facial expressions, tone of voice, etc.). It learns emotional patterns and updates them in real time to detect subtle changes in the user's emotional state.
[2097] System Operation
[2098] In the present invention, the system operates using the following means.
[2099] 1. Handling User Input
[2100] The user inputs data by voice or text, and the device sends the data to the server. For example, the user says, "This transaction is very stressful." The device converts this data into text and sends it to the server.
[2101] 2. Emotion Engine Processing
[2102] The emotion engine in the server recognizes the user's emotions based on the input data. For example, it not only determines that the user is "feeling stressed," but also detects that the user is "very stressed" based on facial expressions and tone of voice.
[2103] 3. Generating counseling messages
[2104] The server generates a counseling message based on the analysis results of the emotion engine. For example, if it determines that the user is feeling extremely stressed, it generates a message such as, "Take a deep breath and you'll be able to relax. Try it out." The generated message is sent to the device and provided to the user.
[2105] 4. Emotion monitoring and countermeasures during transactions and payments
[2106] The server monitors the user's emotions in real time when trading or making payments, and if it determines that the user is feeling stressed, it takes measures to reduce stress. For example, if a user enters "This transaction is very stressful" while trading, the server analyzes this and generates an advice message such as "Try taking deep breaths to relax."
[2107] Specific examples and examples of prompts for generative AI models
[2108] Example 1:
[2109] If the user enters "This transaction is very stressful," the server processes it as follows:
[2110] Emotion analysis result: NEGATIVE
[2111] Score: 0.95
[2112] Counseling message: "You seem to be stressed. Take a deep breath and try to relax a bit."
[2113] Example prompt sentence:
[2114] Create a program that analyzes a user's voice or text input, checks whether the user is feeling anxious or stressed, and provides an appropriate counseling message. Use Transformers as the emotion analysis model, and display a stress relief message if the input indicates a negative emotion.
[2115] The above is a specific embodiment for carrying out this invention. By combining this system with an emotion engine, it is possible to monitor the user's mental health with greater precision and provide appropriate counseling and transaction / payment support.
[2116] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2117] Step 1:
[2118] The user provides voice or text input. For example, the user might say, "This transaction is very stressful." This input is captured by the terminal and converted into text data.
[2119] Step 2:
[2120] The device sends the converted text data to the server, which receives the data and prepares it for sentiment analysis. The input is the user's voice or text data, and the output is the text data sent to the server.
[2121] Step 3:
[2122] The server uses an emotion engine to analyze the incoming data. Specifically, it applies a generative AI model such as Transformers to determine the user's emotion. The input is text data, and the output is an emotion (e.g., NEGATIVE) and its score (e.g., 0.95).
[2123] Step 4:
[2124] The server runs an algorithm that generates a counseling message based on the analysis results. Specifically, if the emotion is "NEGATIVE" and the score is high, a counseling message such as "You seem to be feeling stressed. Take a deep breath and try to relax a bit" is created. The input is the emotion analysis result and score, and the output is the generated counseling message.
[2125] Step 5:
[2126] The server sends the generated counseling message to the terminal. The terminal receives this...
Claims
1. means for receiving a user's voice or text input; means for analyzing the input to determine a user's emotion; means for generating a counseling message based on the emotion; means for providing said counseling message to a user; A system including:
2. means for determining a degree of urgency based on the emotion; a means for requesting a specialist to respond when the urgency is high; The system of claim 1 further comprising:
3. means for receiving a user's practice request for an everyday situation; means for generating a practice scenario based on the request; means for interactively practicing with a user based on the practice scenario; means for generating feedback based on the practice results; The system of claim 1 , comprising:
4. means for matching other users based on the user's requests; a means for allowing the matched users to communicate with each other; means for monitoring the content of said communications and requesting expert intervention as necessary; The system of claim 1 , comprising:
5. A means for providing the system according to any one of claims 1 to 4 for a business.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A