System

The mental care system addresses the lack of tailored support by using user authentication, generative AI, and speech synthesis to provide personalized dialogue, effectively reducing mental stress and isolation.

JP2026017982APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119043
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current systems fail to provide tailored mental care and emotional support to individuals of different age groups, inadequately addressing mental stress and feelings of isolation.

Method used

A mental care system that includes user authentication, question-asking, answer-receiving, and generative AI model-based dialogue content generation, with speech synthesis and data recording for personalized support.

Benefits of technology

Enables real-time detection of mental health risks and provides personalized dialogue to alleviate stress and fatigue, improving user interaction accuracy over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017982000001_ABST
    Figure 2026017982000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A mental care system comprising: means for receiving certification information of a user and performing certification; means for asking a question for confirming a feeling or a state of the user and receiving an answer to the question; means for analyzing the answer of the user and generating an optimal dialogue content for the user; means for providing the generated dialogue content to the user using a speech synthesis technology; and means for recording the dialogue content and a response of the user and learning with a generative AI model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, people of all ages, from young people to the elderly, suffer from mental stress and feelings of isolation. In these circumstances, there is a need for systems that can provide appropriate dialogue and support to address the unique challenges faced by each generation. For example, it is necessary to provide mental care and emotional support tailored to the needs of young people who are struggling with relationships with parents or friends, middle-aged people who have lost self-esteem at work or at home, and elderly people who tend to become isolated due to the need for caregiving. However, current systems are not able to adequately address these challenges, and more effective solutions are needed. [Means for solving the problem]

[0005] The present invention provides a mental care system that includes a means for receiving and authenticating a user's authentication information, a means for asking questions to ascertain the user's feelings and state and receiving the answers, and a means for analyzing the user's answers and generating optimal dialogue content for the user. This system includes a means for providing the generated dialogue content to the user using speech synthesis technology, and further includes a means for recording the dialogue content and the user's responses and learning them using a generative AI model, thereby enabling the system to continuously provide personalized dialogue to the user. This makes it possible to detect the user's mental health risks in advance and take appropriate measures.

[0006] "User authentication information" refers to information used to verify the identity of a user, such as a user name and password, that the user enters when accessing a system.

[0007] The "means for performing authentication" is a means for verifying the identity of a user based on the user's authentication information.

[0008] The "means for asking questions" is a means for presenting specific questions to the user in order to understand the user's feelings and state.

[0009] "Means for receiving answers" refers to the means for receiving and acquiring answers entered by users in response to questions.

[0010] A "generative AI model" is a model that uses artificial intelligence technology to analyze user input information and generate appropriate dialogue content.

[0011] "Means for generating dialogue content" refers to a means for utilizing a generative AI model to create dialogue content that matches the user's state and mood.

[0012] "Speech synthesis technology" is a technology that converts text information into speech and provides it to the user as natural speech.

[0013] "Means for providing using speech synthesis technology" refers to means for conveying the generated dialogue content to the user as voice.

[0014] The "means for recording the dialogue content and user responses" refers to a means for saving the dialogue history and user responses in the system in a database or the like.

[0015] "Means for learning with a generative AI model" refers to a means for updating a generative AI model based on recorded data to improve the quality of the next interaction. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention is a mental care system that provides mental support to users. This system has a series of functions, such as user authentication, grasping of feelings and conditions, generation of dialogue content, dialogue provision by voice synthesis, and recording and learning of dialogue data.

[0038] Overall system overview

[0039] The system is primarily composed of three elements: a server, a terminal, and a user. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[0040] 1. Initial Setup and User Authentication

[0041] 1. Start the device

[0042] When the terminal is started, it presents the user with a login screen, where the user enters their authentication information (username and password), which is then sent from the terminal to the server.

[0043] 2. Authentication Process

[0044] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[0045] 2. Understanding the user's status

[0046] 1. Posing the Question

[0047] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[0048] 2. Receiving and sending responses

[0049] The user inputs their feelings and state into the terminal, and the response is sent from the terminal to the server.

[0050] 3. Analysis of responses

[0051] The server analyzes the user's responses using a generative AI model, and based on the results of this analysis, evaluates the user's current mental state.

[0052] 3. Generating and providing conversation content

[0053] 1. Conversation content generation

[0054] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's mood and state.

[0055] 2. Sending conversation content

[0056] The generated dialogue content is transmitted from the server to the terminal.

[0057] 3. Provision by voice synthesis

[0058] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in a voice that is conveyed to the user in a natural conversational style.

[0059] 4. Data recording and learning

[0060] 1. Data recording

[0061] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which records them in a database.

[0062] 2. Training the generative AI model

[0063] The server uses the recorded data to update the generative AI model to improve the accuracy of the next interaction. This learning allows the system to better respond to the user's individual needs.

[0064] Specific examples

[0065] If the user responds "I'm a little tired"

[0066] 1. Terminal: "How are you feeling today?"

[0067] 2. User: "I'm a little tired."

[0068] 3. The device sends this information to the server.

[0069] 4. The server uses the generative AI model to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[0070] 5. The device will read this aloud using speech synthesis technology.

[0071] 6. The user responds, "I want to go to the park."

[0072] 7. The device sends this response to the server.

[0073] 8. The server generates new dialogue such as "That's a good idea. Let's plan what we want to do at the park." and sends it to the device.

[0074] 9. The device will read this aloud using speech synthesis technology and the conversation will continue.

[0075] In this way, the system of the present invention responds to the user's feelings and state in real time and provides appropriate dialogue, thereby providing psychological support.

[0076] The processing flow will be explained below.

[0077] Step 1:

[0078] When the terminal is started, it displays a login screen to the user.

[0079] Step 2:

[0080] The user enters a username and password into the terminal.

[0081] Step 3:

[0082] The terminal sends the user's authentication information to the server.

[0083] Step 4:

[0084] The server checks the received authentication information against the database to verify that the user is a legitimate user.

[0085] Step 5:

[0086] The server sends the authentication result to the terminal. If the authentication is successful, an authentication success message is sent to the terminal and the user is successfully logged in. If the authentication is unsuccessful, an error message is sent to the terminal.

[0087] Step 6:

[0088] Once the login is successful, the terminal will ask the user questions to confirm their current state of mind and condition.

[0089] Step 7:

[0090] The user inputs their feelings and state into the terminal and responds.

[0091] Step 8:

[0092] The terminal sends the user's answer to the server.

[0093] Step 9:

[0094] The server uses a generative AI model to analyze the user's responses and evaluate the user's mental state.

[0095] Step 10:

[0096] The server generates dialogue content based on the analysis results using an AI model.

[0097] Step 11:

[0098] The server transmits the generated dialogue content to the terminal.

[0099] Step 12:

[0100] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in voice.

[0101] Step 13:

[0102] The user responds to the terminal.

[0103] Step 14:

[0104] The terminal sends the user's response to the server.

[0105] Step 15:

[0106] Based on the user's response, the server generates new dialogue content using a generative AI model and sends it to the device.

[0107] Step 16:

[0108] The terminal provides new dialogue content to the user using voice synthesis technology.

[0109] Step 17:

[0110] The terminal sequentially transmits the contents of the dialogue with the user and the user's responses to the server.

[0111] Step 18:

[0112] The server records the dialogue and the user's responses in a database.

[0113] Step 19:

[0114] The server uses the recorded data to update the generative AI model and learn to improve the accuracy of the next interaction.

[0115] Example 1

[0116] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0117] In modern society, many people suffer from mental stress and fatigue, and appropriate mental care is needed to alleviate these burdens. However, conventional mental care systems have had difficulty accurately grasping the user's mental state and providing dialogue content tailored to individual needs. In addition, the methods for providing dialogue content are limited, making it difficult to achieve natural and effective support.

[0118] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0119] In this invention, the server includes means for receiving and authenticating a user's authentication information, means for presenting questions to ascertain the user's mental state and receiving the answers, means for analyzing the user's answers using a generative AI model and generating appropriate dialogue content, means for providing the generated dialogue content to the user using speech synthesis technology, and means for recording the dialogue content and the user's responses and learning the dialogue content with the generative AI model, thereby enabling the provision of appropriate psychological support to the user in real time.

[0120] "User" refers to an individual who uses the system to receive psychological support.

[0121] "Authentication information" is information used to identify a user and authorize access to a system, and typically includes a username and password.

[0122] "Means" refers to a method, technique, or device for accomplishing a particular function.

[0123] "Mental state" refers to the user's emotional and psychological state, including stress and fatigue.

[0124] "Question" refers to a query posed by the system to ascertain the user's mental state.

[0125] A "generative AI model" refers to a model that uses machine learning and artificial intelligence technology to analyze data and automatically generate dialogue content.

[0126] "Analysis" refers to the process of evaluating the user's mental state based on their answers and deriving appropriate dialogue content.

[0127] "Dialogue content" refers to the advice and response text provided to the user.

[0128] "Speech synthesis technology" refers to technology for converting text data into speech.

[0129] "Recording" refers to the act of the system saving the content of interactions and responses with the user.

[0130] "Learning" refers to the process by which the system improves the generated AI model based on past data, improving the accuracy of the next interaction.

[0131] The present invention provides a mental care system that provides mental support to users. This system has a series of functions, such as user authentication, understanding of emotions and conditions, generating dialogue content, providing dialogue through voice synthesis, and recording and learning dialogue data.

[0132] Overall system overview

[0133] The system is primarily composed of three elements: a server, a terminal, and a user. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[0134] Hardware and software used

[0135] Server: Responsible for data processing and operation of the generative AI model, authenticating users, analyzing responses, and generating dialogue content.

[0136] Terminal: Presents questions to the user, receives answers from the user, and provides the generated dialogue content through voice synthesis.

[0137] Generative AI model: Uses artificial intelligence techniques to generate dialogue content based on user input data.

[0138] Speech synthesis technology: A speech synthesis engine is used to convert text data into speech.

[0139] Specific examples

[0140] If the user answers "I'm a little tired," the specific sequence of events is as follows:

[0141] 1. Terminal: Display the question to the user: "How are you feeling today?"

[0142] 2. User: Type "I'm a little tired."

[0143] 3. The device sends this information to the server.

[0144] 4. The server uses the generative AI model to analyze the answer "I'm a little tired" and assess the user's condition.

[0145] 5. Based on the analysis results, the server generates dialogue content such as, "We will provide you with some advice to help you relax."

[0146] 6. The server sends the generated dialogue content to the terminal.

[0147] 7. The device uses speech synthesis technology to read out the generated dialogue in a natural voice, saying, "We'll give you some advice to help you relax."

[0148] 8. The user may ask further questions about how to relax, in which case the device sends a response to the server, which again uses the generative AI model to generate appropriate dialogue.

[0149] When implementing this system, it is also important to consider the prompt sentences to generate prepared questions. For example, the following prompt sentences can be used in response to input from the user:

[0150] Example prompt sentence:

[0151] "I'm a little tired. How can I relax?"

[0152] The mental care system of the present invention provides mental support by grasping the user's feelings and state in real time and providing appropriate dialogue, allowing the user to easily receive advice and counseling to reduce stress and fatigue in daily life.

[0153] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0154] Program processing flow

[0155] Step 1: Initial configuration and user authentication

[0156] 1.1 Starting the terminal

[0157] When the terminal is started, it displays a login screen. The user enters a username and password. This authentication information is sent from the terminal to the server.

[0158] Input: Username, Password

[0159] Output: Sending authentication information

[0160] 1.2 Authentication process

[0161] The server checks the received authentication information against its database, checking the username and password to ensure the user is a valid user, and sending a success message if authentication is successful, or an error message if authentication is unsuccessful.

[0162] Input: Credentials

[0163] Data processing: Database collation

[0164] Output: Authentication success message or error message

[0165] 1.3 Displaying authentication results

[0166] The terminal receives the message from the server and displays the success or failure of the authentication to the user.

[0167] Input: Authentication success message or error message

[0168] Output: Display of authentication result

[0169] Step 2: Understanding the user's state

[0170] 2.1 Posing the Question

[0171] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[0172] Input: Login successful

[0173] Output: Question posed

[0174] 2.2 Receiving a response

[0175] The user inputs their feelings and state in response to the questions presented, and the answers are sent from the device to the server.

[0176] Input: User's answer

[0177] Output: Sending the answer

[0178] 2.3 Analysis of responses

[0179] The server inputs the received answers into a generative AI model for analysis. For example, if the user answers something like "I'm a little tired," it analyzes that information to evaluate the user's current mental state.

[0180] Input: User's answer

[0181] Data Computation: Analysis with Generative AI Models

[0182] Output: Mental state assessment

[0183] Step 3: Generate and provide conversation content

[0184] 3.1 Conversation content generation

[0185] Based on the evaluation results, the server uses a generative AI model to generate optimal dialogue content for the user. For example, if the user is evaluated as tired, the server generates dialogue content such as "We will give you advice on how to relax."

[0186] Input: Mental status assessment

[0187] Data Computation: Dialogue Content Generation with Generative AI Models

[0188] Output: Generated dialogue

[0189] 3.2 Transmission of dialogue content

[0190] The server transmits the generated dialogue content to the terminal.

[0191] Input: Generated dialogue

[0192] Output: Sending dialogue

[0193] 3.3 Provision by voice synthesis

[0194] The terminal converts the received dialogue content into voice using speech synthesis technology and provides it to the user.

[0195] Input: Generated dialogue

[0196] Data processing: voice synthesis

[0197] Output: Provides dialogue via voice

[0198] Step 4: Data recording and learning

[0199] 4.1 Data recording

[0200] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which then records them in a database.

[0201] Input: Dialogue content, user response

[0202] Output: Data recording

[0203] 4.2 Training generative AI models

[0204] The server uses the recorded data to update the generative AI model, which improves the accuracy of the next interaction.

[0205] Input: Recorded data

[0206] Data Computation: Learning Generative AI Models

[0207] Output: Updated generative AI model

[0208] Through the above processing steps, the mental care system of the present invention has the ability to grasp the user's feelings and state in real time and provide appropriate dialogue.

[0209] (Application example 1)

[0210] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0211] In modern society, the number of users who need mental support is increasing, but there is a lack of systems that can individually generate dialogue content and provide it in an appropriate format.In addition, while there is a demand for effective mental care using smart devices, there is currently no system that can provide appropriate dialogue and content based on the user's mood and state.This poses the problem that users cannot receive support that is appropriate for their own physical and mental state.

[0212] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0213] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, means for recording the dialogue content and the user's responses and learning them using a generative AI model, and means for providing dialogue content and content generated by AI based on the user's feelings and state. This enables the user to receive support in real time according to their own feelings and state.

[0214] "User credentials" are the identifying information provided by a user to log into a system.

[0215] "Mood" refers to the user's feelings and moods.

[0216] "Condition" refers to the physical or mental condition of a user.

[0217] "Questions" are questions that the system presents to ascertain the user's feelings and state of mind.

[0218] An "answer" is information that a user enters in response to a question.

[0219] "Analysis" refers to the processing of data to evaluate users' responses and understand their sentiments and state of mind.

[0220] "Dialogue content" refers to messages consisting of text or voice used in dialogue with a user.

[0221] "Speech synthesis technology" is a technology that converts text into a voice that sounds like a human voice.

[0222] "User response" is information that the user responds to the dialogue content presented by the system.

[0223] "Recording" means saving the dialogue content and user responses in a database.

[0224] A "generative AI model" is an artificial intelligence model that generates appropriate dialogue content based on user information.

[0225] "Content" refers to information material such as video, music, text, etc., provided to users.

[0226] "Dialogue content generated by AI based on emotions and state" refers to messages created by a generative AI model based on the user's emotions and state.

[0227] This invention is a content distribution system that provides psychological support to users. The system is mainly composed of three elements: a server, a terminal, and a user. The server is responsible for data processing and operation of the generative AI model, and the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[0228] Server configuration and functions

[0229] The server includes means for receiving and authenticating a user's authentication information, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for transmitting the generated dialogue content to the terminal, means for recording the dialogue content and the user's responses and learning using a generative AI model, and means for providing dialogue content and content generated by AI based on the user's feelings and state.

[0230] The server uses Python and Flask as the web framework. The generative AI model is implemented using the OpenAI API. It receives and analyzes user data to generate dialogue content, which is then sent to the device. Furthermore, the dialogue data with the user is recorded in a database and used as training data for the generative AI model.

[0231] Device configuration and functions

[0232] The terminal is a device operated by the user, such as a smartphone or tablet, that provides a variety of functions. The terminal displays a user authentication screen and sends authentication information from the user to the server. If authentication is successful, the terminal displays questions to confirm the user's feelings and state, and sends the user's answers to the server. The dialogue received from the server is read aloud using speech synthesis technology.

[0233] The voice synthesis technology uses the Google Text-to-Speech API, which allows the generated dialogue to be presented to the user in a natural conversational format. The device also displays and plays relaxing content such as videos and music.

[0234] User actions and responses

[0235] Users operate their terminals to log in to the system. After logging in, they can report their own feelings and state by answering questions posed by the system. The system generates and provides appropriate dialogue and content based on this information. Users can receive psychological support by using the dialogue and content provided.

[0236] Specific examples

[0237] When a user logs in to the system and responds, "I'm a little tired," the server sends a dialogue to the device, such as, "We've checked your condition. We'll give you some advice to help you relax." This dialogue is read aloud to the user using voice synthesis technology. If the user then responds, "I'd like to go to the park," the server generates a new dialogue, such as, "That's a good idea. Let's plan what we want to do at the park," and sends it to the device.

[0238] Prompt Sentence Examples

[0239] "User is a little fatigued. Please generate a dialogue that offers a support conversation."

[0240] As described above, the system of the present invention can provide psychological support in real time according to the user's feelings and condition.

[0241] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0242] Step 1:

[0243] When the terminal is started, a login screen is displayed. The user enters authentication information (username and password) which the terminal sends to the server. The input is the username and password, and the output is the authentication information sent to the server.

[0244] Step 2:

[0245] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal. The input is authentication information, and the output is an authentication success message or an error message.

[0246] Step 3:

[0247] If authentication is successful, the device asks the user a question to confirm their current state of mind. The question is in the form of "How are you feeling today?" The input is a successful authentication message from the server, and the output is the question displayed to the user.

[0248] Step 4:

[0249] The user inputs their feelings or state of mind into the terminal. For example, they reply, "I'm a little tired." The input is text describing the user's feelings or state, and the output is the transmission of the reply data to the server.

[0250] Step 5:

[0251] The server inputs the received user responses into a generative AI model for analysis. This analysis evaluates the user's feelings and state and generates appropriate dialogue content. The input is the user's response, and the output is the analysis results and the generated dialogue content. Specifically, it sends a prompt to the OpenAI API and receives a response.

[0252] Step 6:

[0253] The server sends the generated dialogue content to the terminal. A response message containing the generated dialogue content is output. The input is the analysis result of the AI ​​model, and the output is the dialogue content sent to the terminal.

[0254] Step 7:

[0255] The device uses speech synthesis technology to provide the received dialogue to the user. The speech synthesis technology used is the Google Text-to-Speech API. The input is the dialogue from the server, and the output is the result presented to the user as voice.

[0256] Step 8:

[0257] The user responds to the dialogue content heard by voice and inputs the response into the terminal. The input is the user's response text, and the output is the transmission of the response data to the server.

[0258] Step 9:

[0259] The server records the received user responses in a database and uses them as training data for the generative AI model. The input is the user response data, and the output is an updated AI model. Specific operations include recording to the database and retraining the generative AI model.

[0260] As described above, this system performs a series of processes, starting with user authentication, then understanding the user's feelings and state, generating dialogue content using a generative AI model, providing dialogue through voice synthesis, and finally recording data and learning the model. Through this process, it is possible to provide individualized support according to the user's state.

[0261] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0262] This invention is a mental care system that provides mental support to users, and by combining it with an emotion engine, it is possible to recognize the user's emotional state in real time and to engage in dialogue and respond based on that. The system has a series of functions, including user authentication, understanding of feelings and state, emotion recognition, dialogue content generation, dialogue provision through voice synthesis, and dialogue data recording and learning.

[0263] Overall system overview

[0264] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion engine has the function of recognizing the user's emotional state in real time. The user is the entity that interacts with the system through the terminal.

[0265] 1. Initial Setup and User Authentication

[0266] 1. Start the device

[0267] When the terminal is started, it presents the user with a login screen, where the user enters their authentication information (username and password), which is then sent from the terminal to the server.

[0268] 2. Authentication Process

[0269] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[0270] 2. Understanding user state and emotion recognition

[0271] 1. Posing the Question

[0272] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[0273] 2. Receiving and sending responses

[0274] The user inputs their feelings and state into the terminal, and the response is sent from the terminal to the server.

[0275] 3. Emotion recognition

[0276] The emotion engine recognizes emotions from the user's voice and text. The emotion engine analyzes the user's tone of voice and expressions to determine their emotional state.

[0277] 4. Analysis of responses and emotional information

[0278] The server uses a generative AI model to analyze the user's responses and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[0279] 3. Generating and providing conversation content

[0280] 1. Conversation content generation

[0281] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's state of mind and condition, taking into account the emotional information recognized by the emotion engine.

[0282] 2. Sending conversation content

[0283] The generated dialogue content is transmitted from the server to the terminal.

[0284] 3. Provision by voice synthesis

[0285] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in a voice that is conveyed to the user in a natural conversational style.

[0286] 4. Data recording and learning

[0287] 1. Data recording

[0288] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which records them in a database.

[0289] 2. Training the generative AI model

[0290] The server uses the recorded data to update the generative AI model to improve the accuracy of the next interaction. This learning allows the system to better respond to the user's individual needs.

[0291] Specific examples

[0292] If the user responds "I'm a little tired"

[0293] 1. Terminal: "How are you feeling today?"

[0294] 2. User: "I'm a little tired."

[0295] 3. The device sends this information to the server.

[0296] 4. The emotion engine recognizes "fatigue" from the user's tone of voice.

[0297] 5. The server uses the information obtained from the generated AI model and emotion engine to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[0298] 6. The device will read this aloud using speech synthesis technology.

[0299] 7. The user responds, "I want to go to the park."

[0300] 8. The device sends this response to the server.

[0301] 9. The server generates new dialogue such as "That's a good idea. Let's plan what we want to do at the park." and sends it to the device.

[0302] 10. The device will read this aloud using speech synthesis technology and the conversation will continue.

[0303] In this way, the system of the present invention, which combines an emotion engine, recognizes the user's feelings and state in real time and provides appropriate dialogue to support the user emotionally.Furthermore, by generating dialogue content that reflects the user's emotional state, more personalized responses are possible.

[0304] The processing flow will be explained below.

[0305] Step 1:

[0306] When the terminal is started, it displays a login screen to the user.

[0307] Step 2:

[0308] The user enters a username and password into the terminal.

[0309] Step 3:

[0310] The terminal sends the user's authentication information to the server.

[0311] Step 4:

[0312] The server checks the received authentication information against the database to verify that the user is a legitimate user.

[0313] Step 5:

[0314] The server sends the authentication result to the terminal. If the authentication is successful, an authentication success message is sent to the terminal and the user is successfully logged in. If the authentication is unsuccessful, an error message is sent to the terminal.

[0315] Step 6:

[0316] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[0317] Step 7:

[0318] The user inputs their feelings and state into the terminal and responds.

[0319] Step 8:

[0320] The terminal sends the user's answer to the server.

[0321] Step 9:

[0322] The emotion engine recognizes emotions from the user's voice and text, analyzing the user's tone of voice and expressions to determine their emotional state.

[0323] Step 10:

[0324] The server uses a generative AI model to analyze the user's responses and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[0325] Step 11:

[0326] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's state of mind and condition, taking into account the emotional information recognized by the emotion engine.

[0327] Step 12:

[0328] The server transmits the generated dialogue content to the terminal.

[0329] Step 13:

[0330] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in voice.

[0331] Step 14:

[0332] The user responds to the terminal.

[0333] Step 15:

[0334] The terminal sends the user's response to the server.

[0335] Step 16:

[0336] Based on the user's response, the server generates new dialogue content using a generative AI model and sends it to the device.

[0337] Step 17:

[0338] The terminal provides new dialogue content to the user using voice synthesis technology.

[0339] Step 18:

[0340] The terminal sequentially transmits the contents of the dialogue with the user and the user's responses to the server.

[0341] Step 19:

[0342] The server records the dialogue and the user's responses in a database.

[0343] Step 20:

[0344] The server uses the recorded data to update the generative AI model and learn to improve the accuracy of the next interaction.

[0345] As a specific example, the flow when the user answers "I'm a little tired" is shown below.

[0346] Step 6:

[0347] Terminal: "How are you feeling today?"

[0348] Step 7:

[0349] User: "I'm a little tired."

[0350] Step 8:

[0351] The terminal sends the user's answer to the server.

[0352] Step 9:

[0353] The emotion engine recognizes "fatigue" from the user's tone of voice.

[0354] Step 10:

[0355] The server uses information obtained from the generative AI model and emotion engine to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[0356] Step 12:

[0357] The server transmits the generated dialogue content to the terminal.

[0358] Step 13:

[0359] The device will read this aloud using voice synthesis technology.

[0360] Step 14:

[0361] The user responds, "I want to go to the park."

[0362] Step 15:

[0363] The terminal sends the user's response to the server.

[0364] Step 16:

[0365] The server generates new dialogue content such as "That's a good idea. Let's plan what we want to do in the park," and sends it to the device.

[0366] Step 17:

[0367] The terminal provides new dialogue content to the user using voice synthesis technology.

[0368] In this way, the system of the present invention, which is combined with an emotion engine, recognizes the user's emotional state in real time and provides dialogue accordingly, thereby realizing more personalized psychological support.

[0369] Example 2

[0370] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0371] In modern society, the number of users suffering from mental stress and anxiety is increasing, but there are only a limited number of systems that provide immediate and personalized responses. Conventional systems have difficulty accurately recognizing the user's emotional state and generating appropriate dialogue based on that, making it impossible to provide adequate mental care. Furthermore, technology for recognizing the user's emotions through voice or text and generating dialogue that reflects this is immature, resulting in a lack of satisfactory support for users.

[0372] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0373] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to confirm the user's feelings and state and receiving the answers, means for analyzing the user's answers and recognizing and evaluating the user's emotions using an emotion engine, means for using a generative AI model to generate optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, and means for recording the dialogue content and the user's responses and learning with the generative AI model. This makes it possible to accurately recognize the user's emotional state in real time and generate natural dialogue based on that.

[0374] "User authentication information" refers to the identifying information a user needs to log in to a system, typically a username and password.

[0375] "Emotion engine" refers to an algorithm or software that recognizes and analyzes emotions from a user's voice or text data in real time.

[0376] A "generative AI model" is an artificial intelligence model that learns large amounts of data and generates dialogue content with users, and uses natural language processing technology.

[0377] "Speech synthesis technology" refers to technology for generating natural speech based on text data, enabling information to be conveyed to users through speech.

[0378] "Dialogue content" refers to the response from the system generated in response to the user's question or status, and includes information such as appropriate support and advice.

[0379] "Emotional state" refers to a psychological state that indicates the type and intensity of the emotion the user is currently experiencing, and includes, for example, joy, sadness, anger, fatigue, and the like.

[0380] "Database" refers to a system for storing and managing data necessary for system operation, such as user authentication information, interaction history, and learning data for generative AI models.

[0381] "Real-time" refers to the instantaneous acquisition and analysis of data, generating an appropriate response immediately.

[0382] "Login screen" refers to the screen interface that is displayed to a user to enter their authentication information.

[0383] "Personalized response" refers to providing responses and support that are customized to suit the individual conditions and needs of the user.

[0384] This invention is a mental care system that provides mental support to users, and by combining it with an emotion engine, it is possible to recognize the user's emotional state in real time and to engage in dialogue and respond based on that. This system has a series of functions, including user authentication, understanding of feelings and state, emotion recognition, dialogue content generation, dialogue provision through voice synthesis, and dialogue data recording and learning.

[0385] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion engine has the function of recognizing the user's emotional state in real time. The user is the entity that interacts with the system through the terminal.

[0386] Hardware and software used

[0387] Hardware: Terminals (PCs, smartphones), servers

[0388] Software: Emotion engine, generative AI model, speech synthesis engine, database management system

[0389] Program processing

[0390] Initial Setup and User Authentication

[0391] When the terminal is started, it displays a login screen to the user. The user enters their authentication information (username and password), and the terminal sends this information to the server. The server compares the received authentication information with a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[0392] Understanding user status and recognizing emotions

[0393] After successful login, the device asks the user questions to ascertain their current mood and state (e.g., "How are you feeling today?"). The user inputs their mood and state into the device, and the answer is sent from the device to the server. The server uses an emotion engine to recognize emotions from the user's voice and text. The emotion engine analyzes the user's tone of voice and expressions to determine their emotional state. The server uses a generative AI model to analyze the user's answers and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[0394] Conversation content generation and provision

[0395] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's feelings and state. The emotional information recognized by the emotion engine is also taken into consideration during this process. The generated dialogue content is sent from the server to the device, which then uses speech synthesis technology to provide the generated dialogue content to the user in a voice format. This voice is conveyed to the user in a natural conversational format.

[0396] Data recording and learning

[0397] The device continuously transmits the user's interactions and responses to the server, which records them in a database and uses the recorded data to update the generative AI model. This learning allows the system to better respond to the user's individual needs.

[0398] Additional examples of specific actions

[0399] For example, if the user responds, "I'm a little tired," the process goes something like this: The device asks, "How are you feeling today?" The user responds, "I'm a little tired," and the device sends that information to the server. The emotion engine recognizes "fatigue" from the user's tone of voice, and the server uses the information obtained from the generative AI model and the emotion engine to generate a dialogue message saying, "We've checked your condition and will provide you with some advice to help you relax." The device reads this aloud using speech synthesis technology. The user responds, "I'd like to go to the park," and the device sends that response to the server. The server generates a new dialogue message saying, "That's a good idea. Let's plan what we want to do at the park," sends it to the device, and the device continues reading this aloud using speech synthesis technology.

[0400] Prompt Sentence Examples

[0401] "Generate a dialogue for when the user responds that they are a little tired."

[0402] "Provide relaxing advice based on the user's emotional state."

[0403] In this way, the present invention can recognize the user's feelings and state in real time and provide appropriate dialogue in response to them, thereby providing psychological support.

[0404] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0405] Mental health care system processing steps

[0406] Initial Setup and User Authentication

[0407] Step 1:

[0408] Starting the terminal

[0409] When the terminal is started, it displays a login screen to the user. The input is powering on, and the output is displaying the login screen.

[0410] Step 2:

[0411] Enter your authentication information

[0412] A user enters authentication information (username and password) into a login screen. The input is the authentication information, and the output is its preparation for submission.

[0413] Step 3:

[0414] Sending authentication information

[0415] The terminal sends authentication information to the server. The input is the authentication information and the output is a request sent to the server.

[0416] Step 4:

[0417] Authentication verification

[0418] The server checks the received authentication information against a database to verify the user's validity. The input is the authentication information, and the output is the authentication result (success or failure).

[0419] Step 5:

[0420] Return and display of authentication results

[0421] The server returns the authentication result to the terminal, which displays it to the user. The input is the authentication result, and the output is an indication of authentication success or failure. If successful, login is complete.

[0422] Understanding user status and recognizing emotions

[0423] Step 6:

[0424] Posing the Question

[0425] The terminal displays a question to the user to confirm their mood or state. For example, "How are you feeling today?" The input is a successful login status, and the output is the display of the question.

[0426] Step 7:

[0427] Enter your answer

[0428] The user inputs their feelings and state. The input is the user's answer to a question, and the output is the answer ready to be sent.

[0429] Step 8:

[0430] Submit your answer

[0431] The terminal sends the user's answer to the server. The input is the user's answer, and the output is the response sent to the server.

[0432] Step 9:

[0433] Performing emotion recognition

[0434] The server inputs the user's response into the emotion engine to recognize the user's emotion. The input is the user's response, and the output is the recognized emotion information.

[0435] Step 10:

[0436] Emotional information analysis

[0437] The server performs analysis using the results of the emotion engine and the generative AI model. The input is emotion information and the generative AI model, and the output is an evaluation of the user's mental state.

[0438] Conversation content generation and provision

[0439] Step 11:

[0440] Conversation generation

[0441] The server uses a generative AI model to generate optimal dialogue for the user. The input is the mental state assessment, and the output is the generated dialogue.

[0442] Step 12:

[0443] Sending conversation transcripts

[0444] The generated dialogue content is sent from the server to the terminal. The input is the generated dialogue content, and the output is the content sent to the terminal.

[0445] Step 13:

[0446] Provided by voice synthesis

[0447] The terminal uses speech synthesis technology to communicate the dialogue content to the user by voice. The input is the generated dialogue content, and the output is the voice output to the user.

[0448] Data recording and learning

[0449] Step 14:

[0450] Data recording

[0451] The terminal sends the dialogue content and the user's response to the server, which records it in a database. The input is the dialogue content and the user's response, and the output is the record in the database.

[0452] Step 15:

[0453] Training generative AI models

[0454] The server uses the recorded data to train the generative AI model to improve the accuracy of the next interaction. The input is the recorded data, and the output is an updated generative AI model.

[0455] (Application example 2)

[0456] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0457] In modern work environments, especially in factories, long hours of monotonous work can cause mental stress for workers. This can lead to reduced work efficiency and increased risk of health problems. Conventional mental health systems face the challenge of being unable to recognize and respond to workers' emotional states in real time. The present invention aims to solve these challenges and provide an advanced mental health system for supporting the mental health of workers.

[0458] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0459] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, means for recording the dialogue content and the user's responses and learning it with a generative AI model, means for converting the user's voice into text using speech recognition technology, emotion recognition means for detecting the user's emotional state in real time, and means for dynamically changing the dialogue content based on the responses. This allows the system to recognize the emotional state of workers in real time and provide appropriate dialogue, thereby providing psychological support and improving work efficiency and maintaining health.

[0460] "User authentication" is the process of receiving user authentication information and verifying that the user is legitimate.

[0461] "Mood confirmation" is a process of asking questions to confirm the user's mood or mental state.

[0462] "Answer analysis" is the process of analyzing a user's answer and evaluating its content.

[0463] "Dialogue content generation" is a process of generating optimal dialogue content for the user based on the analysis results.

[0464] "Speech synthesis technology" is a technology that outputs text information as voice.

[0465] "Data logging" is the process of saving the dialogue and user responses.

[0466] A "generative AI model" is an artificial intelligence model that learns from data and generates appropriate dialogue content.

[0467] "Speech recognition technology" is a technology that converts voice data into text data.

[0468] "Emotion recognition means" is a technology that detects the user's emotional state in real time from their voice and facial expressions.

[0469] The "dynamic dialogue change means" is a process that changes the dialogue content based on the user's response.

[0470] The present invention relates to a mental care system that combines user authentication, emotional state confirmation, emotion recognition, dialogue content generation, voice synthesis, data recording, and learning. Specific embodiments of each element will be described below.

[0471] Overall system configuration

[0472] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion recognition engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion recognition engine has the function of recognizing the user's emotional state in real time, and the user is the entity that interacts with the system through the terminal.

[0473] server

[0474] The server centrally processes data and runs the generative AI model. The server receives the user's authentication information and checks it against the database to verify that the user is legitimate. If authentication is successful, the server returns a successful authentication message to the user.

[0475] Terminal

[0476] The terminal accepts input from the user and communicates with the server. When the user inputs their feelings or state into the terminal, this information is sent to the server. If login is successful, the terminal displays questions to the user to confirm their current feelings or state.

[0477] Emotion Recognition Engine

[0478] The emotion recognition engine works in conjunction with voice recognition technology to analyze the user's tone and expressions to recognize their emotional state, which is then sent to the server and used to assess the user's mental state.

[0479] User

[0480] The user interacts with the system through the device. When the user inputs their feelings and state, the information is sent to the server and analyzed by an emotion recognition engine. The server then uses a generative AI model to generate optimal dialogue content and send it to the device.

[0481] Detailed process flow

[0482] 1. User authentication:

[0483] The user enters a username and password into the terminal and sends them to the server.

[0484] The server checks the authentication information against a database to verify the user is a valid user.

[0485] If the authentication is successful, the server sends an authentication success message to the terminal, completing the login.

[0486] 2. Emotional validation and emotional recognition:

[0487] If authentication is successful, the terminal presents the user with questions to confirm their feelings and state of mind, and receives their answers.

[0488] The emotion recognition engine analyzes the user's voice and recognizes their emotional state in real time.

[0489] The server uses a generative AI model to analyze the user's responses and emotional information to assess their mental state.

[0490] 3. Generating and providing dialogue content:

[0491] The server uses a generative AI model based on the analysis results to generate optimal dialogue content.

[0492] The generated dialogue content is sent from the server to the terminal and provided to the user using voice synthesis technology.

[0493] 4. Data recording and learning:

[0494] The dialogue and the user's responses are sent to the server and recorded in a database.

[0495] The recorded data is used to train the generative AI model to improve the accuracy of the next interaction.

[0496] Hardware and software used

[0497] Server: Database, generative AI models (e.g., Hugging Face Transformers)

[0498] Devices: Smartphones, PCs, tablets

[0499] Emotion recognition engine: Speech recognition technology (e.g., Google Speech Recognition), emotion analysis (e.g., Sentiment Analysis Pipeline)

[0500] Speech synthesis: Pyttsx3 (e.g., a Python text-to-speech library)

[0501] Examples and prompts

[0502] Example 1:

[0503] Please enter your username: user123

[0504] Please enter your password: password123

[0505] Authentication successful. How are you feeling today?

[0506] (User speaks): "I'm a little tired."

[0507] The system responds: "You sound a little down. Let me know if there's anything I can do to help."

[0508] Example 2:

[0509] Please enter your username: sampleuser

[0510] Please enter your password: mypassword

[0511] Authentication successful. How are you feeling today?

[0512] (User speaks): "I feel great today."

[0513] The system responds: "Great! Keep it up!"

[0514] This allows the system to recognize the emotional state of workers in real time and provide appropriate dialogue to provide psychological support, thereby improving work efficiency and maintaining health.

[0515] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0516] Step 1:

[0517] User Authentication

[0518] Input: The user enters a username and password into the terminal.

[0519] Specific operation: The device sends these authentication information to the server.

[0520] Data processing / calculation: The server checks the authentication information against a database to verify that the user is legitimate.

[0521] Output: If the authentication is successful, the server returns an authentication success message to the terminal; if the authentication fails, it returns an error message.

[0522] Step 2:

[0523] Confirmation of feelings

[0524] Input: Receives information that authentication was successful.

[0525] Specific behavior: The device presents the user with a question to ascertain their current state of mind (e.g., "How are you feeling today?").

[0526] Data processing / calculation: Receive user responses via voice or text input.

[0527] Output: The device sends the user's answer to the server.

[0528] Step 3:

[0529] emotion recognition

[0530] Input: The server receives the user's answer and voice data.

[0531] How it works: The emotion recognition engine analyzes the user's tone of voice and expressions to recognize their emotional state in real time.

[0532] Data processing / calculation: The analysis results are sent to the server, and the emotion data is evaluated using a generative AI model.

[0533] Output: The recognized emotional state and analysis results are output.

[0534] Step 4:

[0535] Dialogue content generation

[0536] Input: Analyzed user emotional state and sentiment data.

[0537] Specific operation: The server uses a generative AI model to generate optimal dialogue content based on the user's emotions and feelings.

[0538] Data processing / calculation: The AI ​​model generates dialogue content based on emotion recognition results and sentiment data.

[0539] Output: The generated dialogue content is sent from the server to the terminal.

[0540] Step 5:

[0541] Dialogue provision

[0542] Input: The generated dialogue.

[0543] Specific operation: The device uses speech synthesis technology to provide the generated dialogue content to the user via voice.

[0544] Data processing / calculation: Converting text data into audio data.

[0545] Output: The dialogue is presented to the user via audio.

[0546] Step 6:

[0547] Data Recording

[0548] Input: Dialogue and user response data.

[0549] Specific operation: The dialogue content sent from the terminal to the server and the user's response data are recorded in a database.

[0550] Data processing / calculation: Recorded data is used as training data for generative AI models.

[0551] Output: Interaction data stored in a database and an updated AI model.

[0552] This allows the system to recognize the user's emotional state in real time and provide appropriate dialogue, thereby providing psychological support, improving work efficiency, and maintaining health.

[0553] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0554] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0555] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0556] [Second embodiment]

[0557] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0558] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0559] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0560] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0561] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0562] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0563] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0564] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0565] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0566] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0567] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0568] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0569] The present invention is a mental care system that provides mental support to users. This system has a series of functions, such as user authentication, grasping of feelings and conditions, generation of dialogue content, dialogue provision by voice synthesis, and recording and learning of dialogue data.

[0570] Overall system overview

[0571] The system is primarily composed of three elements: a server, a terminal, and a user. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[0572] 1. Initial Setup and User Authentication

[0573] 1. Start the device

[0574] When the terminal is started, it presents the user with a login screen, where the user enters their authentication information (username and password), which is then sent from the terminal to the server.

[0575] 2. Authentication Process

[0576] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[0577] 2. Understanding the user's status

[0578] 1. Posing the Question

[0579] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[0580] 2. Receiving and sending responses

[0581] The user inputs their feelings and state into the terminal, and the response is sent from the terminal to the server.

[0582] 3. Analysis of responses

[0583] The server analyzes the user's responses using a generative AI model, and based on the results of this analysis, evaluates the user's current mental state.

[0584] 3. Generating and providing conversation content

[0585] 1. Conversation content generation

[0586] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's mood and state.

[0587] 2. Sending conversation content

[0588] The generated dialogue content is transmitted from the server to the terminal.

[0589] 3. Provision by voice synthesis

[0590] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in a voice that is conveyed to the user in a natural conversational style.

[0591] 4. Data recording and learning

[0592] 1. Data recording

[0593] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which records them in a database.

[0594] 2. Training the generative AI model

[0595] The server uses the recorded data to update the generative AI model to improve the accuracy of the next interaction. This learning allows the system to better respond to the user's individual needs.

[0596] Specific examples

[0597] If the user responds "I'm a little tired"

[0598] 1. Terminal: "How are you feeling today?"

[0599] 2. User: "I'm a little tired."

[0600] 3. The device sends this information to the server.

[0601] 4. The server uses the generative AI model to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[0602] 5. The device will read this aloud using speech synthesis technology.

[0603] 6. The user responds, "I want to go to the park."

[0604] 7. The device sends this response to the server.

[0605] 8. The server generates new dialogue such as "That's a good idea. Let's plan what we want to do at the park." and sends it to the device.

[0606] 9. The device will read this aloud using speech synthesis technology and the conversation will continue.

[0607] In this way, the system of the present invention responds to the user's feelings and state in real time and provides appropriate dialogue, thereby providing psychological support.

[0608] The processing flow will be explained below.

[0609] Step 1:

[0610] When the terminal is started, it displays a login screen to the user.

[0611] Step 2:

[0612] The user enters a username and password into the terminal.

[0613] Step 3:

[0614] The terminal sends the user's authentication information to the server.

[0615] Step 4:

[0616] The server checks the received authentication information against the database to verify that the user is a legitimate user.

[0617] Step 5:

[0618] The server sends the authentication result to the terminal. If the authentication is successful, an authentication success message is sent to the terminal and the user is successfully logged in. If the authentication is unsuccessful, an error message is sent to the terminal.

[0619] Step 6:

[0620] Once the login is successful, the terminal will ask the user questions to confirm their current state of mind and condition.

[0621] Step 7:

[0622] The user inputs their feelings and state into the terminal and responds.

[0623] Step 8:

[0624] The terminal sends the user's answer to the server.

[0625] Step 9:

[0626] The server uses a generative AI model to analyze the user's responses and evaluate the user's mental state.

[0627] Step 10:

[0628] The server generates dialogue content based on the analysis results using an AI model.

[0629] Step 11:

[0630] The server transmits the generated dialogue content to the terminal.

[0631] Step 12:

[0632] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in voice.

[0633] Step 13:

[0634] The user responds to the terminal.

[0635] Step 14:

[0636] The terminal sends the user's response to the server.

[0637] Step 15:

[0638] Based on the user's response, the server generates new dialogue content using a generative AI model and sends it to the device.

[0639] Step 16:

[0640] The terminal provides new dialogue content to the user using voice synthesis technology.

[0641] Step 17:

[0642] The terminal sequentially transmits the contents of the dialogue with the user and the user's responses to the server.

[0643] Step 18:

[0644] The server records the dialogue and the user's responses in a database.

[0645] Step 19:

[0646] The server uses the recorded data to update the generative AI model and learn to improve the accuracy of the next interaction.

[0647] Example 1

[0648] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0649] In modern society, many people suffer from mental stress and fatigue, and appropriate mental care is needed to alleviate these burdens. However, conventional mental care systems have had difficulty accurately grasping the user's mental state and providing dialogue content tailored to individual needs. In addition, the methods for providing dialogue content are limited, making it difficult to achieve natural and effective support.

[0650] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0651] In this invention, the server includes means for receiving and authenticating a user's authentication information, means for presenting questions to ascertain the user's mental state and receiving the answers, means for analyzing the user's answers using a generative AI model and generating appropriate dialogue content, means for providing the generated dialogue content to the user using speech synthesis technology, and means for recording the dialogue content and the user's responses and learning the dialogue content with the generative AI model, thereby enabling the provision of appropriate psychological support to the user in real time.

[0652] "User" refers to an individual who uses the system to receive psychological support.

[0653] "Authentication information" is information used to identify a user and authorize access to a system, and typically includes a username and password.

[0654] "Means" refers to a method, technique, or device for accomplishing a particular function.

[0655] "Mental state" refers to the user's emotional and psychological state, including stress and fatigue.

[0656] "Question" refers to a query posed by the system to ascertain the user's mental state.

[0657] A "generative AI model" refers to a model that uses machine learning and artificial intelligence technology to analyze data and automatically generate dialogue content.

[0658] "Analysis" refers to the process of evaluating the user's mental state based on their answers and deriving appropriate dialogue content.

[0659] "Dialogue content" refers to the advice and response text provided to the user.

[0660] "Speech synthesis technology" refers to technology for converting text data into speech.

[0661] "Recording" refers to the act of the system saving the content of interactions and responses with the user.

[0662] "Learning" refers to the process by which the system improves the generated AI model based on past data, improving the accuracy of the next interaction.

[0663] The present invention provides a mental care system that provides mental support to users. This system has a series of functions, such as user authentication, understanding of emotions and conditions, generating dialogue content, providing dialogue through voice synthesis, and recording and learning dialogue data.

[0664] Overall system overview

[0665] The system is primarily composed of three elements: a server, a terminal, and a user. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[0666] Hardware and software used

[0667] Server: Responsible for data processing and operation of the generative AI model, authenticating users, analyzing responses, and generating dialogue content.

[0668] Terminal: Presents questions to the user, receives answers from the user, and provides the generated dialogue content through voice synthesis.

[0669] Generative AI model: Uses artificial intelligence techniques to generate dialogue content based on user input data.

[0670] Speech synthesis technology: A speech synthesis engine is used to convert text data into speech.

[0671] Specific examples

[0672] If the user answers "I'm a little tired," the specific sequence of events is as follows:

[0673] 1. Terminal: Display the question to the user: "How are you feeling today?"

[0674] 2. User: Type "I'm a little tired."

[0675] 3. The device sends this information to the server.

[0676] 4. The server uses the generative AI model to analyze the answer "I'm a little tired" and assess the user's condition.

[0677] 5. Based on the analysis results, the server generates dialogue content such as, "We will provide you with some advice to help you relax."

[0678] 6. The server sends the generated dialogue content to the terminal.

[0679] 7. The device uses speech synthesis technology to read out the generated dialogue in a natural voice, saying, "We'll give you some advice to help you relax."

[0680] 8. The user may ask further questions about how to relax, in which case the device sends a response to the server, which again uses the generative AI model to generate appropriate dialogue.

[0681] When implementing this system, it is also important to consider the prompt sentences to generate prepared questions. For example, the following prompt sentences can be used in response to input from the user:

[0682] Example prompt sentence:

[0683] "I'm a little tired. How can I relax?"

[0684] The mental care system of the present invention provides mental support by grasping the user's feelings and state in real time and providing appropriate dialogue, allowing the user to easily receive advice and counseling to reduce stress and fatigue in daily life.

[0685] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0686] Program processing flow

[0687] Step 1: Initial configuration and user authentication

[0688] 1.1 Starting the terminal

[0689] When the terminal is started, it displays a login screen. The user enters a username and password. This authentication information is sent from the terminal to the server.

[0690] Input: Username, Password

[0691] Output: Sending authentication information

[0692] 1.2 Authentication process

[0693] The server checks the received authentication information against its database, checking the username and password to ensure the user is a valid user, and sending a success message if authentication is successful, or an error message if authentication is unsuccessful.

[0694] Input: Credentials

[0695] Data processing: Database collation

[0696] Output: Authentication success message or error message

[0697] 1.3 Displaying authentication results

[0698] The terminal receives the message from the server and displays the success or failure of the authentication to the user.

[0699] Input: Authentication success message or error message

[0700] Output: Display of authentication result

[0701] Step 2: Understanding the user's state

[0702] 2.1 Posing the Question

[0703] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[0704] Input: Login successful

[0705] Output: Question posed

[0706] 2.2 Receiving a response

[0707] The user inputs their feelings and state in response to the questions presented, and the answers are sent from the device to the server.

[0708] Input: User's answer

[0709] Output: Sending the answer

[0710] 2.3 Analysis of responses

[0711] The server inputs the received answers into a generative AI model for analysis. For example, if the user answers something like "I'm a little tired," it analyzes that information to evaluate the user's current mental state.

[0712] Input: User's answer

[0713] Data Computation: Analysis with Generative AI Models

[0714] Output: Mental state assessment

[0715] Step 3: Generate and provide conversation content

[0716] 3.1 Conversation content generation

[0717] Based on the evaluation results, the server uses a generative AI model to generate optimal dialogue content for the user. For example, if the user is evaluated as tired, the server generates dialogue content such as "We will give you advice on how to relax."

[0718] Input: Mental status assessment

[0719] Data Computation: Dialogue Content Generation with Generative AI Models

[0720] Output: Generated dialogue

[0721] 3.2 Transmission of dialogue content

[0722] The server transmits the generated dialogue content to the terminal.

[0723] Input: Generated dialogue

[0724] Output: Sending dialogue

[0725] 3.3 Provision by voice synthesis

[0726] The terminal converts the received dialogue content into voice using speech synthesis technology and provides it to the user.

[0727] Input: Generated dialogue

[0728] Data processing: voice synthesis

[0729] Output: Provides dialogue via voice

[0730] Step 4: Data recording and learning

[0731] 4.1 Data recording

[0732] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which then records them in a database.

[0733] Input: Dialogue content, user response

[0734] Output: Data recording

[0735] 4.2 Training generative AI models

[0736] The server uses the recorded data to update the generative AI model, which improves the accuracy of the next interaction.

[0737] Input: Recorded data

[0738] Data Computation: Learning Generative AI Models

[0739] Output: Updated generative AI model

[0740] Through the above processing steps, the mental care system of the present invention has the ability to grasp the user's feelings and state in real time and provide appropriate dialogue.

[0741] (Application example 1)

[0742] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0743] In modern society, the number of users who need mental support is increasing, but there is a lack of systems that can individually generate dialogue content and provide it in an appropriate format.In addition, while there is a demand for effective mental care using smart devices, there is currently no system that can provide appropriate dialogue and content based on the user's mood and state.This poses the problem that users cannot receive support that is appropriate for their own physical and mental state.

[0744] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0745] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, means for recording the dialogue content and the user's responses and learning them using a generative AI model, and means for providing dialogue content and content generated by AI based on the user's feelings and state. This enables the user to receive support in real time according to their own feelings and state.

[0746] "User credentials" are the identifying information provided by a user to log into a system.

[0747] "Mood" refers to the user's feelings and moods.

[0748] "Condition" refers to the physical or mental condition of a user.

[0749] "Questions" are questions that the system presents to ascertain the user's feelings and state of mind.

[0750] An "answer" is information that a user enters in response to a question.

[0751] "Analysis" refers to the processing of data to evaluate users' responses and understand their sentiments and state of mind.

[0752] "Dialogue content" refers to messages consisting of text or voice used in dialogue with a user.

[0753] "Speech synthesis technology" is a technology that converts text into a voice that sounds like a human voice.

[0754] "User response" is information that the user responds to the dialogue content presented by the system.

[0755] "Recording" means saving the dialogue content and user responses in a database.

[0756] A "generative AI model" is an artificial intelligence model that generates appropriate dialogue content based on user information.

[0757] "Content" refers to information material such as video, music, text, etc., provided to users.

[0758] "Dialogue content generated by AI based on emotions and state" refers to messages created by a generative AI model based on the user's emotions and state.

[0759] This invention is a content distribution system that provides psychological support to users. The system is mainly composed of three elements: a server, a terminal, and a user. The server is responsible for data processing and operation of the generative AI model, and the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[0760] Server configuration and functions

[0761] The server includes means for receiving and authenticating a user's authentication information, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for transmitting the generated dialogue content to the terminal, means for recording the dialogue content and the user's responses and learning using a generative AI model, and means for providing dialogue content and content generated by AI based on the user's feelings and state.

[0762] The server uses Python and Flask as the web framework. The generative AI model is implemented using the OpenAI API. It receives and analyzes user data to generate dialogue content, which is then sent to the device. Furthermore, the dialogue data with the user is recorded in a database and used as training data for the generative AI model.

[0763] Device configuration and functions

[0764] The terminal is a device operated by the user, such as a smartphone or tablet, that provides a variety of functions. The terminal displays a user authentication screen and sends authentication information from the user to the server. If authentication is successful, the terminal displays questions to confirm the user's feelings and state, and sends the user's answers to the server. The dialogue received from the server is read aloud using speech synthesis technology.

[0765] The voice synthesis technology uses the Google Text-to-Speech API, which allows the generated dialogue to be presented to the user in a natural conversational format. The device also displays and plays relaxing content such as videos and music.

[0766] User actions and responses

[0767] Users operate their terminals to log in to the system. After logging in, they can report their own feelings and state by answering questions posed by the system. The system generates and provides appropriate dialogue and content based on this information. Users can receive psychological support by using the dialogue and content provided.

[0768] Specific examples

[0769] When a user logs in to the system and responds, "I'm a little tired," the server sends a dialogue to the device, such as, "We've checked your condition. We'll give you some advice to help you relax." This dialogue is read aloud to the user using voice synthesis technology. If the user then responds, "I'd like to go to the park," the server generates a new dialogue, such as, "That's a good idea. Let's plan what we want to do at the park," and sends it to the device.

[0770] Prompt Sentence Examples

[0771] "User is a little fatigued. Please generate a dialogue that offers a support conversation."

[0772] As described above, the system of the present invention can provide psychological support in real time according to the user's feelings and condition.

[0773] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0774] Step 1:

[0775] When the terminal is started, a login screen is displayed. The user enters authentication information (username and password) which the terminal sends to the server. The input is the username and password, and the output is the authentication information sent to the server.

[0776] Step 2:

[0777] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal. The input is authentication information, and the output is an authentication success message or an error message.

[0778] Step 3:

[0779] If authentication is successful, the device asks the user a question to confirm their current state of mind. The question is in the form of "How are you feeling today?" The input is a successful authentication message from the server, and the output is the question displayed to the user.

[0780] Step 4:

[0781] The user inputs their feelings or state of mind into the terminal. For example, they reply, "I'm a little tired." The input is text describing the user's feelings or state, and the output is the transmission of the reply data to the server.

[0782] Step 5:

[0783] The server inputs the received user responses into a generative AI model for analysis. This analysis evaluates the user's feelings and state and generates appropriate dialogue content. The input is the user's response, and the output is the analysis results and the generated dialogue content. Specifically, it sends a prompt to the OpenAI API and receives a response.

[0784] Step 6:

[0785] The server sends the generated dialogue content to the terminal. A response message containing the generated dialogue content is output. The input is the analysis result of the AI ​​model, and the output is the dialogue content sent to the terminal.

[0786] Step 7:

[0787] The device uses speech synthesis technology to provide the received dialogue to the user. The speech synthesis technology used is the Google Text-to-Speech API. The input is the dialogue from the server, and the output is the result presented to the user as voice.

[0788] Step 8:

[0789] The user responds to the dialogue content heard by voice and inputs the response into the terminal. The input is the user's response text, and the output is the transmission of the response data to the server.

[0790] Step 9:

[0791] The server records the received user responses in a database and uses them as training data for the generative AI model. The input is the user response data, and the output is an updated AI model. Specific operations include recording to the database and retraining the generative AI model.

[0792] As described above, this system performs a series of processes, starting with user authentication, then understanding the user's feelings and state, generating dialogue content using a generative AI model, providing dialogue through voice synthesis, and finally recording data and learning the model. Through this process, it is possible to provide individualized support according to the user's state.

[0793] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0794] This invention is a mental care system that provides mental support to users, and by combining it with an emotion engine, it is possible to recognize the user's emotional state in real time and to engage in dialogue and respond based on that. The system has a series of functions, including user authentication, understanding of feelings and state, emotion recognition, dialogue content generation, dialogue provision through voice synthesis, and dialogue data recording and learning.

[0795] Overall system overview

[0796] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion engine has the function of recognizing the user's emotional state in real time. The user is the entity that interacts with the system through the terminal.

[0797] 1. Initial Setup and User Authentication

[0798] 1. Start the device

[0799] When the terminal is started, it presents the user with a login screen, where the user enters their authentication information (username and password), which is then sent from the terminal to the server.

[0800] 2. Authentication Process

[0801] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[0802] 2. Understanding user state and emotion recognition

[0803] 1. Posing the Question

[0804] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[0805] 2. Receiving and sending responses

[0806] The user inputs their feelings and state into the terminal, and the response is sent from the terminal to the server.

[0807] 3. Emotion recognition

[0808] The emotion engine recognizes emotions from the user's voice and text. The emotion engine analyzes the user's tone of voice and expressions to determine their emotional state.

[0809] 4. Analysis of responses and emotional information

[0810] The server uses a generative AI model to analyze the user's responses and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[0811] 3. Generating and providing conversation content

[0812] 1. Conversation content generation

[0813] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's state of mind and condition, taking into account the emotional information recognized by the emotion engine.

[0814] 2. Sending conversation content

[0815] The generated dialogue content is transmitted from the server to the terminal.

[0816] 3. Provision by voice synthesis

[0817] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in a voice that is conveyed to the user in a natural conversational style.

[0818] 4. Data recording and learning

[0819] 1. Data recording

[0820] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which records them in a database.

[0821] 2. Training the generative AI model

[0822] The server uses the recorded data to update the generative AI model to improve the accuracy of the next interaction. This learning allows the system to better respond to the user's individual needs.

[0823] Specific examples

[0824] If the user responds "I'm a little tired"

[0825] 1. Terminal: "How are you feeling today?"

[0826] 2. User: "I'm a little tired."

[0827] 3. The device sends this information to the server.

[0828] 4. The emotion engine recognizes "fatigue" from the user's tone of voice.

[0829] 5. The server uses the information obtained from the generated AI model and emotion engine to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[0830] 6. The device will read this aloud using speech synthesis technology.

[0831] 7. The user responds, "I want to go to the park."

[0832] 8. The device sends this response to the server.

[0833] 9. The server generates new dialogue such as "That's a good idea. Let's plan what we want to do at the park." and sends it to the device.

[0834] 10. The device will read this aloud using speech synthesis technology and the conversation will continue.

[0835] In this way, the system of the present invention, which combines an emotion engine, recognizes the user's feelings and state in real time and provides appropriate dialogue to support the user emotionally.Furthermore, by generating dialogue content that reflects the user's emotional state, more personalized responses are possible.

[0836] The processing flow will be explained below.

[0837] Step 1:

[0838] When the terminal is started, it displays a login screen to the user.

[0839] Step 2:

[0840] The user enters a username and password into the terminal.

[0841] Step 3:

[0842] The terminal sends the user's authentication information to the server.

[0843] Step 4:

[0844] The server checks the received authentication information against the database to verify that the user is a legitimate user.

[0845] Step 5:

[0846] The server sends the authentication result to the terminal. If the authentication is successful, an authentication success message is sent to the terminal and the user is successfully logged in. If the authentication is unsuccessful, an error message is sent to the terminal.

[0847] Step 6:

[0848] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[0849] Step 7:

[0850] The user inputs their feelings and state into the terminal and responds.

[0851] Step 8:

[0852] The terminal sends the user's answer to the server.

[0853] Step 9:

[0854] The emotion engine recognizes emotions from the user's voice and text, analyzing the user's tone of voice and expressions to determine their emotional state.

[0855] Step 10:

[0856] The server uses a generative AI model to analyze the user's responses and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[0857] Step 11:

[0858] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's state of mind and condition, taking into account the emotional information recognized by the emotion engine.

[0859] Step 12:

[0860] The server transmits the generated dialogue content to the terminal.

[0861] Step 13:

[0862] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in voice.

[0863] Step 14:

[0864] The user responds to the terminal.

[0865] Step 15:

[0866] The terminal sends the user's response to the server.

[0867] Step 16:

[0868] Based on the user's response, the server generates new dialogue content using a generative AI model and sends it to the device.

[0869] Step 17:

[0870] The terminal provides new dialogue content to the user using voice synthesis technology.

[0871] Step 18:

[0872] The terminal sequentially transmits the contents of the dialogue with the user and the user's responses to the server.

[0873] Step 19:

[0874] The server records the dialogue and the user's responses in a database.

[0875] Step 20:

[0876] The server uses the recorded data to update the generative AI model and learn to improve the accuracy of the next interaction.

[0877] As a specific example, the flow when the user answers "I'm a little tired" is shown below.

[0878] Step 6:

[0879] Terminal: "How are you feeling today?"

[0880] Step 7:

[0881] User: "I'm a little tired."

[0882] Step 8:

[0883] The terminal sends the user's answer to the server.

[0884] Step 9:

[0885] The emotion engine recognizes "fatigue" from the user's tone of voice.

[0886] Step 10:

[0887] The server uses information obtained from the generative AI model and emotion engine to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[0888] Step 12:

[0889] The server transmits the generated dialogue content to the terminal.

[0890] Step 13:

[0891] The device will read this aloud using voice synthesis technology.

[0892] Step 14:

[0893] The user responds, "I want to go to the park."

[0894] Step 15:

[0895] The terminal sends the user's response to the server.

[0896] Step 16:

[0897] The server generates new dialogue content such as "That's a good idea. Let's plan what we want to do in the park," and sends it to the device.

[0898] Step 17:

[0899] The terminal provides new dialogue content to the user using voice synthesis technology.

[0900] In this way, the system of the present invention, which is combined with an emotion engine, recognizes the user's emotional state in real time and provides dialogue accordingly, thereby realizing more personalized psychological support.

[0901] Example 2

[0902] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0903] In modern society, the number of users suffering from mental stress and anxiety is increasing, but there are only a limited number of systems that provide immediate and personalized responses. Conventional systems have difficulty accurately recognizing the user's emotional state and generating appropriate dialogue based on that, making it impossible to provide adequate mental care. Furthermore, technology for recognizing the user's emotions through voice or text and generating dialogue that reflects this is immature, resulting in a lack of satisfactory support for users.

[0904] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0905] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to confirm the user's feelings and state and receiving the answers, means for analyzing the user's answers and recognizing and evaluating the user's emotions using an emotion engine, means for using a generative AI model to generate optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, and means for recording the dialogue content and the user's responses and learning with the generative AI model. This makes it possible to accurately recognize the user's emotional state in real time and generate natural dialogue based on that.

[0906] "User authentication information" refers to the identifying information a user needs to log in to a system, typically a username and password.

[0907] "Emotion engine" refers to an algorithm or software that recognizes and analyzes emotions from a user's voice or text data in real time.

[0908] A "generative AI model" is an artificial intelligence model that learns large amounts of data and generates dialogue content with users, and uses natural language processing technology.

[0909] "Speech synthesis technology" refers to technology for generating natural speech based on text data, enabling information to be conveyed to users through speech.

[0910] "Dialogue content" refers to the response from the system generated in response to the user's question or status, and includes information such as appropriate support and advice.

[0911] "Emotional state" refers to a psychological state that indicates the type and intensity of the emotion the user is currently experiencing, and includes, for example, joy, sadness, anger, fatigue, and the like.

[0912] "Database" refers to a system for storing and managing data necessary for system operation, such as user authentication information, interaction history, and learning data for generative AI models.

[0913] "Real-time" refers to the instantaneous acquisition and analysis of data, generating an appropriate response immediately.

[0914] "Login screen" refers to the screen interface that is displayed to a user to enter their authentication information.

[0915] "Personalized response" refers to providing responses and support that are customized to suit the individual conditions and needs of the user.

[0916] This invention is a mental care system that provides mental support to users, and by combining it with an emotion engine, it is possible to recognize the user's emotional state in real time and to engage in dialogue and respond based on that. This system has a series of functions, including user authentication, understanding of feelings and state, emotion recognition, dialogue content generation, dialogue provision through voice synthesis, and dialogue data recording and learning.

[0917] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion engine has the function of recognizing the user's emotional state in real time. The user is the entity that interacts with the system through the terminal.

[0918] Hardware and software used

[0919] Hardware: Terminals (PCs, smartphones), servers

[0920] Software: Emotion engine, generative AI model, speech synthesis engine, database management system

[0921] Program processing

[0922] Initial Setup and User Authentication

[0923] When the terminal is started, it displays a login screen to the user. The user enters their authentication information (username and password), and the terminal sends this information to the server. The server compares the received authentication information with a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[0924] Understanding user status and recognizing emotions

[0925] After successful login, the device asks the user questions to ascertain their current mood and state (e.g., "How are you feeling today?"). The user inputs their mood and state into the device, and the answer is sent from the device to the server. The server uses an emotion engine to recognize emotions from the user's voice and text. The emotion engine analyzes the user's tone of voice and expressions to determine their emotional state. The server uses a generative AI model to analyze the user's answers and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[0926] Conversation content generation and provision

[0927] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's feelings and state. The emotional information recognized by the emotion engine is also taken into consideration during this process. The generated dialogue content is sent from the server to the device, which then uses speech synthesis technology to provide the generated dialogue content to the user in a voice format. This voice is conveyed to the user in a natural conversational format.

[0928] Data recording and learning

[0929] The device continuously transmits the user's interactions and responses to the server, which records them in a database and uses the recorded data to update the generative AI model. This learning allows the system to better respond to the user's individual needs.

[0930] Additional examples of specific actions

[0931] For example, if the user responds, "I'm a little tired," the process goes something like this: The device asks, "How are you feeling today?" The user responds, "I'm a little tired," and the device sends that information to the server. The emotion engine recognizes "fatigue" from the user's tone of voice, and the server uses the information obtained from the generative AI model and the emotion engine to generate a dialogue message saying, "We've checked your condition and will provide you with some advice to help you relax." The device reads this aloud using speech synthesis technology. The user responds, "I'd like to go to the park," and the device sends that response to the server. The server generates a new dialogue message saying, "That's a good idea. Let's plan what we want to do at the park," sends it to the device, and the device continues reading this aloud using speech synthesis technology.

[0932] Prompt Sentence Examples

[0933] "Generate a dialogue for when the user responds that they are a little tired."

[0934] "Provide relaxing advice based on the user's emotional state."

[0935] In this way, the present invention can recognize the user's feelings and state in real time and provide appropriate dialogue in response to them, thereby providing psychological support.

[0936] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0937] Mental health care system processing steps

[0938] Initial Setup and User Authentication

[0939] Step 1:

[0940] Starting the terminal

[0941] When the terminal is started, it displays a login screen to the user. The input is powering on, and the output is displaying the login screen.

[0942] Step 2:

[0943] Enter your authentication information

[0944] A user enters authentication information (username and password) into a login screen. The input is the authentication information, and the output is its preparation for submission.

[0945] Step 3:

[0946] Sending authentication information

[0947] The terminal sends authentication information to the server. The input is the authentication information and the output is a request sent to the server.

[0948] Step 4:

[0949] Authentication verification

[0950] The server checks the received authentication information against a database to verify the user's validity. The input is the authentication information, and the output is the authentication result (success or failure).

[0951] Step 5:

[0952] Return and display of authentication results

[0953] The server returns the authentication result to the terminal, which displays it to the user. The input is the authentication result, and the output is an indication of authentication success or failure. If successful, login is complete.

[0954] Understanding user status and recognizing emotions

[0955] Step 6:

[0956] Posing the Question

[0957] The terminal displays a question to the user to confirm their mood or state. For example, "How are you feeling today?" The input is a successful login status, and the output is the display of the question.

[0958] Step 7:

[0959] Enter your answer

[0960] The user inputs their feelings and state. The input is the user's answer to a question, and the output is the answer ready to be sent.

[0961] Step 8:

[0962] Submit your answer

[0963] The terminal sends the user's answer to the server. The input is the user's answer, and the output is the response sent to the server.

[0964] Step 9:

[0965] Performing emotion recognition

[0966] The server inputs the user's response into the emotion engine to recognize the user's emotion. The input is the user's response, and the output is the recognized emotion information.

[0967] Step 10:

[0968] Emotional information analysis

[0969] The server performs analysis using the results of the emotion engine and the generative AI model. The input is emotion information and the generative AI model, and the output is an evaluation of the user's mental state.

[0970] Conversation content generation and provision

[0971] Step 11:

[0972] Conversation generation

[0973] The server uses a generative AI model to generate optimal dialogue for the user. The input is the mental state assessment, and the output is the generated dialogue.

[0974] Step 12:

[0975] Sending conversation transcripts

[0976] The generated dialogue content is sent from the server to the terminal. The input is the generated dialogue content, and the output is the content sent to the terminal.

[0977] Step 13:

[0978] Provided by voice synthesis

[0979] The terminal uses speech synthesis technology to communicate the dialogue content to the user by voice. The input is the generated dialogue content, and the output is the voice output to the user.

[0980] Data recording and learning

[0981] Step 14:

[0982] Data recording

[0983] The terminal sends the dialogue content and the user's response to the server, which records it in a database. The input is the dialogue content and the user's response, and the output is the record in the database.

[0984] Step 15:

[0985] Training generative AI models

[0986] The server uses the recorded data to train the generative AI model to improve the accuracy of the next interaction. The input is the recorded data, and the output is an updated generative AI model.

[0987] (Application example 2)

[0988] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0989] In modern work environments, especially in factories, long hours of monotonous work can cause mental stress for workers. This can lead to reduced work efficiency and increased risk of health problems. Conventional mental health systems face the challenge of being unable to recognize and respond to workers' emotional states in real time. The present invention aims to solve these challenges and provide an advanced mental health system for supporting the mental health of workers.

[0990] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0991] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, means for recording the dialogue content and the user's responses and learning it with a generative AI model, means for converting the user's voice into text using speech recognition technology, emotion recognition means for detecting the user's emotional state in real time, and means for dynamically changing the dialogue content based on the responses. This allows the system to recognize the emotional state of workers in real time and provide appropriate dialogue, thereby providing psychological support and improving work efficiency and maintaining health.

[0992] "User authentication" is the process of receiving user authentication information and verifying that the user is legitimate.

[0993] "Mood confirmation" is a process of asking questions to confirm the user's mood or mental state.

[0994] "Answer analysis" is the process of analyzing a user's answer and evaluating its content.

[0995] "Dialogue content generation" is a process of generating optimal dialogue content for the user based on the analysis results.

[0996] "Speech synthesis technology" is a technology that outputs text information as voice.

[0997] "Data logging" is the process of saving the dialogue and user responses.

[0998] A "generative AI model" is an artificial intelligence model that learns from data and generates appropriate dialogue content.

[0999] "Speech recognition technology" is a technology that converts voice data into text data.

[1000] "Emotion recognition means" is a technology that detects the user's emotional state in real time from their voice and facial expressions.

[1001] The "dynamic dialogue change means" is a process that changes the dialogue content based on the user's response.

[1002] The present invention relates to a mental care system that combines user authentication, emotional state confirmation, emotion recognition, dialogue content generation, voice synthesis, data recording, and learning. Specific embodiments of each element will be described below.

[1003] Overall system configuration

[1004] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion recognition engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion recognition engine has the function of recognizing the user's emotional state in real time, and the user is the entity that interacts with the system through the terminal.

[1005] server

[1006] The server centrally processes data and runs the generative AI model. The server receives the user's authentication information and checks it against the database to verify that the user is legitimate. If authentication is successful, the server returns a successful authentication message to the user.

[1007] Terminal

[1008] The terminal accepts input from the user and communicates with the server. When the user inputs their feelings or state into the terminal, this information is sent to the server. If login is successful, the terminal displays questions to the user to confirm their current feelings or state.

[1009] Emotion Recognition Engine

[1010] The emotion recognition engine works in conjunction with voice recognition technology to analyze the user's tone and expressions to recognize their emotional state, which is then sent to the server and used to assess the user's mental state.

[1011] User

[1012] The user interacts with the system through the device. When the user inputs their feelings and state, the information is sent to the server and analyzed by an emotion recognition engine. The server then uses a generative AI model to generate optimal dialogue content and send it to the device.

[1013] Detailed process flow

[1014] 1. User authentication:

[1015] The user enters a username and password into the terminal and sends them to the server.

[1016] The server checks the authentication information against a database to verify the user is a valid user.

[1017] If the authentication is successful, the server sends an authentication success message to the terminal, completing the login.

[1018] 2. Emotional validation and emotional recognition:

[1019] If authentication is successful, the terminal presents the user with questions to confirm their feelings and state of mind, and receives their answers.

[1020] The emotion recognition engine analyzes the user's voice and recognizes their emotional state in real time.

[1021] The server uses a generative AI model to analyze the user's responses and emotional information to assess their mental state.

[1022] 3. Generating and providing dialogue content:

[1023] The server uses a generative AI model based on the analysis results to generate optimal dialogue content.

[1024] The generated dialogue content is sent from the server to the terminal and provided to the user using voice synthesis technology.

[1025] 4. Data recording and learning:

[1026] The dialogue and the user's responses are sent to the server and recorded in a database.

[1027] The recorded data is used to train the generative AI model to improve the accuracy of the next interaction.

[1028] Hardware and software used

[1029] Server: Database, generative AI models (e.g., Hugging Face Transformers)

[1030] Devices: Smartphones, PCs, tablets

[1031] Emotion recognition engine: Speech recognition technology (e.g., Google Speech Recognition), emotion analysis (e.g., Sentiment Analysis Pipeline)

[1032] Speech synthesis: Pyttsx3 (e.g., a Python text-to-speech library)

[1033] Examples and prompts

[1034] Example 1:

[1035] Please enter your username: user123

[1036] Please enter your password: password123

[1037] Authentication successful. How are you feeling today?

[1038] (User speaks): "I'm a little tired."

[1039] The system responds: "You sound a little down. Let me know if there's anything I can do to help."

[1040] Example 2:

[1041] Please enter your username: sampleuser

[1042] Please enter your password: mypassword

[1043] Authentication successful. How are you feeling today?

[1044] (User speaks): "I feel great today."

[1045] The system responds: "Great! Keep it up!"

[1046] This allows the system to recognize the emotional state of workers in real time and provide appropriate dialogue to provide psychological support, thereby improving work efficiency and maintaining health.

[1047] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1048] Step 1:

[1049] User Authentication

[1050] Input: The user enters a username and password into the terminal.

[1051] Specific operation: The device sends these authentication information to the server.

[1052] Data processing / calculation: The server checks the authentication information against a database to verify that the user is legitimate.

[1053] Output: If the authentication is successful, the server returns an authentication success message to the terminal; if the authentication fails, it returns an error message.

[1054] Step 2:

[1055] Confirmation of feelings

[1056] Input: Receives information that authentication was successful.

[1057] Specific behavior: The device presents the user with a question to ascertain their current state of mind (e.g., "How are you feeling today?").

[1058] Data processing / calculation: Receive user responses via voice or text input.

[1059] Output: The device sends the user's answer to the server.

[1060] Step 3:

[1061] emotion recognition

[1062] Input: The server receives the user's answer and voice data.

[1063] How it works: The emotion recognition engine analyzes the user's tone of voice and expressions to recognize their emotional state in real time.

[1064] Data processing / calculation: The analysis results are sent to the server, and the emotion data is evaluated using a generative AI model.

[1065] Output: The recognized emotional state and analysis results are output.

[1066] Step 4:

[1067] Dialogue content generation

[1068] Input: Analyzed user emotional state and sentiment data.

[1069] Specific operation: The server uses a generative AI model to generate optimal dialogue content based on the user's emotions and feelings.

[1070] Data processing / calculation: The AI ​​model generates dialogue content based on emotion recognition results and sentiment data.

[1071] Output: The generated dialogue content is sent from the server to the terminal.

[1072] Step 5:

[1073] Dialogue provision

[1074] Input: The generated dialogue.

[1075] Specific operation: The device uses speech synthesis technology to provide the generated dialogue content to the user via voice.

[1076] Data processing / calculation: Converting text data into audio data.

[1077] Output: The dialogue is presented to the user via audio.

[1078] Step 6:

[1079] Data Recording

[1080] Input: Dialogue and user response data.

[1081] Specific operation: The dialogue content sent from the terminal to the server and the user's response data are recorded in a database.

[1082] Data processing / calculation: Recorded data is used as training data for generative AI models.

[1083] Output: Interaction data stored in a database and an updated AI model.

[1084] This allows the system to recognize the user's emotional state in real time and provide appropriate dialogue, thereby providing psychological support, improving work efficiency, and maintaining health.

[1085] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1086] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1087] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1088] [Third embodiment]

[1089] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1090] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1091] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1092] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1093] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1094] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1095] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1096] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1097] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1098] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1099] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1100] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1101] The present invention is a mental care system that provides mental support to users. This system has a series of functions, such as user authentication, grasping of feelings and conditions, generation of dialogue content, dialogue provision by voice synthesis, and recording and learning of dialogue data.

[1102] Overall system overview

[1103] The system is primarily composed of three elements: a server, a terminal, and a user. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[1104] 1. Initial Setup and User Authentication

[1105] 1. Start the device

[1106] When the terminal is started, it presents the user with a login screen, where the user enters their authentication information (username and password), which is then sent from the terminal to the server.

[1107] 2. Authentication Process

[1108] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[1109] 2. Understanding the user's status

[1110] 1. Posing the Question

[1111] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[1112] 2. Receiving and sending responses

[1113] The user inputs their feelings and state into the terminal, and the response is sent from the terminal to the server.

[1114] 3. Analysis of responses

[1115] The server analyzes the user's responses using a generative AI model, and based on the results of this analysis, evaluates the user's current mental state.

[1116] 3. Generating and providing conversation content

[1117] 1. Conversation content generation

[1118] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's mood and state.

[1119] 2. Sending conversation content

[1120] The generated dialogue content is transmitted from the server to the terminal.

[1121] 3. Provision by voice synthesis

[1122] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in a voice that is conveyed to the user in a natural conversational style.

[1123] 4. Data recording and learning

[1124] 1. Data recording

[1125] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which records them in a database.

[1126] 2. Training the generative AI model

[1127] The server uses the recorded data to update the generative AI model to improve the accuracy of the next interaction. This learning allows the system to better respond to the user's individual needs.

[1128] Specific examples

[1129] If the user responds "I'm a little tired"

[1130] 1. Terminal: "How are you feeling today?"

[1131] 2. User: "I'm a little tired."

[1132] 3. The device sends this information to the server.

[1133] 4. The server uses the generative AI model to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[1134] 5. The device will read this aloud using speech synthesis technology.

[1135] 6. The user responds, "I want to go to the park."

[1136] 7. The device sends this response to the server.

[1137] 8. The server generates new dialogue such as "That's a good idea. Let's plan what we want to do at the park." and sends it to the device.

[1138] 9. The device will read this aloud using speech synthesis technology and the conversation will continue.

[1139] In this way, the system of the present invention responds to the user's feelings and state in real time and provides appropriate dialogue, thereby providing psychological support.

[1140] The processing flow will be explained below.

[1141] Step 1:

[1142] When the terminal is started, it displays a login screen to the user.

[1143] Step 2:

[1144] The user enters a username and password into the terminal.

[1145] Step 3:

[1146] The terminal sends the user's authentication information to the server.

[1147] Step 4:

[1148] The server checks the received authentication information against the database to verify that the user is a legitimate user.

[1149] Step 5:

[1150] The server sends the authentication result to the terminal. If the authentication is successful, an authentication success message is sent to the terminal and the user is successfully logged in. If the authentication is unsuccessful, an error message is sent to the terminal.

[1151] Step 6:

[1152] Once the login is successful, the terminal will ask the user questions to confirm their current state of mind and condition.

[1153] Step 7:

[1154] The user inputs their feelings and state into the terminal and responds.

[1155] Step 8:

[1156] The terminal sends the user's answer to the server.

[1157] Step 9:

[1158] The server uses a generative AI model to analyze the user's responses and evaluate the user's mental state.

[1159] Step 10:

[1160] The server generates dialogue content based on the analysis results using an AI model.

[1161] Step 11:

[1162] The server transmits the generated dialogue content to the terminal.

[1163] Step 12:

[1164] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in voice.

[1165] Step 13:

[1166] The user responds to the terminal.

[1167] Step 14:

[1168] The terminal sends the user's response to the server.

[1169] Step 15:

[1170] Based on the user's response, the server generates new dialogue content using a generative AI model and sends it to the device.

[1171] Step 16:

[1172] The terminal provides new dialogue content to the user using voice synthesis technology.

[1173] Step 17:

[1174] The terminal sequentially transmits the contents of the dialogue with the user and the user's responses to the server.

[1175] Step 18:

[1176] The server records the dialogue and the user's responses in a database.

[1177] Step 19:

[1178] The server uses the recorded data to update the generative AI model and learn to improve the accuracy of the next interaction.

[1179] Example 1

[1180] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1181] In modern society, many people suffer from mental stress and fatigue, and appropriate mental care is needed to alleviate these burdens. However, conventional mental care systems have had difficulty accurately grasping the user's mental state and providing dialogue content tailored to individual needs. In addition, the methods for providing dialogue content are limited, making it difficult to achieve natural and effective support.

[1182] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1183] In this invention, the server includes means for receiving and authenticating a user's authentication information, means for presenting questions to ascertain the user's mental state and receiving the answers, means for analyzing the user's answers using a generative AI model and generating appropriate dialogue content, means for providing the generated dialogue content to the user using speech synthesis technology, and means for recording the dialogue content and the user's responses and learning the dialogue content with the generative AI model, thereby enabling the provision of appropriate psychological support to the user in real time.

[1184] "User" refers to an individual who uses the system to receive psychological support.

[1185] "Authentication information" is information used to identify a user and authorize access to a system, and typically includes a username and password.

[1186] "Means" refers to a method, technique, or device for accomplishing a particular function.

[1187] "Mental state" refers to the user's emotional and psychological state, including stress and fatigue.

[1188] "Question" refers to a query posed by the system to ascertain the user's mental state.

[1189] A "generative AI model" refers to a model that uses machine learning and artificial intelligence technology to analyze data and automatically generate dialogue content.

[1190] "Analysis" refers to the process of evaluating the user's mental state based on their answers and deriving appropriate dialogue content.

[1191] "Dialogue content" refers to the advice and response text provided to the user.

[1192] "Speech synthesis technology" refers to technology for converting text data into speech.

[1193] "Recording" refers to the act of the system saving the content of interactions and responses with the user.

[1194] "Learning" refers to the process by which the system improves the generated AI model based on past data, improving the accuracy of the next interaction.

[1195] The present invention provides a mental care system that provides mental support to users. This system has a series of functions, such as user authentication, understanding of emotions and conditions, generating dialogue content, providing dialogue through voice synthesis, and recording and learning dialogue data.

[1196] Overall system overview

[1197] The system is primarily composed of three elements: a server, a terminal, and a user. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[1198] Hardware and software used

[1199] Server: Responsible for data processing and operation of the generative AI model, authenticating users, analyzing responses, and generating dialogue content.

[1200] Terminal: Presents questions to the user, receives answers from the user, and provides the generated dialogue content through voice synthesis.

[1201] Generative AI model: Uses artificial intelligence techniques to generate dialogue content based on user input data.

[1202] Speech synthesis technology: A speech synthesis engine is used to convert text data into speech.

[1203] Specific examples

[1204] If the user answers "I'm a little tired," the specific sequence of events is as follows:

[1205] 1. Terminal: Display the question to the user: "How are you feeling today?"

[1206] 2. User: Type "I'm a little tired."

[1207] 3. The device sends this information to the server.

[1208] 4. The server uses the generative AI model to analyze the answer "I'm a little tired" and assess the user's condition.

[1209] 5. Based on the analysis results, the server generates dialogue content such as, "We will provide you with some advice to help you relax."

[1210] 6. The server sends the generated dialogue content to the terminal.

[1211] 7. The device uses speech synthesis technology to read out the generated dialogue in a natural voice, saying, "We'll give you some advice to help you relax."

[1212] 8. The user may ask further questions about how to relax, in which case the device sends a response to the server, which again uses the generative AI model to generate appropriate dialogue.

[1213] When implementing this system, it is also important to consider the prompt sentences to generate prepared questions. For example, the following prompt sentences can be used in response to input from the user:

[1214] Example prompt sentence:

[1215] "I'm a little tired. How can I relax?"

[1216] The mental care system of the present invention provides mental support by grasping the user's feelings and state in real time and providing appropriate dialogue, allowing the user to easily receive advice and counseling to reduce stress and fatigue in daily life.

[1217] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1218] Program processing flow

[1219] Step 1: Initial configuration and user authentication

[1220] 1.1 Starting the terminal

[1221] When the terminal is started, it displays a login screen. The user enters a username and password. This authentication information is sent from the terminal to the server.

[1222] Input: Username, Password

[1223] Output: Sending authentication information

[1224] 1.2 Authentication process

[1225] The server checks the received authentication information against its database, checking the username and password to ensure the user is a valid user, and sending a success message if authentication is successful, or an error message if authentication is unsuccessful.

[1226] Input: Credentials

[1227] Data processing: Database collation

[1228] Output: Authentication success message or error message

[1229] 1.3 Displaying authentication results

[1230] The terminal receives the message from the server and displays the success or failure of the authentication to the user.

[1231] Input: Authentication success message or error message

[1232] Output: Display of authentication result

[1233] Step 2: Understanding the user's state

[1234] 2.1 Posing the Question

[1235] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[1236] Input: Login successful

[1237] Output: Question posed

[1238] 2.2 Receiving a response

[1239] The user inputs their feelings and state in response to the questions presented, and the answers are sent from the device to the server.

[1240] Input: User's answer

[1241] Output: Sending the answer

[1242] 2.3 Analysis of responses

[1243] The server inputs the received answers into a generative AI model for analysis. For example, if the user answers something like "I'm a little tired," it analyzes that information to evaluate the user's current mental state.

[1244] Input: User's answer

[1245] Data Computation: Analysis with Generative AI Models

[1246] Output: Mental state assessment

[1247] Step 3: Generate and provide conversation content

[1248] 3.1 Conversation content generation

[1249] Based on the evaluation results, the server uses a generative AI model to generate optimal dialogue content for the user. For example, if the user is evaluated as tired, the server generates dialogue content such as "We will give you advice on how to relax."

[1250] Input: Mental status assessment

[1251] Data Computation: Dialogue Content Generation with Generative AI Models

[1252] Output: Generated dialogue

[1253] 3.2 Transmission of dialogue content

[1254] The server transmits the generated dialogue content to the terminal.

[1255] Input: Generated dialogue

[1256] Output: Sending dialogue

[1257] 3.3 Provision by voice synthesis

[1258] The terminal converts the received dialogue content into voice using speech synthesis technology and provides it to the user.

[1259] Input: Generated dialogue

[1260] Data processing: voice synthesis

[1261] Output: Provides dialogue via voice

[1262] Step 4: Data recording and learning

[1263] 4.1 Data recording

[1264] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which then records them in a database.

[1265] Input: Dialogue content, user response

[1266] Output: Data recording

[1267] 4.2 Training generative AI models

[1268] The server uses the recorded data to update the generative AI model, which improves the accuracy of the next interaction.

[1269] Input: Recorded data

[1270] Data Computation: Learning Generative AI Models

[1271] Output: Updated generative AI model

[1272] Through the above processing steps, the mental care system of the present invention has the ability to grasp the user's feelings and state in real time and provide appropriate dialogue.

[1273] (Application example 1)

[1274] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1275] In modern society, the number of users who need mental support is increasing, but there is a lack of systems that can individually generate dialogue content and provide it in an appropriate format.In addition, while there is a demand for effective mental care using smart devices, there is currently no system that can provide appropriate dialogue and content based on the user's mood and state.This poses the problem that users cannot receive support that is appropriate for their own physical and mental state.

[1276] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1277] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, means for recording the dialogue content and the user's responses and learning them using a generative AI model, and means for providing dialogue content and content generated by AI based on the user's feelings and state. This enables the user to receive support in real time according to their own feelings and state.

[1278] "User credentials" are the identifying information provided by a user to log into a system.

[1279] "Mood" refers to the user's feelings and moods.

[1280] "Condition" refers to the physical or mental condition of a user.

[1281] "Questions" are questions that the system presents to ascertain the user's feelings and state of mind.

[1282] An "answer" is information that a user enters in response to a question.

[1283] "Analysis" refers to the processing of data to evaluate users' responses and understand their sentiments and state of mind.

[1284] "Dialogue content" refers to messages consisting of text or voice used in dialogue with a user.

[1285] "Speech synthesis technology" is a technology that converts text into a voice that sounds like a human voice.

[1286] "User response" is information that the user responds to the dialogue content presented by the system.

[1287] "Recording" means saving the dialogue content and user responses in a database.

[1288] A "generative AI model" is an artificial intelligence model that generates appropriate dialogue content based on user information.

[1289] "Content" refers to information material such as video, music, text, etc., provided to users.

[1290] "Dialogue content generated by AI based on emotions and state" refers to messages created by a generative AI model based on the user's emotions and state.

[1291] This invention is a content distribution system that provides psychological support to users. The system is mainly composed of three elements: a server, a terminal, and a user. The server is responsible for data processing and operation of the generative AI model, and the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[1292] Server configuration and functions

[1293] The server includes means for receiving and authenticating a user's authentication information, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for transmitting the generated dialogue content to the terminal, means for recording the dialogue content and the user's responses and learning using a generative AI model, and means for providing dialogue content and content generated by AI based on the user's feelings and state.

[1294] The server uses Python and Flask as the web framework. The generative AI model is implemented using the OpenAI API. It receives and analyzes user data to generate dialogue content, which is then sent to the device. Furthermore, the dialogue data with the user is recorded in a database and used as training data for the generative AI model.

[1295] Device configuration and functions

[1296] The terminal is a device operated by the user, such as a smartphone or tablet, that provides a variety of functions. The terminal displays a user authentication screen and sends authentication information from the user to the server. If authentication is successful, the terminal displays questions to confirm the user's feelings and state, and sends the user's answers to the server. The dialogue received from the server is read aloud using speech synthesis technology.

[1297] The voice synthesis technology uses the Google Text-to-Speech API, which allows the generated dialogue to be presented to the user in a natural conversational format. The device also displays and plays relaxing content such as videos and music.

[1298] User actions and responses

[1299] Users operate their terminals to log in to the system. After logging in, they can report their own feelings and state by answering questions posed by the system. The system generates and provides appropriate dialogue and content based on this information. Users can receive psychological support by using the dialogue and content provided.

[1300] Specific examples

[1301] When a user logs in to the system and responds, "I'm a little tired," the server sends a dialogue to the device, such as, "We've checked your condition. We'll give you some advice to help you relax." This dialogue is read aloud to the user using voice synthesis technology. If the user then responds, "I'd like to go to the park," the server generates a new dialogue, such as, "That's a good idea. Let's plan what we want to do at the park," and sends it to the device.

[1302] Prompt Sentence Examples

[1303] "User is a little fatigued. Please generate a dialogue that offers a support conversation."

[1304] As described above, the system of the present invention can provide psychological support in real time according to the user's feelings and condition.

[1305] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1306] Step 1:

[1307] When the terminal is started, a login screen is displayed. The user enters authentication information (username and password) which the terminal sends to the server. The input is the username and password, and the output is the authentication information sent to the server.

[1308] Step 2:

[1309] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal. The input is authentication information, and the output is an authentication success message or an error message.

[1310] Step 3:

[1311] If authentication is successful, the device asks the user a question to confirm their current state of mind. The question is in the form of "How are you feeling today?" The input is a successful authentication message from the server, and the output is the question displayed to the user.

[1312] Step 4:

[1313] The user inputs their feelings or state of mind into the terminal. For example, they reply, "I'm a little tired." The input is text describing the user's feelings or state, and the output is the transmission of the reply data to the server.

[1314] Step 5:

[1315] The server inputs the received user responses into a generative AI model for analysis. This analysis evaluates the user's feelings and state and generates appropriate dialogue content. The input is the user's response, and the output is the analysis results and the generated dialogue content. Specifically, it sends a prompt to the OpenAI API and receives a response.

[1316] Step 6:

[1317] The server sends the generated dialogue content to the terminal. A response message containing the generated dialogue content is output. The input is the analysis result of the AI ​​model, and the output is the dialogue content sent to the terminal.

[1318] Step 7:

[1319] The device uses speech synthesis technology to provide the received dialogue to the user. The speech synthesis technology used is the Google Text-to-Speech API. The input is the dialogue from the server, and the output is the result presented to the user as voice.

[1320] Step 8:

[1321] The user responds to the dialogue content heard by voice and inputs the response into the terminal. The input is the user's response text, and the output is the transmission of the response data to the server.

[1322] Step 9:

[1323] The server records the received user responses in a database and uses them as training data for the generative AI model. The input is the user response data, and the output is an updated AI model. Specific operations include recording to the database and retraining the generative AI model.

[1324] As described above, this system performs a series of processes, starting with user authentication, then understanding the user's feelings and state, generating dialogue content using a generative AI model, providing dialogue through voice synthesis, and finally recording data and learning the model. Through this process, it is possible to provide individualized support according to the user's state.

[1325] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1326] This invention is a mental care system that provides mental support to users, and by combining it with an emotion engine, it is possible to recognize the user's emotional state in real time and to engage in dialogue and respond based on that. The system has a series of functions, including user authentication, understanding of feelings and state, emotion recognition, dialogue content generation, dialogue provision through voice synthesis, and dialogue data recording and learning.

[1327] Overall system overview

[1328] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion engine has the function of recognizing the user's emotional state in real time. The user is the entity that interacts with the system through the terminal.

[1329] 1. Initial Setup and User Authentication

[1330] 1. Start the device

[1331] When the terminal is started, it presents the user with a login screen, where the user enters their authentication information (username and password), which is then sent from the terminal to the server.

[1332] 2. Authentication Process

[1333] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[1334] 2. Understanding user state and emotion recognition

[1335] 1. Posing the Question

[1336] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[1337] 2. Receiving and sending responses

[1338] The user inputs their feelings and state into the terminal, and the response is sent from the terminal to the server.

[1339] 3. Emotion recognition

[1340] The emotion engine recognizes emotions from the user's voice and text. The emotion engine analyzes the user's tone of voice and expressions to determine their emotional state.

[1341] 4. Analysis of responses and emotional information

[1342] The server uses a generative AI model to analyze the user's responses and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[1343] 3. Generating and providing conversation content

[1344] 1. Conversation content generation

[1345] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's state of mind and condition, taking into account the emotional information recognized by the emotion engine.

[1346] 2. Sending conversation content

[1347] The generated dialogue content is transmitted from the server to the terminal.

[1348] 3. Provision by voice synthesis

[1349] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in a voice that is conveyed to the user in a natural conversational style.

[1350] 4. Data recording and learning

[1351] 1. Data recording

[1352] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which records them in a database.

[1353] 2. Training the generative AI model

[1354] The server uses the recorded data to update the generative AI model to improve the accuracy of the next interaction. This learning allows the system to better respond to the user's individual needs.

[1355] Specific examples

[1356] If the user responds "I'm a little tired"

[1357] 1. Terminal: "How are you feeling today?"

[1358] 2. User: "I'm a little tired."

[1359] 3. The device sends this information to the server.

[1360] 4. The emotion engine recognizes "fatigue" from the user's tone of voice.

[1361] 5. The server uses the information obtained from the generated AI model and emotion engine to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[1362] 6. The device will read this aloud using speech synthesis technology.

[1363] 7. The user responds, "I want to go to the park."

[1364] 8. The device sends this response to the server.

[1365] 9. The server generates new dialogue such as "That's a good idea. Let's plan what we want to do at the park." and sends it to the device.

[1366] 10. The device will read this aloud using speech synthesis technology and the conversation will continue.

[1367] In this way, the system of the present invention, which combines an emotion engine, recognizes the user's feelings and state in real time and provides appropriate dialogue to support the user emotionally.Furthermore, by generating dialogue content that reflects the user's emotional state, more personalized responses are possible.

[1368] The processing flow will be explained below.

[1369] Step 1:

[1370] When the terminal is started, it displays a login screen to the user.

[1371] Step 2:

[1372] The user enters a username and password into the terminal.

[1373] Step 3:

[1374] The terminal sends the user's authentication information to the server.

[1375] Step 4:

[1376] The server checks the received authentication information against the database to verify that the user is a legitimate user.

[1377] Step 5:

[1378] The server sends the authentication result to the terminal. If the authentication is successful, an authentication success message is sent to the terminal and the user is successfully logged in. If the authentication is unsuccessful, an error message is sent to the terminal.

[1379] Step 6:

[1380] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[1381] Step 7:

[1382] The user inputs their feelings and state into the terminal and responds.

[1383] Step 8:

[1384] The terminal sends the user's answer to the server.

[1385] Step 9:

[1386] The emotion engine recognizes emotions from the user's voice and text, analyzing the user's tone of voice and expressions to determine their emotional state.

[1387] Step 10:

[1388] The server uses a generative AI model to analyze the user's responses and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[1389] Step 11:

[1390] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's state of mind and condition, taking into account the emotional information recognized by the emotion engine.

[1391] Step 12:

[1392] The server transmits the generated dialogue content to the terminal.

[1393] Step 13:

[1394] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in voice.

[1395] Step 14:

[1396] The user responds to the terminal.

[1397] Step 15:

[1398] The terminal sends the user's response to the server.

[1399] Step 16:

[1400] Based on the user's response, the server generates new dialogue content using a generative AI model and sends it to the device.

[1401] Step 17:

[1402] The terminal provides new dialogue content to the user using voice synthesis technology.

[1403] Step 18:

[1404] The terminal sequentially transmits the contents of the dialogue with the user and the user's responses to the server.

[1405] Step 19:

[1406] The server records the dialogue and the user's responses in a database.

[1407] Step 20:

[1408] The server uses the recorded data to update the generative AI model and learn to improve the accuracy of the next interaction.

[1409] As a specific example, the flow when the user answers "I'm a little tired" is shown below.

[1410] Step 6:

[1411] Terminal: "How are you feeling today?"

[1412] Step 7:

[1413] User: "I'm a little tired."

[1414] Step 8:

[1415] The terminal sends the user's answer to the server.

[1416] Step 9:

[1417] The emotion engine recognizes "fatigue" from the user's tone of voice.

[1418] Step 10:

[1419] The server uses information obtained from the generative AI model and emotion engine to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[1420] Step 12:

[1421] The server transmits the generated dialogue content to the terminal.

[1422] Step 13:

[1423] The device will read this aloud using voice synthesis technology.

[1424] Step 14:

[1425] The user responds, "I want to go to the park."

[1426] Step 15:

[1427] The terminal sends the user's response to the server.

[1428] Step 16:

[1429] The server generates new dialogue content such as "That's a good idea. Let's plan what we want to do in the park," and sends it to the device.

[1430] Step 17:

[1431] The terminal provides new dialogue content to the user using voice synthesis technology.

[1432] In this way, the system of the present invention, which is combined with an emotion engine, recognizes the user's emotional state in real time and provides dialogue accordingly, thereby realizing more personalized psychological support.

[1433] Example 2

[1434] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1435] In modern society, the number of users suffering from mental stress and anxiety is increasing, but there are only a limited number of systems that provide immediate and personalized responses. Conventional systems have difficulty accurately recognizing the user's emotional state and generating appropriate dialogue based on that, making it impossible to provide adequate mental care. Furthermore, technology for recognizing the user's emotions through voice or text and generating dialogue that reflects this is immature, resulting in a lack of satisfactory support for users.

[1436] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1437] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to confirm the user's feelings and state and receiving the answers, means for analyzing the user's answers and recognizing and evaluating the user's emotions using an emotion engine, means for using a generative AI model to generate optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, and means for recording the dialogue content and the user's responses and learning with the generative AI model. This makes it possible to accurately recognize the user's emotional state in real time and generate natural dialogue based on that.

[1438] "User authentication information" refers to the identifying information a user needs to log in to a system, typically a username and password.

[1439] "Emotion engine" refers to an algorithm or software that recognizes and analyzes emotions from a user's voice or text data in real time.

[1440] A "generative AI model" is an artificial intelligence model that learns large amounts of data and generates dialogue content with users, and uses natural language processing technology.

[1441] "Speech synthesis technology" refers to technology for generating natural speech based on text data, enabling information to be conveyed to users through speech.

[1442] "Dialogue content" refers to the response from the system generated in response to the user's question or status, and includes information such as appropriate support and advice.

[1443] "Emotional state" refers to a psychological state that indicates the type and intensity of the emotion the user is currently experiencing, and includes, for example, joy, sadness, anger, fatigue, and the like.

[1444] "Database" refers to a system for storing and managing data necessary for system operation, such as user authentication information, interaction history, and learning data for generative AI models.

[1445] "Real-time" refers to the instantaneous acquisition and analysis of data, generating an appropriate response immediately.

[1446] "Login screen" refers to the screen interface that is displayed to a user to enter their authentication information.

[1447] "Personalized response" refers to providing responses and support that are customized to suit the individual conditions and needs of the user.

[1448] This invention is a mental care system that provides mental support to users, and by combining it with an emotion engine, it is possible to recognize the user's emotional state in real time and to engage in dialogue and respond based on that. This system has a series of functions, including user authentication, understanding of feelings and state, emotion recognition, dialogue content generation, dialogue provision through voice synthesis, and dialogue data recording and learning.

[1449] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion engine has the function of recognizing the user's emotional state in real time. The user is the entity that interacts with the system through the terminal.

[1450] Hardware and software used

[1451] Hardware: Terminals (PCs, smartphones), servers

[1452] Software: Emotion engine, generative AI model, speech synthesis engine, database management system

[1453] Program processing

[1454] Initial Setup and User Authentication

[1455] When the terminal is started, it displays a login screen to the user. The user enters their authentication information (username and password), and the terminal sends this information to the server. The server compares the received authentication information with a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[1456] Understanding user status and recognizing emotions

[1457] After successful login, the device asks the user questions to ascertain their current mood and state (e.g., "How are you feeling today?"). The user inputs their mood and state into the device, and the answer is sent from the device to the server. The server uses an emotion engine to recognize emotions from the user's voice and text. The emotion engine analyzes the user's tone of voice and expressions to determine their emotional state. The server uses a generative AI model to analyze the user's answers and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[1458] Conversation content generation and provision

[1459] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's feelings and state. The emotional information recognized by the emotion engine is also taken into consideration during this process. The generated dialogue content is sent from the server to the device, which then uses speech synthesis technology to provide the generated dialogue content to the user in a voice format. This voice is conveyed to the user in a natural conversational format.

[1460] Data recording and learning

[1461] The device continuously transmits the user's interactions and responses to the server, which records them in a database and uses the recorded data to update the generative AI model. This learning allows the system to better respond to the user's individual needs.

[1462] Additional examples of specific actions

[1463] For example, if the user responds, "I'm a little tired," the process goes something like this: The device asks, "How are you feeling today?" The user responds, "I'm a little tired," and the device sends that information to the server. The emotion engine recognizes "fatigue" from the user's tone of voice, and the server uses the information obtained from the generative AI model and the emotion engine to generate a dialogue message saying, "We've checked your condition and will provide you with some advice to help you relax." The device reads this aloud using speech synthesis technology. The user responds, "I'd like to go to the park," and the device sends that response to the server. The server generates a new dialogue message saying, "That's a good idea. Let's plan what we want to do at the park," sends it to the device, and the device continues reading this aloud using speech synthesis technology.

[1464] Prompt Sentence Examples

[1465] "Generate a dialogue for when the user responds that they are a little tired."

[1466] "Provide relaxing advice based on the user's emotional state."

[1467] In this way, the present invention can recognize the user's feelings and state in real time and provide appropriate dialogue in response to them, thereby providing psychological support.

[1468] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1469] Mental health care system processing steps

[1470] Initial Setup and User Authentication

[1471] Step 1:

[1472] Starting the terminal

[1473] When the terminal is started, it displays a login screen to the user. The input is powering on, and the output is displaying the login screen.

[1474] Step 2:

[1475] Enter your authentication information

[1476] A user enters authentication information (username and password) into a login screen. The input is the authentication information, and the output is its preparation for submission.

[1477] Step 3:

[1478] Sending authentication information

[1479] The terminal sends authentication information to the server. The input is the authentication information and the output is a request sent to the server.

[1480] Step 4:

[1481] Authentication verification

[1482] The server checks the received authentication information against a database to verify the user's validity. The input is the authentication information, and the output is the authentication result (success or failure).

[1483] Step 5:

[1484] Return and display of authentication results

[1485] The server returns the authentication result to the terminal, which displays it to the user. The input is the authentication result, and the output is an indication of authentication success or failure. If successful, login is complete.

[1486] Understanding user status and recognizing emotions

[1487] Step 6:

[1488] Posing the Question

[1489] The terminal displays a question to the user to confirm their mood or state. For example, "How are you feeling today?" The input is a successful login status, and the output is the display of the question.

[1490] Step 7:

[1491] Enter your answer

[1492] The user inputs their feelings and state. The input is the user's answer to a question, and the output is the answer ready to be sent.

[1493] Step 8:

[1494] Submit your answer

[1495] The terminal sends the user's answer to the server. The input is the user's answer, and the output is the response sent to the server.

[1496] Step 9:

[1497] Performing emotion recognition

[1498] The server inputs the user's response into the emotion engine to recognize the user's emotion. The input is the user's response, and the output is the recognized emotion information.

[1499] Step 10:

[1500] Emotional information analysis

[1501] The server performs analysis using the results of the emotion engine and the generative AI model. The input is emotion information and the generative AI model, and the output is an evaluation of the user's mental state.

[1502] Conversation content generation and provision

[1503] Step 11:

[1504] Conversation generation

[1505] The server uses a generative AI model to generate optimal dialogue for the user. The input is the mental state assessment, and the output is the generated dialogue.

[1506] Step 12:

[1507] Sending conversation transcripts

[1508] The generated dialogue content is sent from the server to the terminal. The input is the generated dialogue content, and the output is the content sent to the terminal.

[1509] Step 13:

[1510] Provided by voice synthesis

[1511] The terminal uses speech synthesis technology to communicate the dialogue content to the user by voice. The input is the generated dialogue content, and the output is the voice output to the user.

[1512] Data recording and learning

[1513] Step 14:

[1514] Data recording

[1515] The terminal sends the dialogue content and the user's response to the server, which records it in a database. The input is the dialogue content and the user's response, and the output is the record in the database.

[1516] Step 15:

[1517] Training generative AI models

[1518] The server uses the recorded data to train the generative AI model to improve the accuracy of the next interaction. The input is the recorded data, and the output is an updated generative AI model.

[1519] (Application example 2)

[1520] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1521] In modern work environments, especially in factories, long hours of monotonous work can cause mental stress for workers. This can lead to reduced work efficiency and increased risk of health problems. Conventional mental health systems face the challenge of being unable to recognize and respond to workers' emotional states in real time. The present invention aims to solve these challenges and provide an advanced mental health system for supporting the mental health of workers.

[1522] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1523] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, means for recording the dialogue content and the user's responses and learning it with a generative AI model, means for converting the user's voice into text using speech recognition technology, emotion recognition means for detecting the user's emotional state in real time, and means for dynamically changing the dialogue content based on the responses. This allows the system to recognize the emotional state of workers in real time and provide appropriate dialogue, thereby providing psychological support and improving work efficiency and maintaining health.

[1524] "User authentication" is the process of receiving user authentication information and verifying that the user is legitimate.

[1525] "Mood confirmation" is a process of asking questions to confirm the user's mood or mental state.

[1526] "Answer analysis" is the process of analyzing a user's answer and evaluating its content.

[1527] "Dialogue content generation" is a process of generating optimal dialogue content for the user based on the analysis results.

[1528] "Speech synthesis technology" is a technology that outputs text information as voice.

[1529] "Data logging" is the process of saving the dialogue and user responses.

[1530] A "generative AI model" is an artificial intelligence model that learns from data and generates appropriate dialogue content.

[1531] "Speech recognition technology" is a technology that converts voice data into text data.

[1532] "Emotion recognition means" is a technology that detects the user's emotional state in real time from their voice and facial expressions.

[1533] The "dynamic dialogue change means" is a process that changes the dialogue content based on the user's response.

[1534] The present invention relates to a mental care system that combines user authentication, emotional state confirmation, emotion recognition, dialogue content generation, voice synthesis, data recording, and learning. Specific embodiments of each element will be described below.

[1535] Overall system configuration

[1536] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion recognition engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion recognition engine has the function of recognizing the user's emotional state in real time, and the user is the entity that interacts with the system through the terminal.

[1537] server

[1538] The server centrally processes data and runs the generative AI model. The server receives the user's authentication information and checks it against the database to verify that the user is legitimate. If authentication is successful, the server returns a successful authentication message to the user.

[1539] Terminal

[1540] The terminal accepts input from the user and communicates with the server. When the user inputs their feelings or state into the terminal, this information is sent to the server. If login is successful, the terminal displays questions to the user to confirm their current feelings or state.

[1541] Emotion Recognition Engine

[1542] The emotion recognition engine works in conjunction with voice recognition technology to analyze the user's tone and expressions to recognize their emotional state, which is then sent to the server and used to assess the user's mental state.

[1543] User

[1544] The user interacts with the system through the device. When the user inputs their feelings and state, the information is sent to the server and analyzed by an emotion recognition engine. The server then uses a generative AI model to generate optimal dialogue content and send it to the device.

[1545] Detailed process flow

[1546] 1. User authentication:

[1547] The user enters a username and password into the terminal and sends them to the server.

[1548] The server checks the authentication information against a database to verify the user is a valid user.

[1549] If the authentication is successful, the server sends an authentication success message to the terminal, completing the login.

[1550] 2. Emotional validation and emotional recognition:

[1551] If authentication is successful, the terminal presents the user with questions to confirm their feelings and state of mind, and receives their answers.

[1552] The emotion recognition engine analyzes the user's voice and recognizes their emotional state in real time.

[1553] The server uses a generative AI model to analyze the user's responses and emotional information to assess their mental state.

[1554] 3. Generating and providing dialogue content:

[1555] The server uses a generative AI model based on the analysis results to generate optimal dialogue content.

[1556] The generated dialogue content is sent from the server to the terminal and provided to the user using voice synthesis technology.

[1557] 4. Data recording and learning:

[1558] The dialogue and the user's responses are sent to the server and recorded in a database.

[1559] The recorded data is used to train the generative AI model to improve the accuracy of the next interaction.

[1560] Hardware and software used

[1561] Server: Database, generative AI models (e.g., Hugging Face Transformers)

[1562] Devices: Smartphones, PCs, tablets

[1563] Emotion recognition engine: Speech recognition technology (e.g., Google Speech Recognition), emotion analysis (e.g., Sentiment Analysis Pipeline)

[1564] Speech synthesis: Pyttsx3 (e.g., a Python text-to-speech library)

[1565] Examples and prompts

[1566] Example 1:

[1567] Please enter your username: user123

[1568] Please enter your password: password123

[1569] Authentication successful. How are you feeling today?

[1570] (User speaks): "I'm a little tired."

[1571] The system responds: "You sound a little down. Let me know if there's anything I can do to help."

[1572] Example 2:

[1573] Please enter your username: sampleuser

[1574] Please enter your password: mypassword

[1575] Authentication successful. How are you feeling today?

[1576] (User speaks): "I feel great today."

[1577] The system responds: "Great! Keep it up!"

[1578] This allows the system to recognize the emotional state of workers in real time and provide appropriate dialogue to provide psychological support, thereby improving work efficiency and maintaining health.

[1579] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1580] Step 1:

[1581] User Authentication

[1582] Input: The user enters a username and password into the terminal.

[1583] Specific operation: The device sends these authentication information to the server.

[1584] Data processing / calculation: The server checks the authentication information against a database to verify that the user is legitimate.

[1585] Output: If the authentication is successful, the server returns an authentication success message to the terminal; if the authentication fails, it returns an error message.

[1586] Step 2:

[1587] Confirmation of feelings

[1588] Input: Receives information that authentication was successful.

[1589] Specific behavior: The device presents the user with a question to ascertain their current state of mind (e.g., "How are you feeling today?").

[1590] Data processing / calculation: Receive user responses via voice or text input.

[1591] Output: The device sends the user's answer to the server.

[1592] Step 3:

[1593] emotion recognition

[1594] Input: The server receives the user's answer and voice data.

[1595] How it works: The emotion recognition engine analyzes the user's tone of voice and expressions to recognize their emotional state in real time.

[1596] Data processing / calculation: The analysis results are sent to the server, and the emotion data is evaluated using a generative AI model.

[1597] Output: The recognized emotional state and analysis results are output.

[1598] Step 4:

[1599] Dialogue content generation

[1600] Input: Analyzed user emotional state and sentiment data.

[1601] Specific operation: The server uses a generative AI model to generate optimal dialogue content based on the user's emotions and feelings.

[1602] Data processing / calculation: The AI ​​model generates dialogue content based on emotion recognition results and sentiment data.

[1603] Output: The generated dialogue content is sent from the server to the terminal.

[1604] Step 5:

[1605] Dialogue provision

[1606] Input: The generated dialogue.

[1607] Specific operation: The device uses speech synthesis technology to provide the generated dialogue content to the user via voice.

[1608] Data processing / calculation: Converting text data into audio data.

[1609] Output: The dialogue is presented to the user via audio.

[1610] Step 6:

[1611] Data Recording

[1612] Input: Dialogue and user response data.

[1613] Specific operation: The dialogue content sent from the terminal to the server and the user's response data are recorded in a database.

[1614] Data processing / calculation: Recorded data is used as training data for generative AI models.

[1615] Output: Interaction data stored in a database and an updated AI model.

[1616] This allows the system to recognize the user's emotional state in real time and provide appropriate dialogue, thereby providing psychological support, improving work efficiency, and maintaining health.

[1617] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1618] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1619] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1620] [Fourth embodiment]

[1621] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1622] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1623] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1624] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1625] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1626] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1627] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1628] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1629] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1630] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1631] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1632] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1633] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1634] The present invention is a mental care system that provides mental support to users. This system has a series of functions, such as user authentication, grasping of feelings and conditions, generation of dialogue content, dialogue provision by voice synthesis, and recording and learning of dialogue data.

[1635] Overall system overview

[1636] The system is primarily composed of three elements: a server, a terminal, and a user. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[1637] 1. Initial Setup and User Authentication

[1638] 1. Start the device

[1639] When the terminal is started, it presents the user with a login screen, where the user enters their authentication information (username and password), which is then sent from the terminal to the server.

[1640] 2. Authentication Process

[1641] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[1642] 2. Understanding the user's status

[1643] 1. Posing the Question

[1644] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[1645] 2. Receiving and sending responses

[1646] The user inputs their feelings and state into the terminal, and the response is sent from the terminal to the server.

[1647] 3. Analysis of responses

[1648] The server analyzes the user's responses using a generative AI model, and based on the results of this analysis, evaluates the user's current mental state.

[1649] 3. Generating and providing conversation content

[1650] 1. Conversation content generation

[1651] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's mood and state.

[1652] 2. Sending conversation content

[1653] The generated dialogue content is transmitted from the server to the terminal.

[1654] 3. Provision by voice synthesis

[1655] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in a voice that is conveyed to the user in a natural conversational style.

[1656] 4. Data recording and learning

[1657] 1. Data recording

[1658] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which records them in a database.

[1659] 2. Training the generative AI model

[1660] The server uses the recorded data to update the generative AI model to improve the accuracy of the next interaction. This learning allows the system to better respond to the user's individual needs.

[1661] Specific examples

[1662] If the user responds "I'm a little tired"

[1663] 1. Terminal: "How are you feeling today?"

[1664] 2. User: "I'm a little tired."

[1665] 3. The device sends this information to the server.

[1666] 4. The server uses the generative AI model to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[1667] 5. The device will read this aloud using speech synthesis technology.

[1668] 6. The user responds, "I want to go to the park."

[1669] 7. The device sends this response to the server.

[1670] 8. The server generates new dialogue such as "That's a good idea. Let's plan what we want to do at the park." and sends it to the device.

[1671] 9. The device will read this aloud using speech synthesis technology and the conversation will continue.

[1672] In this way, the system of the present invention responds to the user's feelings and state in real time and provides appropriate dialogue, thereby providing psychological support.

[1673] The processing flow will be explained below.

[1674] Step 1:

[1675] When the terminal is started, it displays a login screen to the user.

[1676] Step 2:

[1677] The user enters a username and password into the terminal.

[1678] Step 3:

[1679] The terminal sends the user's authentication information to the server.

[1680] Step 4:

[1681] The server checks the received authentication information against the database to verify that the user is a legitimate user.

[1682] Step 5:

[1683] The server sends the authentication result to the terminal. If the authentication is successful, an authentication success message is sent to the terminal and the user is successfully logged in. If the authentication is unsuccessful, an error message is sent to the terminal.

[1684] Step 6:

[1685] Once the login is successful, the terminal will ask the user questions to confirm their current state of mind and condition.

[1686] Step 7:

[1687] The user inputs their feelings and state into the terminal and responds.

[1688] Step 8:

[1689] The terminal sends the user's answer to the server.

[1690] Step 9:

[1691] The server uses a generative AI model to analyze the user's responses and evaluate the user's mental state.

[1692] Step 10:

[1693] The server generates dialogue content based on the analysis results using an AI model.

[1694] Step 11:

[1695] The server transmits the generated dialogue content to the terminal.

[1696] Step 12:

[1697] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in voice.

[1698] Step 13:

[1699] The user responds to the terminal.

[1700] Step 14:

[1701] The terminal sends the user's response to the server.

[1702] Step 15:

[1703] Based on the user's response, the server generates new dialogue content using a generative AI model and sends it to the device.

[1704] Step 16:

[1705] The terminal provides new dialogue content to the user using voice synthesis technology.

[1706] Step 17:

[1707] The terminal sequentially transmits the contents of the dialogue with the user and the user's responses to the server.

[1708] Step 18:

[1709] The server records the dialogue and the user's responses in a database.

[1710] Step 19:

[1711] The server uses the recorded data to update the generative AI model and learn to improve the accuracy of the next interaction.

[1712] Example 1

[1713] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1714] In modern society, many people suffer from mental stress and fatigue, and appropriate mental care is needed to alleviate these burdens. However, conventional mental care systems have had difficulty accurately grasping the user's mental state and providing dialogue content tailored to individual needs. In addition, the methods for providing dialogue content are limited, making it difficult to achieve natural and effective support.

[1715] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1716] In this invention, the server includes means for receiving and authenticating a user's authentication information, means for presenting questions to ascertain the user's mental state and receiving the answers, means for analyzing the user's answers using a generative AI model and generating appropriate dialogue content, means for providing the generated dialogue content to the user using speech synthesis technology, and means for recording the dialogue content and the user's responses and learning the dialogue content with the generative AI model, thereby enabling the provision of appropriate psychological support to the user in real time.

[1717] "User" refers to an individual who uses the system to receive psychological support.

[1718] "Authentication information" is information used to identify a user and authorize access to a system, and typically includes a username and password.

[1719] "Means" refers to a method, technique, or device for accomplishing a particular function.

[1720] "Mental state" refers to the user's emotional and psychological state, including stress and fatigue.

[1721] "Question" refers to a query posed by the system to ascertain the user's mental state.

[1722] A "generative AI model" refers to a model that uses machine learning and artificial intelligence technology to analyze data and automatically generate dialogue content.

[1723] "Analysis" refers to the process of evaluating the user's mental state based on their answers and deriving appropriate dialogue content.

[1724] "Dialogue content" refers to the advice and response text provided to the user.

[1725] "Speech synthesis technology" refers to technology for converting text data into speech.

[1726] "Recording" refers to the act of the system saving the content of interactions and responses with the user.

[1727] "Learning" refers to the process by which the system improves the generated AI model based on past data, improving the accuracy of the next interaction.

[1728] The present invention provides a mental care system that provides mental support to users. This system has a series of functions, such as user authentication, understanding of emotions and conditions, generating dialogue content, providing dialogue through voice synthesis, and recording and learning dialogue data.

[1729] Overall system overview

[1730] The system is primarily composed of three elements: a server, a terminal, and a user. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[1731] Hardware and software used

[1732] Server: Responsible for data processing and operation of the generative AI model, authenticating users, analyzing responses, and generating dialogue content.

[1733] Terminal: Presents questions to the user, receives answers from the user, and provides the generated dialogue content through voice synthesis.

[1734] Generative AI model: Uses artificial intelligence techniques to generate dialogue content based on user input data.

[1735] Speech synthesis technology: A speech synthesis engine is used to convert text data into speech.

[1736] Specific examples

[1737] If the user answers "I'm a little tired," the specific sequence of events is as follows:

[1738] 1. Terminal: Display the question to the user: "How are you feeling today?"

[1739] 2. User: Type "I'm a little tired."

[1740] 3. The device sends this information to the server.

[1741] 4. The server uses the generative AI model to analyze the answer "I'm a little tired" and assess the user's condition.

[1742] 5. Based on the analysis results, the server generates dialogue content such as, "We will provide you with some advice to help you relax."

[1743] 6. The server sends the generated dialogue content to the terminal.

[1744] 7. The device uses speech synthesis technology to read out the generated dialogue in a natural voice, saying, "We'll give you some advice to help you relax."

[1745] 8. The user may ask further questions about how to relax, in which case the device sends a response to the server, which again uses the generative AI model to generate appropriate dialogue.

[1746] When implementing this system, it is also important to consider the prompt sentences to generate prepared questions. For example, the following prompt sentences can be used in response to input from the user:

[1747] Example prompt sentence:

[1748] "I'm a little tired. How can I relax?"

[1749] The mental care system of the present invention provides mental support by grasping the user's feelings and state in real time and providing appropriate dialogue, allowing the user to easily receive advice and counseling to reduce stress and fatigue in daily life.

[1750] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1751] Program processing flow

[1752] Step 1: Initial configuration and user authentication

[1753] 1.1 Starting the terminal

[1754] When the terminal is started, it displays a login screen. The user enters a username and password. This authentication information is sent from the terminal to the server.

[1755] Input: Username, Password

[1756] Output: Sending authentication information

[1757] 1.2 Authentication process

[1758] The server checks the received authentication information against its database, checking the username and password to ensure the user is a valid user, and sending a success message if authentication is successful, or an error message if authentication is unsuccessful.

[1759] Input: Credentials

[1760] Data processing: Database collation

[1761] Output: Authentication success message or error message

[1762] 1.3 Displaying authentication results

[1763] The terminal receives the message from the server and displays the success or failure of the authentication to the user.

[1764] Input: Authentication success message or error message

[1765] Output: Display of authentication result

[1766] Step 2: Understanding the user's state

[1767] 2.1 Posing the Question

[1768] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[1769] Input: Login successful

[1770] Output: Question posed

[1771] 2.2 Receiving a response

[1772] The user inputs their feelings and state in response to the questions presented, and the answers are sent from the device to the server.

[1773] Input: User's answer

[1774] Output: Sending the answer

[1775] 2.3 Analysis of responses

[1776] The server inputs the received answers into a generative AI model for analysis. For example, if the user answers something like "I'm a little tired," it analyzes that information to evaluate the user's current mental state.

[1777] Input: User's answer

[1778] Data Computation: Analysis with Generative AI Models

[1779] Output: Mental state assessment

[1780] Step 3: Generate and provide conversation content

[1781] 3.1 Conversation content generation

[1782] Based on the evaluation results, the server uses a generative AI model to generate optimal dialogue content for the user. For example, if the user is evaluated as tired, the server generates dialogue content such as "We will give you advice on how to relax."

[1783] Input: Mental status assessment

[1784] Data Computation: Dialogue Content Generation with Generative AI Models

[1785] Output: Generated dialogue

[1786] 3.2 Transmission of dialogue content

[1787] The server transmits the generated dialogue content to the terminal.

[1788] Input: Generated dialogue

[1789] Output: Sending dialogue

[1790] 3.3 Provision by voice synthesis

[1791] The terminal converts the received dialogue content into voice using speech synthesis technology and provides it to the user.

[1792] Input: Generated dialogue

[1793] Data processing: voice synthesis

[1794] Output: Provides dialogue via voice

[1795] Step 4: Data recording and learning

[1796] 4.1 Data recording

[1797] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which then records them in a database.

[1798] Input: Dialogue content, user response

[1799] Output: Data recording

[1800] 4.2 Training generative AI models

[1801] The server uses the recorded data to update the generative AI model, which improves the accuracy of the next interaction.

[1802] Input: Recorded data

[1803] Data Computation: Learning Generative AI Models

[1804] Output: Updated generative AI model

[1805] Through the above processing steps, the mental care system of the present invention has the ability to grasp the user's feelings and state in real time and provide appropriate dialogue.

[1806] (Application example 1)

[1807] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1808] In modern society, the number of users who need mental support is increasing, but there is a lack of systems that can individually generate dialogue content and provide it in an appropriate format.In addition, while there is a demand for effective mental care using smart devices, there is currently no system that can provide appropriate dialogue and content based on the user's mood and state.This poses the problem that users cannot receive support that is appropriate for their own physical and mental state.

[1809] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1810] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, means for recording the dialogue content and the user's responses and learning them using a generative AI model, and means for providing dialogue content and content generated by AI based on the user's feelings and state. This enables the user to receive support in real time according to their own feelings and state.

[1811] "User credentials" are the identifying information provided by a user to log into a system.

[1812] "Mood" refers to the user's feelings and moods.

[1813] "Condition" refers to the physical or mental condition of a user.

[1814] "Questions" are questions that the system presents to ascertain the user's feelings and state of mind.

[1815] An "answer" is information that a user enters in response to a question.

[1816] "Analysis" refers to the processing of data to evaluate users' responses and understand their sentiments and state of mind.

[1817] "Dialogue content" refers to messages consisting of text or voice used in dialogue with a user.

[1818] "Speech synthesis technology" is a technology that converts text into a voice that sounds like a human voice.

[1819] "User response" is information that the user responds to the dialogue content presented by the system.

[1820] "Recording" means saving the dialogue content and user responses in a database.

[1821] A "generative AI model" is an artificial intelligence model that generates appropriate dialogue content based on user information.

[1822] "Content" refers to information material such as video, music, text, etc., provided to users.

[1823] "Dialogue content generated by AI based on emotions and state" refers to messages created by a generative AI model based on the user's emotions and state.

[1824] This invention is a content distribution system that provides psychological support to users. The system is mainly composed of three elements: a server, a terminal, and a user. The server is responsible for data processing and operation of the generative AI model, and the terminal is a device that provides an interface with the user. The user is the entity that interacts with the system through the terminal.

[1825] Server configuration and functions

[1826] The server includes means for receiving and authenticating a user's authentication information, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for transmitting the generated dialogue content to the terminal, means for recording the dialogue content and the user's responses and learning using a generative AI model, and means for providing dialogue content and content generated by AI based on the user's feelings and state.

[1827] The server uses Python and Flask as the web framework. The generative AI model is implemented using the OpenAI API. It receives and analyzes user data to generate dialogue content, which is then sent to the device. Furthermore, the dialogue data with the user is recorded in a database and used as training data for the generative AI model.

[1828] Device configuration and functions

[1829] The terminal is a device operated by the user, such as a smartphone or tablet, that provides a variety of functions. The terminal displays a user authentication screen and sends authentication information from the user to the server. If authentication is successful, the terminal displays questions to confirm the user's feelings and state, and sends the user's answers to the server. The dialogue received from the server is read aloud using speech synthesis technology.

[1830] The voice synthesis technology uses the Google Text-to-Speech API, which allows the generated dialogue to be presented to the user in a natural conversational format. The device also displays and plays relaxing content such as videos and music.

[1831] User actions and responses

[1832] Users operate their terminals to log in to the system. After logging in, they can report their own feelings and state by answering questions posed by the system. The system generates and provides appropriate dialogue and content based on this information. Users can receive psychological support by using the dialogue and content provided.

[1833] Specific examples

[1834] When a user logs in to the system and responds, "I'm a little tired," the server sends a dialogue to the device, such as, "We've checked your condition. We'll give you some advice to help you relax." This dialogue is read aloud to the user using voice synthesis technology. If the user then responds, "I'd like to go to the park," the server generates a new dialogue, such as, "That's a good idea. Let's plan what we want to do at the park," and sends it to the device.

[1835] Prompt Sentence Examples

[1836] "User is a little fatigued. Please generate a dialogue that offers a support conversation."

[1837] As described above, the system of the present invention can provide psychological support in real time according to the user's feelings and condition.

[1838] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1839] Step 1:

[1840] When the terminal is started, a login screen is displayed. The user enters authentication information (username and password) which the terminal sends to the server. The input is the username and password, and the output is the authentication information sent to the server.

[1841] Step 2:

[1842] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal. The input is authentication information, and the output is an authentication success message or an error message.

[1843] Step 3:

[1844] If authentication is successful, the device asks the user a question to confirm their current state of mind. The question is in the form of "How are you feeling today?" The input is a successful authentication message from the server, and the output is the question displayed to the user.

[1845] Step 4:

[1846] The user inputs their feelings or state of mind into the terminal. For example, they reply, "I'm a little tired." The input is text describing the user's feelings or state, and the output is the transmission of the reply data to the server.

[1847] Step 5:

[1848] The server inputs the received user responses into a generative AI model for analysis. This analysis evaluates the user's feelings and state and generates appropriate dialogue content. The input is the user's response, and the output is the analysis results and the generated dialogue content. Specifically, it sends a prompt to the OpenAI API and receives a response.

[1849] Step 6:

[1850] The server sends the generated dialogue content to the terminal. A response message containing the generated dialogue content is output. The input is the analysis result of the AI ​​model, and the output is the dialogue content sent to the terminal.

[1851] Step 7:

[1852] The device uses speech synthesis technology to provide the received dialogue to the user. The speech synthesis technology used is the Google Text-to-Speech API. The input is the dialogue from the server, and the output is the result presented to the user as voice.

[1853] Step 8:

[1854] The user responds to the dialogue content heard by voice and inputs the response into the terminal. The input is the user's response text, and the output is the transmission of the response data to the server.

[1855] Step 9:

[1856] The server records the received user responses in a database and uses them as training data for the generative AI model. The input is the user response data, and the output is an updated AI model. Specific operations include recording to the database and retraining the generative AI model.

[1857] As described above, this system performs a series of processes, starting with user authentication, then understanding the user's feelings and state, generating dialogue content using a generative AI model, providing dialogue through voice synthesis, and finally recording data and learning the model. Through this process, it is possible to provide individualized support according to the user's state.

[1858] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1859] This invention is a mental care system that provides mental support to users, and by combining it with an emotion engine, it is possible to recognize the user's emotional state in real time and to engage in dialogue and respond based on that. The system has a series of functions, including user authentication, understanding of feelings and state, emotion recognition, dialogue content generation, dialogue provision through voice synthesis, and dialogue data recording and learning.

[1860] Overall system overview

[1861] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion engine has the function of recognizing the user's emotional state in real time. The user is the entity that interacts with the system through the terminal.

[1862] 1. Initial Setup and User Authentication

[1863] 1. Start the device

[1864] When the terminal is started, it presents the user with a login screen, where the user enters their authentication information (username and password), which is then sent from the terminal to the server.

[1865] 2. Authentication Process

[1866] The server checks the received authentication information against a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[1867] 2. Understanding user state and emotion recognition

[1868] 1. Posing the Question

[1869] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[1870] 2. Receiving and sending responses

[1871] The user inputs their feelings and state into the terminal, and the response is sent from the terminal to the server.

[1872] 3. Emotion recognition

[1873] The emotion engine recognizes emotions from the user's voice and text. The emotion engine analyzes the user's tone of voice and expressions to determine their emotional state.

[1874] 4. Analysis of responses and emotional information

[1875] The server uses a generative AI model to analyze the user's responses and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[1876] 3. Generating and providing conversation content

[1877] 1. Conversation content generation

[1878] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's state of mind and condition, taking into account the emotional information recognized by the emotion engine.

[1879] 2. Sending conversation content

[1880] The generated dialogue content is transmitted from the server to the terminal.

[1881] 3. Provision by voice synthesis

[1882] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in a voice that is conveyed to the user in a natural conversational style.

[1883] 4. Data recording and learning

[1884] 1. Data recording

[1885] The terminal sequentially transmits the contents of the conversation with the user and the user's responses to the server, which records them in a database.

[1886] 2. Training the generative AI model

[1887] The server uses the recorded data to update the generative AI model to improve the accuracy of the next interaction. This learning allows the system to better respond to the user's individual needs.

[1888] Specific examples

[1889] If the user responds "I'm a little tired"

[1890] 1. Terminal: "How are you feeling today?"

[1891] 2. User: "I'm a little tired."

[1892] 3. The device sends this information to the server.

[1893] 4. The emotion engine recognizes "fatigue" from the user's tone of voice.

[1894] 5. The server uses the information obtained from the generated AI model and emotion engine to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[1895] 6. The device will read this aloud using speech synthesis technology.

[1896] 7. The user responds, "I want to go to the park."

[1897] 8. The device sends this response to the server.

[1898] 9. The server generates new dialogue such as "That's a good idea. Let's plan what we want to do at the park." and sends it to the device.

[1899] 10. The device will read this aloud using speech synthesis technology and the conversation will continue.

[1900] In this way, the system of the present invention, which combines an emotion engine, recognizes the user's feelings and state in real time and provides appropriate dialogue to support the user emotionally.Furthermore, by generating dialogue content that reflects the user's emotional state, more personalized responses are possible.

[1901] The processing flow will be explained below.

[1902] Step 1:

[1903] When the terminal is started, it displays a login screen to the user.

[1904] Step 2:

[1905] The user enters a username and password into the terminal.

[1906] Step 3:

[1907] The terminal sends the user's authentication information to the server.

[1908] Step 4:

[1909] The server checks the received authentication information against the database to verify that the user is a legitimate user.

[1910] Step 5:

[1911] The server sends the authentication result to the terminal. If the authentication is successful, an authentication success message is sent to the terminal and the user is successfully logged in. If the authentication is unsuccessful, an error message is sent to the terminal.

[1912] Step 6:

[1913] Once the login is successful, the device will ask the user questions to ascertain their current state of mind (e.g., "How are you feeling today?").

[1914] Step 7:

[1915] The user inputs their feelings and state into the terminal and responds.

[1916] Step 8:

[1917] The terminal sends the user's answer to the server.

[1918] Step 9:

[1919] The emotion engine recognizes emotions from the user's voice and text, analyzing the user's tone of voice and expressions to determine their emotional state.

[1920] Step 10:

[1921] The server uses a generative AI model to analyze the user's responses and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[1922] Step 11:

[1923] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's state of mind and condition, taking into account the emotional information recognized by the emotion engine.

[1924] Step 12:

[1925] The server transmits the generated dialogue content to the terminal.

[1926] Step 13:

[1927] The terminal uses speech synthesis technology to provide the generated dialogue content to the user in voice.

[1928] Step 14:

[1929] The user responds to the terminal.

[1930] Step 15:

[1931] The terminal sends the user's response to the server.

[1932] Step 16:

[1933] Based on the user's response, the server generates new dialogue content using a generative AI model and sends it to the device.

[1934] Step 17:

[1935] The terminal provides new dialogue content to the user using voice synthesis technology.

[1936] Step 18:

[1937] The terminal sequentially transmits the contents of the dialogue with the user and the user's responses to the server.

[1938] Step 19:

[1939] The server records the dialogue and the user's responses in a database.

[1940] Step 20:

[1941] The server uses the recorded data to update the generative AI model and learn to improve the accuracy of the next interaction.

[1942] As a specific example, the flow when the user answers "I'm a little tired" is shown below.

[1943] Step 6:

[1944] Terminal: "How are you feeling today?"

[1945] Step 7:

[1946] User: "I'm a little tired."

[1947] Step 8:

[1948] The terminal sends the user's answer to the server.

[1949] Step 9:

[1950] The emotion engine recognizes "fatigue" from the user's tone of voice.

[1951] Step 10:

[1952] The server uses information obtained from the generative AI model and emotion engine to generate dialogue such as, "We have checked your condition. We will provide you with some advice to help you relax."

[1953] Step 12:

[1954] The server transmits the generated dialogue content to the terminal.

[1955] Step 13:

[1956] The device will read this aloud using voice synthesis technology.

[1957] Step 14:

[1958] The user responds, "I want to go to the park."

[1959] Step 15:

[1960] The terminal sends the user's response to the server.

[1961] Step 16:

[1962] The server generates new dialogue content such as "That's a good idea. Let's plan what we want to do in the park," and sends it to the device.

[1963] Step 17:

[1964] The terminal provides new dialogue content to the user using voice synthesis technology.

[1965] In this way, the system of the present invention, which is combined with an emotion engine, recognizes the user's emotional state in real time and provides dialogue accordingly, thereby realizing more personalized psychological support.

[1966] Example 2

[1967] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1968] In modern society, the number of users suffering from mental stress and anxiety is increasing, but there are only a limited number of systems that provide immediate and personalized responses. Conventional systems have difficulty accurately recognizing the user's emotional state and generating appropriate dialogue based on that, making it impossible to provide adequate mental care. Furthermore, technology for recognizing the user's emotions through voice or text and generating dialogue that reflects this is immature, resulting in a lack of satisfactory support for users.

[1969] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1970] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to confirm the user's feelings and state and receiving the answers, means for analyzing the user's answers and recognizing and evaluating the user's emotions using an emotion engine, means for using a generative AI model to generate optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, and means for recording the dialogue content and the user's responses and learning with the generative AI model. This makes it possible to accurately recognize the user's emotional state in real time and generate natural dialogue based on that.

[1971] "User authentication information" refers to the identifying information a user needs to log in to a system, typically a username and password.

[1972] "Emotion engine" refers to an algorithm or software that recognizes and analyzes emotions from a user's voice or text data in real time.

[1973] A "generative AI model" is an artificial intelligence model that learns large amounts of data and generates dialogue content with users, and uses natural language processing technology.

[1974] "Speech synthesis technology" refers to technology for generating natural speech based on text data, enabling information to be conveyed to users through speech.

[1975] "Dialogue content" refers to the response from the system generated in response to the user's question or status, and includes information such as appropriate support and advice.

[1976] "Emotional state" refers to a psychological state that indicates the type and intensity of the emotion the user is currently experiencing, and includes, for example, joy, sadness, anger, fatigue, and the like.

[1977] "Database" refers to a system for storing and managing data necessary for system operation, such as user authentication information, interaction history, and learning data for generative AI models.

[1978] "Real-time" refers to the instantaneous acquisition and analysis of data, generating an appropriate response immediately.

[1979] "Login screen" refers to the screen interface that is displayed to a user to enter their authentication information.

[1980] "Personalized response" refers to providing responses and support that are customized to suit the individual conditions and needs of the user.

[1981] This invention is a mental care system that provides mental support to users, and by combining it with an emotion engine, it is possible to recognize the user's emotional state in real time and to engage in dialogue and respond based on that. This system has a series of functions, including user authentication, understanding of feelings and state, emotion recognition, dialogue content generation, dialogue provision through voice synthesis, and dialogue data recording and learning.

[1982] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion engine has the function of recognizing the user's emotional state in real time. The user is the entity that interacts with the system through the terminal.

[1983] Hardware and software used

[1984] Hardware: Terminals (PCs, smartphones), servers

[1985] Software: Emotion engine, generative AI model, speech synthesis engine, database management system

[1986] Program processing

[1987] Initial Setup and User Authentication

[1988] When the terminal is started, it displays a login screen to the user. The user enters their authentication information (username and password), and the terminal sends this information to the server. The server compares the received authentication information with a database to verify that the user is a valid user. If authentication is successful, the server returns an authentication success message to the terminal, and the user is successfully logged in. If authentication fails, the server returns an error message, which the terminal displays to the user.

[1989] Understanding user status and recognizing emotions

[1990] After successful login, the device asks the user questions to ascertain their current mood and state (e.g., "How are you feeling today?"). The user inputs their mood and state into the device, and the answer is sent from the device to the server. The server uses an emotion engine to recognize emotions from the user's voice and text. The emotion engine analyzes the user's tone of voice and expressions to determine their emotional state. The server uses a generative AI model to analyze the user's answers and the emotional information obtained from the emotion engine to evaluate the user's current mental state.

[1991] Conversation content generation and provision

[1992] Based on the analysis results, the server uses a generative AI model to generate optimal dialogue content tailored to the user's feelings and state. The emotional information recognized by the emotion engine is also taken into consideration during this process. The generated dialogue content is sent from the server to the device, which then uses speech synthesis technology to provide the generated dialogue content to the user in a voice format. This voice is conveyed to the user in a natural conversational format.

[1993] Data recording and learning

[1994] The device continuously transmits the user's interactions and responses to the server, which records them in a database and uses the recorded data to update the generative AI model. This learning allows the system to better respond to the user's individual needs.

[1995] Additional examples of specific actions

[1996] For example, if the user responds, "I'm a little tired," the process goes something like this: The device asks, "How are you feeling today?" The user responds, "I'm a little tired," and the device sends that information to the server. The emotion engine recognizes "fatigue" from the user's tone of voice, and the server uses the information obtained from the generative AI model and the emotion engine to generate a dialogue message saying, "We've checked your condition and will provide you with some advice to help you relax." The device reads this aloud using speech synthesis technology. The user responds, "I'd like to go to the park," and the device sends that response to the server. The server generates a new dialogue message saying, "That's a good idea. Let's plan what we want to do at the park," sends it to the device, and the device continues reading this aloud using speech synthesis technology.

[1997] Prompt Sentence Examples

[1998] "Generate a dialogue for when the user responds that they are a little tired."

[1999] "Provide relaxing advice based on the user's emotional state."

[2000] In this way, the present invention can recognize the user's feelings and state in real time and provide appropriate dialogue in response to them, thereby providing psychological support.

[2001] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2002] Mental health care system processing steps

[2003] Initial Setup and User Authentication

[2004] Step 1:

[2005] Starting the terminal

[2006] When the terminal is started, it displays a login screen to the user. The input is powering on, and the output is displaying the login screen.

[2007] Step 2:

[2008] Enter your authentication information

[2009] A user enters authentication information (username and password) into a login screen. The input is the authentication information, and the output is its preparation for submission.

[2010] Step 3:

[2011] Sending authentication information

[2012] The terminal sends authentication information to the server. The input is the authentication information and the output is a request sent to the server.

[2013] Step 4:

[2014] Authentication verification

[2015] The server checks the received authentication information against a database to verify the user's validity. The input is the authentication information, and the output is the authentication result (success or failure).

[2016] Step 5:

[2017] Return and display of authentication results

[2018] The server returns the authentication result to the terminal, which displays it to the user. The input is the authentication result, and the output is an indication of authentication success or failure. If successful, login is complete.

[2019] Understanding user status and recognizing emotions

[2020] Step 6:

[2021] Posing the Question

[2022] The terminal displays a question to the user to confirm their mood or state. For example, "How are you feeling today?" The input is a successful login status, and the output is the display of the question.

[2023] Step 7:

[2024] Enter your answer

[2025] The user inputs their feelings and state. The input is the user's answer to a question, and the output is the answer ready to be sent.

[2026] Step 8:

[2027] Submit your answer

[2028] The terminal sends the user's answer to the server. The input is the user's answer, and the output is the response sent to the server.

[2029] Step 9:

[2030] Performing emotion recognition

[2031] The server inputs the user's response into the emotion engine to recognize the user's emotion. The input is the user's response, and the output is the recognized emotion information.

[2032] Step 10:

[2033] Emotional information analysis

[2034] The server performs analysis using the results of the emotion engine and the generative AI model. The input is emotion information and the generative AI model, and the output is an evaluation of the user's mental state.

[2035] Conversation content generation and provision

[2036] Step 11:

[2037] Conversation generation

[2038] The server uses a generative AI model to generate optimal dialogue for the user. The input is the mental state assessment, and the output is the generated dialogue.

[2039] Step 12:

[2040] Sending conversation transcripts

[2041] The generated dialogue content is sent from the server to the terminal. The input is the generated dialogue content, and the output is the content sent to the terminal.

[2042] Step 13:

[2043] Provided by voice synthesis

[2044] The terminal uses speech synthesis technology to communicate the dialogue content to the user by voice. The input is the generated dialogue content, and the output is the voice output to the user.

[2045] Data recording and learning

[2046] Step 14:

[2047] Data recording

[2048] The terminal sends the dialogue content and the user's response to the server, which records it in a database. The input is the dialogue content and the user's response, and the output is the record in the database.

[2049] Step 15:

[2050] Training generative AI models

[2051] The server uses the recorded data to train the generative AI model to improve the accuracy of the next interaction. The input is the recorded data, and the output is an updated generative AI model.

[2052] (Application example 2)

[2053] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2054] In modern work environments, especially in factories, long hours of monotonous work can cause mental stress for workers. This can lead to reduced work efficiency and increased risk of health problems. Conventional mental health systems face the challenge of being unable to recognize and respond to workers' emotional states in real time. The present invention aims to solve these challenges and provide an advanced mental health system for supporting the mental health of workers.

[2055] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2056] In this invention, the server includes means for receiving user authentication information and performing authentication, means for asking questions to ascertain the user's feelings and state and receiving the answers, means for analyzing the user's answers and generating optimal dialogue content for the user, means for providing the generated dialogue content to the user using speech synthesis technology, means for recording the dialogue content and the user's responses and learning it with a generative AI model, means for converting the user's voice into text using speech recognition technology, emotion recognition means for detecting the user's emotional state in real time, and means for dynamically changing the dialogue content based on the responses. This allows the system to recognize the emotional state of workers in real time and provide appropriate dialogue, thereby providing psychological support and improving work efficiency and maintaining health.

[2057] "User authentication" is the process of receiving user authentication information and verifying that the user is legitimate.

[2058] "Mood confirmation" is a process of asking questions to confirm the user's mood or mental state.

[2059] "Answer analysis" is the process of analyzing a user's answer and evaluating its content.

[2060] "Dialogue content generation" is a process of generating optimal dialogue content for the user based on the analysis results.

[2061] "Speech synthesis technology" is a technology that outputs text information as voice.

[2062] "Data logging" is the process of saving the dialogue and user responses.

[2063] A "generative AI model" is an artificial intelligence model that learns from data and generates appropriate dialogue content.

[2064] "Speech recognition technology" is a technology that converts voice data into text data.

[2065] "Emotion recognition means" is a technology that detects the user's emotional state in real time from their voice and facial expressions.

[2066] The "dynamic dialogue change means" is a process that changes the dialogue content based on the user's response.

[2067] The present invention relates to a mental care system that combines user authentication, emotional state confirmation, emotion recognition, dialogue content generation, voice synthesis, data recording, and learning. Specific embodiments of each element will be described below.

[2068] Overall system configuration

[2069] The system is primarily composed of four elements: a server, a terminal, a user, and an emotion recognition engine. The server is the central location responsible for data processing and operation of the generative AI model, while the terminal is a device that provides an interface with the user. The emotion recognition engine has the function of recognizing the user's emotional state in real time, and the user is the entity that interacts with the system through the terminal.

[2070] server

[2071] The server centrally processes data and runs the generative AI model. The server receives the user's authentication information and checks it against the database to verify that the user is legitimate. If authentication is successful, the server returns a successful authentication message to the user.

[2072] Terminal

[2073] The terminal accepts input from the user and communicates with the server. When the user inputs their feelings or state into the terminal, this information is sent to the server. If login is successful, the terminal displays questions to the user to confirm their current feelings or state.

[2074] Emotion Recognition Engine

[2075] The emotion recognition engine works in conjunction with voice recognition technology to analyze the user's tone and expressions to recognize their emotional state, which is then sent to the server and used to assess the user's mental state.

[2076] User

[2077] The user interacts with the system through the device. When the user inputs their feelings and state, the information is sent to the server and analyzed by an emotion recognition engine. The server then uses a generative AI model to generate optimal dialogue content and send it to the device.

[2078] Detailed process flow

[2079] 1. User authentication:

[2080] The user enters a username and password into the terminal and sends them to the server.

[2081] The server checks the authentication information against a database to verify the user is a valid user.

[2082] If the authentication is successful, the server sends an authentication success message to the terminal, completing the login.

[2083] 2. Emotional validation and emotional recognition:

[2084] If authentication is successful, the terminal presents the user with questions to confirm their feelings and state of mind, and receives their answers.

[2085] The emotion recognition engine analyzes the user's voice and recognizes their emotional state in real time.

[2086] The server uses a generative AI model to analyze the user's responses and emotional information to assess their mental state.

[2087] 3. Generating and providing dialogue content:

[2088] The server uses a generative AI model based on the analysis results to generate optimal dialogue content.

[2089] The generated dialogue content is sent from the server to the terminal and provided to the user using voice synthesis technology.

[2090] 4. Data recording and learning:

[2091] The dialogue and the user's responses are sent to the server and recorded in a database.

[2092] The recorded data is used to train the generative AI model to improve the accuracy of the next interaction.

[2093] Hardware and software used

[2094] Server: Database, generative AI models (e.g., Hugging Face Transformers)

[2095] Devices: Smartphones, PCs, tablets

[2096] Emotion recognition engine: Speech recognition technology (e.g., Google Speech Recognition), emotion analysis (e.g., Sentiment Analysis Pipeline)

[2097] Speech synthesis: Pyttsx3 (e.g., a Python text-to-speech library)

[2098] Examples and prompts

[2099] Example 1:

[2100] Please enter your username: user123

[2101] Please enter your password: password123

[2102] Authentication successful. How are you feeling today?

[2103] (User speaks): "I'm a little tired."

[2104] The system responds: "You sound a little down. Let me know if there's anything I can do to help."

[2105] Example 2:

[2106] Please enter your username: sampleuser

[2107] Please enter your password: mypassword

[2108] Authentication successful. How are you feeling today?

[2109] (User speaks): "I feel great today."

[2110] The system responds: "Great! Keep it up!"

[2111] This allows the system to recognize the emotional state of workers in real time and provide appropriate dialogue to provide psychological support, thereby improving work efficiency and maintaining health.

[2112] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2113] Step 1:

[2114] User Authentication

[2115] Input: The user enters a username and password into the terminal.

[2116] Specific operation: The device sends these authentication information to the server.

[2117] Data processing / calculation: The server checks the authentication information against a database to verify that the user is legitimate.

[2118] Output: If the authentication is successful, the server returns an authentication success message to the terminal; if the authentication fails, it returns an error message.

[2119] Step 2:

[2120] Confirmation of feelings

[2121] Input: Receives information that authentication was successful.

[2122] Specific behavior: The device presents the user with a question to ascertain their current state of mind (e.g., "How are you feeling today?").

[2123] Data processing / calculation: Receive user responses via voice or text input.

[2124] Output: The device sends the user's answer to the server.

[2125] Step 3:

[2126] emotion recognition

[2127] Input: The server receives the user's answer and voice data.

[2128] How it works: The emotion recognition engine analyzes the user's tone of voice and expressions to recognize their emotional state in real time.

[2129] Data processing / calculation: The analysis results are sent to the server, and the emotion data is evaluated using a generative AI model.

[2130] Output: The recognized emotional state and analysis results are output.

[2131] Step 4:

[2132] Dialogue content generation

[2133] Input: Analyzed user emotional state and sentiment data.

[2134] Specific operation: The server uses a generative AI model to generate optimal dialogue content based on the user's emotions and feelings.

[2135] Data processing / calculation: The AI ​​model generates dialogue content based on emotion recognition results and sentiment data.

[2136] Output: The generated dialogue content is sent from the server to the terminal.

[2137] Step 5:

[2138] Dialogue provision

[2139] Input: The generated dialogue.

[2140] Specific operation: The device uses speech synthesis technology to provide the generated dialogue content to the user via voice.

[2141] Data processing / calculation: Converting text data into audio data.

[2142] Output: The dialogue is presented to the user via audio.

[2143] Step 6:

[2144] Data Recording

[2145] Input: Dialogue and user response data.

[2146] Specific operation: The dialogue content sent from the terminal to the server and the user's response data are recorded in a database.

[2147] Data processing / calculation: Recorded data is used as training data for generative AI models.

[2148] Output: Interaction data stored in a database and an updated AI model.

[2149] This allows the system to recognize the user's emotional state in real time and provide appropriate dialogue, thereby providing psychological support, improving work efficiency, and maintaining health.

[2150] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2151] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2152] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2153] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2154] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2155] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2156] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2157] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2158] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2159] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2160] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2161] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2162] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2163] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2164] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2165] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2166] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2167] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2168] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2169] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2170] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2171] The following is further disclosed regarding the above embodiment.

[2172] (Claim 1)

[2173] means for receiving and authenticating a user's authentication information;

[2174] means for asking questions to ascertain the user's state of mind and state and receiving answers thereto;

[2175] A means for analyzing the user's response and generating optimal dialogue content for the user;

[2176] a means for providing the generated dialogue content to a user using a speech synthesis technology;

[2177] A means of recording the dialogue and user responses and training the generative AI model;

[2178] mental health care system, including

[2179] (Claim 2)

[2180] 10. The system of claim 1, wherein a generative AI model is used to analyze user responses and generate appropriate dialogue content.

[2181] (Claim 3)

[2182] 10. The system of claim 1, wherein speech synthesis technology is used to provide the generated dialogue content to the user in voice.

[2183] "Example 1"

[2184] (Claim 1)

[2185] means for receiving and authenticating a user's authentication information;

[2186] means for presenting questions to ascertain the user's mental state and receiving answers thereto;

[2187] A means for analyzing the user's response using a generative AI model and generating appropriate dialogue content for the user;

[2188] a means for providing the generated dialogue content to a user using a speech synthesis technology;

[2189] A means of recording the dialogue and user responses and training the generative AI model;

[2190] A system including:

[2191] (Claim 2)

[2192] The system of claim 1, which uses a generative AI model to analyze user responses and generate appropriate dialogue content.

[2193] (Claim 3)

[2194] 10. The system according to claim 1, wherein the dialogue content generated using speech synthesis technology is provided to the user by voice.

[2195] "Application Example 1"

[2196] (Claim 1)

[2197] means for receiving and authenticating a user's authentication information;

[2198] means for asking questions to ascertain the user's state of mind and state and receiving answers thereto;

[2199] A means for analyzing the user's response and generating optimal dialogue content for the user;

[2200] a means for providing the generated dialogue content to a user using a speech synthesis technology;

[2201] A means of recording the dialogue and user responses and training the generative AI model;

[2202] A means to provide dialogue and content generated by AI based on the user's feelings and state,

[2203] A system including:

[2204] (Claim 2)

[2205] 10. The system of claim 1, wherein a generative AI model is used to analyze user responses and generate appropriate dialogue content.

[2206] (Claim 3)

[2207] 10. The system of claim 1, wherein speech synthesis technology is used to provide the generated dialogue content to the user in voice.

[2208] "Example 2: Combining Emotion Engines"

[2209] (Claim 1)

[2210] means for receiving and authenticating a user's authentication information;

[2211] means for asking questions to ascertain the user's state of mind and state and receiving answers thereto;

[2212] means for analyzing the user's responses and recognizing and assessing the user's emotions using an emotion engine;

[2213] a means for using the generative AI model to generate optimal dialogue content for the user;

[2214] a means for providing the generated dialogue content to a user using a speech synthesis technology;

[2215] A means of recording the dialogue and user responses and training the generative AI model;

[2216] A system including:

[2217] (Claim 2)

[2218] The system of claim 1, wherein a generative AI model is used to integrate and analyze the user's responses and the emotion engine's recognition results to generate appropriate dialogue content.

[2219] (Claim 3)

[2220] 10. The system of claim 1, wherein speech synthesis technology is used to provide the generated dialogue content to the user in voice.

[2221] "Application example 2 when combining emotion engines"

[2222] (Claim 1)

[2223] means for receiving and authenticating a user's authentication information;

[2224] means for asking questions to ascertain the user's state of mind and state and receiving answers thereto;

[2225] A means for analyzing the user's response and generating optimal dialogue content for the user;

[2226] a means for providing the generated dialogue content to a user using a speech synthesis technology;

[2227] A means of recording the dialogue and user responses and training the generative AI model;

[2228] means for converting a user's speech into text using speech recognition technology;

[2229] emotion recognition means for detecting the emotional state of a user in real time;

[2230] means for dynamically changing the dialogue content based on the response;

[2231] mental health care system, including

[2232] (Claim 2)

[2233] 10. The system of claim 1, wherein a generative AI model is used to analyze user responses and generate appropriate dialogue content.

[2234] (Claim 3)

[2235] 10. The system of claim 1, wherein speech synthesis technology is used to provide the generated dialogue content to the user in voice. [Explanation of symbols]

[2236] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving and authenticating a user's authentication information; means for asking questions to ascertain the user's state of mind and state and receiving answers thereto; A means for analyzing the user's response and generating optimal dialogue content for the user; a means for providing the generated dialogue content to a user using a speech synthesis technology; A means of recording the dialogue and user responses and training the generative AI model; mental health care system, including

2. 10. The system of claim 1, wherein a generative AI model is used to analyze user responses and generate appropriate dialogue content.

3. 10. The system of claim 1, wherein speech synthesis technology is used to provide the generated dialogue content to the user by voice.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A