Information Processing System and Information Processing Method

The information processing system addresses the challenge of interacting with users with mental disorders by analyzing their mental state and generating empathetic dialogue, enhancing user interaction.

JP7705198B1Active Publication Date: 2025-07-09IMBESIDEYOU INC

Patent Information

Application Number
JP2025050994
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-09
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

It is difficult to have a smooth conversation with a user having a mental disorder.

Method used

An information processing system that includes an image acquisition unit, a mental state estimation unit, a conversation generation unit, and an output unit to interact effectively with users by analyzing their mental state and generating empathetic dialogue.

Benefits of technology

The system enables effective interaction with users, particularly those with mental illnesses, by generating personalized and empathetic conversations using a large language model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705198000001_ABST
    Figure 0007705198000001_ABST
Patent Text Reader

Abstract

Enable effective interaction with the user. 【Solution means】An information processing system, comprising: an image acquisition unit that acquires an image of the user speaking; a mental state estimation unit that analyzes the image and estimates the mental state of the user; a conversation generation unit that generates a line to gain empathy from the user based on the estimated mental state; and an output unit that outputs the line.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system and an information processing method.

Background Art

[0002] A conversation program using artificial intelligence has been provided (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is difficult to have a smooth conversation with a user having a mental disorder.

[0005] The present invention has been made in view of such a background, and an object thereof is to provide a technology capable of effectively interacting with a user.

Means for Solving the Problems

[0006] The main invention of the present invention for solving the above problems is an information processing system including an image acquisition unit that acquires an image of a user speaking, a mental state estimation unit that analyzes the image and estimates the mental state of the user, a conversation generation unit that generates a line for obtaining empathy from the user based on the estimated mental state, and an output unit that outputs the line.

[0007] Other problems disclosed in the present application and solutions therefor will be clarified by the embodiments of the invention and the drawings.

Effects of the Invention

[0008] According to the present invention, it is possible to effectively interact with a user.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Modes for Carrying Out the Invention

[0010] <Overview of the System> Hereinafter, an information processing system according to an embodiment of the present invention will be described. The information processing system of this embodiment attempts to interact with a user, and particularly assumes that it will conduct counseling for users with mental illnesses. As will be described later, the information processing system generates a script for speaking to the user using a large language model. In the information processing system of this embodiment, the large language model is made to create a script for interacting with the user in the following manner. (1) Icebreaker and Information Collection Start with a conversation that can be answered easily, and collect information about the user. (2) Empathy Generation Conduct a conversation that makes the user feel empathy. (3) Review Judge the validity of the script using a large language model, and regenerate the script as necessary. (4) Reservation for the Next Conversation At the end of the session, summarize the conversation so far, determine the timing for notifying the user of the next conversation using a large language model, and store the personal information that has appeared in the conversation so far.

[0011] Hereinafter, the information processing system of this embodiment will be described.

[0012] FIG. 1 is a diagram showing an example of the overall configuration of an information processing system. The information processing system of the present embodiment includes a management server 2. The management server 2 is communicably connected to the user terminal 1 via a communication network. The communication network is, for example, the Internet and is constructed by a public telephone line network, a mobile phone line network, a wireless communication path, Ethernet (registered trademark), or the like.

[0013] The user terminal 1 is a computer operated by a user. The user terminal 1 can be, for example, a smartphone, a tablet computer, a personal computer, or the like.

[0014] The management server 2 may be a general-purpose computer such as a workstation or a personal computer, or may be logically realized by cloud computing.

[0015] <Management Server> FIG. 2 is a diagram showing an example of the hardware configuration of the management server 2. Note that the illustrated configuration is an example, and other configurations may be used. The management server 2 includes a CPU 201, a memory 202, a storage device 203, a communication interface 204, an input device 205, and an output device 206. The storage device 203 stores various data and programs, such as a hard disk drive, a solid state drive, or a flash memory. The communication interface 204 is an interface for connecting to a communication network, such as an adapter for connecting to Ethernet (registered trademark), a modem for connecting to a public telephone network, a wireless communication device for performing wireless communication, or a USB (Universal Serial Bus) connector or an RS232C connector for serial communication. The input device 205 inputs data, such as a keyboard, a mouse, a touch panel, a button, or a microphone. The output device 206 outputs data, such as a display, a printer, or a speaker. Each functional unit of the management server 2 described below is realized by the CPU 201 reading a program stored in the storage device 203 into the memory 202 and executing it, and each storage unit of the management server 2 is realized as a part of the storage area provided by the memory 202 and the storage device 203.

[0016] FIG. 3 is a diagram showing an example of the software configuration of the management server 2. The management server 2 includes a personal information storage unit 231, an image acquisition unit 211, a mental state estimation unit 212, a conversation generation unit 213, an output unit 214, a conversation acquisition unit 215, a determination unit 216, a notification setting unit 217, a summary notification unit 218, and a personal information extraction unit 219.

[0017] <Storage unit> The personal information storage unit 231 stores the user's personal information. The personal information storage unit 231 stores the personal information in association with information for identifying the user (e.g., user ID). The personal information includes the user's name, gender, address, etc. The personal information broadly includes the user's attributes, such as the user's personality and emotional expression.

[0018] The personal information storage unit 231 receives the personal information extracted by the personal information extraction unit 219 described later and stores it as structured data. The personal information storage unit 231 can be implemented in the form of, for example, a relational database, a NoSQL database, a key-value store, etc. The personal information stored in the personal information storage unit 231 is associated with the past conversation history with the user and is used as reference information when the conversation generation unit 213 generates a more personalized line. Also, the personal information storage unit 231 can hold the update history of personal information and is capable of tracking changes in the user's personal information over time.

[0019] <Functional unit> The image acquisition unit 211 acquires an image (a video in this embodiment) of the user speaking. The image acquisition unit 211 can, for example, receive the video captured by the camera provided in the user terminal 1 from the user terminal 1 by streaming. The image acquisition unit 211 can acquire, together with the image, the voice of the user's speech that has been collected. The image acquisition unit 211 can, for example, receive the voice collected by the microphone provided in the user terminal 1 from the user terminal 1 by streaming. The image acquisition unit 211 can receive a moving image including both video and audio from the user terminal 1.

[0020] The mental state estimation unit 212 estimates the mental state of the user. In this embodiment, the "mental state" is a concept that broadly includes the user's psychological and mental states such as not only the emotional state but also the cognitive state, thinking pattern, state of attention, level of motivation, level of stress, degree of fatigue, level of arousal, and the manifestation state of symptoms related to mental disorders. The mental state can include both a temporary state (state aspect) and a relatively persistent characteristic (characteristic aspect).

[0021] The mental state estimation unit 212 can analyze the moving image to detect changes in the user's biological reaction.

[0022] The mental state estimation unit 212 can, for example, separate a moving image into a set of images (a collection of frame images) and audio, and analyze changes in the biological reaction from each. For example, the mental state estimation unit 212 can analyze changes in the biological reaction related to at least one of the user's facial expression, eye line, pulse, and facial movement by analyzing the user's face image using the frame images separated from the moving image. Also, the mental state estimation unit 212 can analyze changes in the biological reaction related to at least one of the user's speech content and voice quality by analyzing the audio separated from the moving image.

[0023] When a person's emotion changes, it appears as a change in the biological reaction such as facial expression, eye line, pulse, facial movement, speech content, and voice quality. In the present embodiment, the change in the user's emotion is analyzed by analyzing the change in the user's biological reaction. The emotion analyzed in the present embodiment is, as an example, the degree of pleasure / displeasure. In the present embodiment, the mental state estimation unit 212 can calculate a biological reaction index value reflecting the content of the change in the biological reaction by quantifying the change in the biological reaction according to a predetermined standard.

[0024] The analysis of the change in facial expression is performed, for example, as follows. That is, for each frame image, the face region is specified from the frame image, and the facial expressions specified are classified into a plurality according to an image analysis model that has been machine-learned in advance. Then, based on the classification result, it is analyzed whether a positive facial expression change has occurred, a negative facial expression change has occurred, and how large the facial expression change is between consecutive frame images, and a facial expression change index value corresponding to the analysis result can be calculated.

[0025] Analysis of the change in the line of sight is performed, for example, as follows. That is, for each frame image, the eye region is specified from within the frame image, and by analyzing the directions of both eyes, it is analyzed where the user is looking. For example, it is analyzed whether the user is looking at the face of the speaker being displayed, the shared material being displayed, outside the screen, etc. Also, it may be analyzed whether the movement of the line of sight is large or small, whether the frequency of movement is high or low, etc. The change in the line of sight is also related to the user's concentration. The mental state estimation unit 212 can calculate a line-of-sight change index value according to the analysis result of the change in the line of sight.

[0026] Analysis of the change in the pulse is performed, for example, as follows. That is, for each frame image, the face region is specified from within the frame image. Then, using a learned image analysis model that captures the numerical value of the face color information (G of RGB), the change in the G color on the face surface is analyzed. By arranging the results along the time axis, a waveform representing the change in the color information is formed, and the pulse is specified from this waveform. When a person is tense, the pulse becomes faster, and when the person is calm, the pulse becomes slower. The biological reaction analysis unit 213 can calculate a pulse change index value according to the analysis result of the change in the pulse.

[0027] Analysis of the change in the movement of the face is performed, for example, as follows. That is, for each frame image, the face region is specified from within the frame image, and by analyzing the direction of the face, it is analyzed where the user is looking. For example, it is analyzed whether the user is looking at the face of the speaker being displayed, the shared material being displayed, outside the screen, etc. Also, it may be analyzed whether the movement of the face is large or small, whether the frequency of movement is high or low, etc. The movement of the face and the movement of the line of sight may be analyzed together. For example, it may be analyzed whether the user is looking straight at the face of the speaker being displayed, looking with an upward or downward glance, looking obliquely, etc. The mental state estimation unit 212 can calculate a face orientation change index value according to the analysis result of the change in the direction of the face.

[0028] The analysis of the speech content is performed as follows, for example. That is, the biological reaction analysis unit 213 converts the speech into a character string by performing known speech recognition processing on the speech for a specified time (for example, a time of about 30 to 150 seconds), and morphological analyzes the character string to remove words that are unnecessary for representing conversations such as particles and articles. Then, the remaining words are vectorized, and it is analyzed whether a positive emotional change has occurred, whether a negative emotional change has occurred, and the magnitude of the emotional change that has occurred, and a speech content index value corresponding to the analysis result can be calculated.

[0029] The analysis of the voice quality is performed as follows, for example. That is, the biological reaction analysis unit 12 identifies the acoustic characteristics of the speech by performing known speech analysis processing on the speech for a specified time (for example, a time of about 30 to 150 seconds). Then, based on the acoustic characteristics, it is analyzed whether a positive voice quality change has occurred, whether a negative voice quality change has occurred, and the magnitude of the voice quality change that has occurred, and a voice quality change index value corresponding to the analysis result can be calculated.

[0030] The mental state estimation unit 212 calculates a biological reaction index value using at least one of the facial expression change index value, eye line change index value, pulse change index value, face orientation change index value, speech content index value, and voice quality change index value calculated as described above. For example, the biological reaction index value can be calculated by performing weighted calculation on the facial expression change index value, eye line change index value, pulse change index value, face orientation change index value, speech content index value, and voice quality change index value.

[0031] The mental state estimation unit 212 can estimate a mental disorder according to the above response from the patient and the change in the biological reaction (biological reaction index value). The estimation unit 214 can estimate the mental disorder suffered by the patient by giving the received response and the detected change in the biological reaction to a disease model stored in advance.

[0032] The mental state estimation unit 212 can determine the probability that a patient has multiple mental disorders based on the reliability of the inference by the disease model. Specifically, it uses the probability of belonging to each class of mental disorder calculated in the output layer of the disease model. Usually, the disease model receives input data (the patient's answers and biometric reaction index values) and probabilistically outputs to which class of mental disorder the data belongs. For example, the probability of depression is 70%, the probability of bipolar disorder is 20%, and the probability of schizophrenia is 10%, and the probabilities of suffering from multiple mental disorders are calculated. Note that the reliability may be used as "probability". That is, the non-linear probability of suffering can be called probability.

[0033] The mental state estimation unit 212 can evaluate the certainty of the diagnosis based on the probability of the estimated mental disorder. The reliability of the estimation result can be evaluated by the absolute value of the probability of the mental disorder with the highest probability. For example, when the probability of depression is 90% or more, it can be determined that the possibility of depression is very high. On the other hand, when the probability of the mental disorder with the highest probability is about 50%, it can be determined that the certainty of the diagnosis is not very high.

[0034] The mental state estimation unit 212 can suggest the possibility of comorbidities based on the probability of the estimated mental disorder. When the probabilities of suffering from two or more mental disorders are both high, it can be suggested that those disorders may coexist.

[0035] In the above manner, the mental state estimation unit 212 can estimate the first mental state analyzed from the image and the second mental state analyzed from the voice.

[0036] The conversation generation unit 213 generates lines for the system to the user. The conversation generation unit 213 can generate lines to gain the user's empathy based on the mental state estimated by the mental state estimation unit 212. The conversation generation unit 213 can generate lines by providing a large language model with a prompt including information representing the estimated mental state and an instruction to generate lines based on the information. Also, the conversation generation unit 213 may generate lines by providing a large language model with a prompt including information representing the first and second mental states and an instruction to generate lines based on this information.

[0037] When the conversation generation unit 213 generates lines to gain the user's empathy, the following specific language patterns and psychological approaches can be used.

[0038] First, the conversation generation unit 213 can generate lines using the technique of reflective listening. Reflective listening is a technique that paraphrases and returns the user's statement content and emotions so that the user feels understood. For example, the following language patterns can be used.

[0039] (1) Simple reflection: Repeat the user's statement almost as it is. For example, when the user says "I'm very tired today," respond with "You're very tired today, aren't you."

[0040] (2) Emotion reflection: Read the emotion from the user's statement and verbalize it. For example, when the user says "There are too many requests from my boss and I can't handle them," respond with "There are so many requests from your boss that you feel overwhelmed, don't you."

[0041] (3) Summary reflection: Summarize and return the user's multiple statements. For example, after a long conversation, respond with "So, it means that the work burden and family problems overlap and you're running out of mental capacity, right."

[0042] The conversation generation unit 213 can generate lines using verification techniques. Verification means justifying the user's feelings and experiences and recognizing them as understandable and natural. For example, the following language patterns can be used.

[0043] (1) Justification of feelings: Use expressions such as "It's natural to feel anxious in such a situation" and "I can well understand that you feel that way".

[0044] (2) Generalization: Use expressions such as "Many people feel the same way in such a situation" and "That's not an unusual reaction".

[0045] (3) Expression of understanding: Use expressions such as "If I were in your position, I might feel the same way" and "I can well understand the difficulty of that situation".

[0046] The conversation generation unit 213 can generate lines using self-disclosure techniques. Self-disclosure is a technique for reducing the psychological distance from the user by appropriately sharing the system's own "experiences" and "feelings". However, the system's self-disclosure is fictional and is used only as a means to gain the user's empathy. For example, the following language patterns can be used.

[0047] (1) Sharing of similar experiences: Use expressions such as "I have also experienced a similar situation before" and "I have heard such stories from many people".

[0048] (2) Sharing of feelings: Use expressions such as "Hearing your story, I also feel a little sad" and "Hearing about your success, I am also happy".

[0049] The dialogue generation unit 213 can generate lines including non-verbal elements. The non-verbal elements are expressed as instructions such as "(nods)" and "(pauses for a while)" in text, and are realized as actual actions during speech synthesis and visual expression. For example, the following non-verbal elements can be specified.

[0050] (1) Use expressions such as "I see (while nodding)" and "Let me think for a while (pausing for a while), yes...".

[0051] (2) Specify the tone and speed of the voice: Use expressions such as "It's okay (in a gentle voice)" and "Let's think one by one slowly".

[0052] (3) Specify emotional expressions: Use expressions such as "That must have been tough (with a sympathetic expression)" and "What a great result (while smiling)".

[0053] The dialogue generation unit 213 can generate lines by appropriately combining the above techniques according to the user's mental state and the context of the dialogue. For example, when the user is expressing sadness or a sense of loss, lines mainly using the techniques of reflective listening and verification can be generated. Also, when the user is expressing anxiety or fear, lines combining the technique of verification and non-verbal elements can be generated.

[0054] Also, the dialogue generation unit 213 can use different techniques according to the progress of the conversation with the user.

[0055] The "lines" generated by the dialogue generation unit 213 do not have to be limited to mere text information. In addition to the text information indicating the utterance content, the lines can include information on non-verbal elements and emotional expressions as follows. (1) Voice pattern information: Information such as the pitch, speed, rhythm, intonation, and the way of taking pauses of the voice. For example, when the user is estimated to be depressed, it includes information specifying to generate lines in a slow and gentle tone. (2) Emotional expression information: Information that specifies the emotions (such as empathy, encouragement, sense of security, etc.) embedded in the lines. This information can be used to adjust the voice characteristics reflecting emotions in the subsequent speech synthesis process. (3) Non-verbal expression information: Information that specifies physical expressions such as facial expressions, gestures, postures, etc. when using physical output devices such as avatars or robots. For example, it includes instructions such as "while nodding" and "while smiling". (4) Timing information: Information that specifies the timing of uttering the lines or the intervals between multiple lines. It can specify whether to respond immediately to the user's speech or to respond after a short pause.

[0056] The conversation generation unit 213 can appropriately adjust the information regarding these non-verbal elements and emotional expressions according to the mental state of the user. For example, when the user is presumed to be in an anxious state, it is possible to generate lines specified to speak slowly with a calm and low tone of voice while taking an appropriate pause. Also, when the user is presumed to be feeling sad, it is possible to generate lines specified to give a sympathetic response with a warm tone of voice at an appropriate timing.

[0057] In addition, the conversation generation unit 213 generates questions for estimating the second mental state so that the content can be answered more easily as it gets closer to the start time of the session with the user. The second mental state can be estimated by analyzing the answers to these questions.

[0058] The large language model (LLM) used in this embodiment is a natural language processing model pre-trained with a large amount of text data. The large language model is based on, for example, a Transformer-type neural network architecture and can understand the context using a self-attention mechanism and perform text generation.

[0059] Examples of large language models that can be used in this embodiment include, for example, GPT (Generative Pre-trained Transformer) series models, LLaMA (Large Language Model Meta AI), PaLM (Pathways Language Model), Claude, Bard, Gemini, etc. In this embodiment, these existing models may be used as they are, or models that have been fine-tuned to specialize in specific tasks (for example, understanding mental states and generating empathetic responses) may be used.

[0060] The input to the large language model is a text-based instruction called a prompt. In this embodiment, the prompt may include information representing the user's mental state, the conversation history, information regarding the role and purpose of the system, and specific instructions (for example, "Please generate an empathetic line based on the following mental state").

[0061] Examples of methods for implementing a large language model include the following. (1) Using cloud-based APIs: A method of accessing a large language model through an API provided by an external cloud service. In this method, there is no need to prepare high-performance hardware in-house, and scalability can also be ensured. (2) On-premises implementation: A method of implementing a large language model on the company's own server. In this method, high levels of data privacy and security can be ensured. (3) Implementation on edge devices: A method of operating a lightweight large language model on edge devices (such as user terminals). In this method, network latency can be minimized, and it can also be used in an offline environment.

[0062] In this embodiment, any of the above implementation methods can be adopted, and an appropriate method can be selected according to the requirements of the system and the operating environment.

[0063] In this embodiment, in order to control the output of the large language model, parameters such as temperature, Top-P, and Top-K can be adjusted. Temperature is a parameter that controls the diversity of the output. With a low value (e.g., 0.2), a deterministic response is generated, and with a high value (e.g., 0.8), a more creative response is generated. In this embodiment, a medium temperature setting (e.g., 0.5 - 0.7) is used to give appropriate creativity when generating dialogue lines, and a low temperature setting (e.g., 0.1 - 0.3) is preferably used for tasks that require judgments such as validity determination.

[0064] The output unit 214 outputs dialogue lines. The output unit 214 can output the generated question to the user. For example, the output unit 214 can provide the dialogue lines to a speech synthesis engine to generate voice data of the dialogue lines and transmit the generated voice data to the user terminal 1.

[0065] The output unit 214 can also output dialogue lines using an avatar or a virtual character. By outputting dialogue lines through a virtual character that changes expressions and gestures according to the user's mental state, more human-like interactions can be realized. For example, when it is determined that the mental state of the user estimated by the mental state estimation unit 212 is depressed, the virtual character can be made to speak the dialogue lines while showing a kind expression and empathetic gestures. Conversely, when it is determined that the mental state of the user is excited, the virtual character can be made to speak the dialogue lines while showing a bright expression and lively gestures.

[0066] The output unit 214 can also customize the appearance and personality of the character according to the user's preferences and treatment purposes. For example, the user can be enabled to select a character with an age group, gender, and appearance characteristics that the user is likely to feel familiar with, or according to the user's treatment purpose, select from multiple character types such as a character with more empathetic personality characteristics or a character with more guiding personality characteristics. The setting information of the appearance and personality of the character can be stored in the personal information storage unit 231, and the optimal character setting for each user can be called and used.

[0067] The output unit 214 can also control the motion pattern of the virtual character according to the content and intention of the dialogue. For example, when asking a question, a gesture of tilting the head can be added, or when emphasizing an important point, a gesture of moving the hand can be added. Also, according to the emotional nuance of the dialogue, the tone of voice, the speaking speed, the change of expression, etc. can be adjusted. By appropriately combining these non-verbal elements, subtle nuances and emotions that are difficult to convey only by text or voice can be effectively transmitted, and the construction of a trust relationship with the user can be promoted.

[0068] The conversation acquisition unit 215 acquires the conversation content spoken by the user. The conversation acquisition unit 215 may acquire the conversation content analyzed from the voice by the mental state estimation unit 212, or may analyze the voice instead of the mental state estimation unit 212 to acquire the conversation content. The conversation acquisition unit 215 can acquire the conversation content including the answer to the generated question from the voice after the conversation generation unit 213 generates the question.

[0069] The determination unit 216 determines the validity of the dialogue line. The determination unit 216 can determine the validity of the dialogue line based on the second mental state and the conversation content. The determination unit 216 can give a prompt including information indicating the second mental state, the conversation content, the dialogue line, and an instruction to determine the validity of the dialogue line based on the second mental state and the conversation content to a large language model to generate the validity. Note that instead of or in addition to the second mental state, the first mental state may be used. In this case, the information indicating the first and second mental states may be included in the prompt, and the instruction may be such that it determines the validity of the dialogue line based on the first and second mental states and the conversation content.

[0070] As criteria for the determination unit 216 to determine the validity of the dialogue line, the following multiple viewpoints can be included. (1) Psychological safety: Is there any possibility that the dialogue line will cause psychological harm to the user? For example, evaluate whether expressions that strengthen suicidal thoughts, overly negative expressions, expressions that may worsen the user's state, etc. are included. (2) Appropriateness of empathy: Does the dialogue line show appropriate empathy for the user's mental state? For example, evaluate whether an inappropriate bright reaction is made when the user expresses sadness, or conversely, whether an overly serious reaction is made to a minor concern. (3) Consistency of context: Is the dialogue line consistent with the flow and context of the conversation? Evaluate whether it is not contradictory to the past conversation content and whether there is no abrupt topic change. (4) Degree of personalization: Does the dialogue line appropriately reflect the user's personal information and past conversation content? Evaluate whether it is not a too general response but personalized content tailored to the user's specific situation. (5) Therapeutic value: Is the dialogue line valuable from a therapeutic perspective? Evaluate whether it includes therapeutic elements such as not only empathy but also cognitive restructuring, promotion of problem-solving, and improvement of self-efficacy. (6) Cultural appropriateness: Does the dialogue line take into account the user's cultural background and values? Evaluate whether it includes culturally inappropriate expressions or assumptions.

[0071] For example, the determination unit 216 may calculate a numerical score from 0 to 10 for each determination criterion, and calculate a comprehensive score after weighting each criterion. For example, a high weight (e.g., 0.3) can be assigned to psychological safety, a medium weight (e.g., 0.2) to the appropriateness of empathy, and weights of about 0.1 to 0.15 to the other criteria respectively. If the comprehensive score is below a predetermined threshold (e.g., 7.0), the line is determined to be inappropriate.

[0072] When determining the validity of a line using a large language model, the determination unit 216 can also use a structured prompt as follows. "As an expert in mental healthcare, you will evaluate the validity of the lines generated by the counseling system. Please evaluate based on the following information. User's mental state: [Detailed information on mental state] Conversation history: [Most recent conversation content] Line to be evaluated: [Generated line] Please evaluate on a scale from 0 to 10 based on the following criteria: 1. Psychological safety (weight: 0.3): Does this line have the potential to cause psychological harm to the user? 2. Appropriateness of empathy (weight: 0.2): Does this line show appropriate empathy for the user's mental state? 3. Context consistency (weight: 0.15): Is this line consistent with the flow and context of the conversation? 4. Degree of personalization (weight: 0.15): Does this line appropriately reflect the user's personal information and past conversations? 5. Therapeutic value (weight: 0.1): Does this line have value from a therapeutic perspective? 6. Cultural appropriateness (weight: 0.1): Does this line take into account the user's cultural background and values? Briefly explain the reasons for the evaluation for each criterion, and finally calculate the weighted comprehensive score. Also, if improvements are needed, please propose specific points for improvement."

[0073] Note that the determination unit 216 may be instructed to determine validity based on only one determination criterion. For example, an instruction can be input to the prompt to judge the appropriateness of empathy.

[0074] The dialogue generation unit 213 can recreate the lines according to the validity.

[0075] As a method for the dialogue generation unit 213 to recreate the lines, the following multiple approaches can be adopted.

[0076] (1) Problem-specific regeneration: A method of performing regeneration focusing on the specific problems identified by the determination unit 216. For example, when the score of "appropriateness of empathy" is low, the dialogue generation unit 213 gives a prompt with the following correction instructions added to the original lines to the large language model. "The following lines lack empathy for the user's mental state. Please correct them to show deeper empathy according to the user's [specific emotions or situations]: [original lines]"

[0077] (2) Step-by-step improvement regeneration: A method of dealing with multiple problems in order from the problems with higher priority when there are multiple problems. For example, first generate lines with the problem of "psychological safety" corrected, and then improve the "appropriateness of empathy" of those lines step by step.

[0078] (3) Complete regeneration type: A method of generating anew from scratch while referring to the original lines. Create a new prompt that reflects the evaluation results by the determination unit 216 in detail and give it to the large language model. For example: "Please generate new lines based on the following conversation history and the user's mental state. There were [specific problems] in the previously generated lines. Please pay special attention to [points to be improved] in the new lines."

[0079] (4) Multiple candidate generation type: A method of generating multiple candidate scripts, each evaluated by the determination unit 216, and adopting the one with the highest score. In this method, the conversation generation unit 213 generates scripts multiple times (for example, 3 to 5 times) using the same prompt, and the determination unit 216 evaluates each candidate.

[0080] In the recreation of the script, the conversation generation unit 213 can apply the following specific improvement strategies.

[0081] (1) Improvement by paraphrasing: Rewrite problematic expressions or phrases into more appropriate ones. For example, change an assertive expression like "You are wrong" to a softer expression like "There may be another way of looking at it".

[0082] (2) Structure modification: Improve the way the message is conveyed by changing the structure of the script. For example, instead of conveying negative content first and then positive content, change to a structure that starts with positive content and carefully conveys negative content.

[0083] (3) Specificity adjustment: Make overly abstract expressions more specific, or conversely, generalize overly specific and restrictive expressions. Select an appropriate level of abstraction according to the user's situation.

[0084] (4) Emotional tone adjustment: Adjust the overall emotional tone of the script. For example, if it is overly bright, adjust it to a calm tone, and if it is overly dark, adjust it to a tone containing hope.

[0085] (5) Addition of personalized elements: Appropriately refer to the user's personal information and past conversation content and incorporate them into the script. For example, refer to past conversations like "Regarding 〇〇 talked about last time, how has it been since then?"

[0086] The conversation generation unit 213 can re-evaluate the recreated lines by the determination unit 216 and confirm that the validity has been improved. As a result of the re-evaluation, if the validity has not been sufficiently improved, it is possible to try a different recreation approach or consider more fundamental problems (for example, the possibility that the estimation of the user's mental state is inaccurate). Also, if the validity does not improve even after attempting multiple recreations, it can be equipped with a function to fallback to a safer and more general response (for example, "Could you please tell me a little more in detail?").

[0087] The notification setting unit 217 sets a plan to notify the user to have a conversation next time based on the conversation content when the session with the user ends. The notification setting unit 217 can set the date and time of the notification in, for example, a reminder application or an alarm application. The notification setting unit 217 can determine the scheduled (date and time) of the notification by giving a prompt including the conversation content and an instruction to consider the date and time when the user should be addressed next based on the conversation content to the large language model.

[0088] The notification setting unit 217 can use a decision algorithm that considers multiple elements to determine the timing of the notification. For example, the notification timing can be calculated based on the following elements.

[0089] (1) Severity of the user's mental state: Adjust the notification frequency according to the severity of the mental state estimated by the mental state estimation unit 212. For example, if it is estimated to be a severe depressive state, set a notification within a short period of 1 to 2 days, and if it is estimated to be a mild anxiety state, set a notification 3 to 7 days later.

[0090] (2) Temporal elements extracted from the conversation content: Consider the schedules and events mentioned during the conversation ("I have an interview tomorrow", "I'm going back to my hometown on the weekend", etc.) and set the notification timing according to the user's life rhythm.

[0091] (3) Past response patterns: Analyze the historical data on how the user has responded to past notifications, and preferentially select the days of the week and time periods with high response rates.

[0092] (4) Optimal intervals based on treatment protocols: Refer to the recommended follow-up intervals based on the standard treatment protocols for specific mental disorders. For example, in cognitive behavioral therapy, a once-a-week session is generally considered standard, so set an interval according to this as the reference value.

[0093] The notification setting unit 217 also has a function to dynamically adjust the notification frequency according to the user's state. This adjustment is carried out based on the following rules:

[0094] (1) Increased frequency during state deterioration: If the user's mental state shows a tendency to deteriorate in consecutive sessions, automatically increase the notification frequency. For example, increase the usually once-a-week notification to two or three times a week.

[0095] (2) Frequency optimization during state improvement: When the user's state shows a tendency to improve, instead of suddenly reducing the notification frequency, gradually widen the interval step by step. For example, gradually adjust from twice a week to once a week, and then to once every two weeks.

[0096] (3) Concentrated support before and after important events: If an important life event (such as an exam, a job interview, a move, etc.) is detected from the user's conversation content, set to send concentrated notifications before and after that event.

[0097] (4) Alternative strategies in case of non-response: If the user does not respond to notifications a certain number of times (for example, two consecutive times), execute alternative strategies such as changing the notification method (text, voice, visual alert, etc.) or the time period.

[0098] The notification setting unit 217 also personalizes the notification contents according to the user's state. For example, if a specific task or goal was set in the previous session, the notification includes a message asking about the progress. Also, by referring to the information stored in the personal information storage unit 231 and incorporating topics related to the user's interests and concerns into the notification contents, the notification setting unit 217 makes efforts to increase the user's willingness to respond.

[0099] The notification setting unit 217 records the results of the notification settings as structured data, and learns the optimal notification pattern for each user over time. This learning data is analyzed using a machine learning algorithm (e.g., random forest or gradient boosting decision tree) to continuously improve the prediction accuracy of the notification timing. The learning model is constructed with the user's reaction to the notification (time to respond, whether or not to participate in a session, activeness during a session, etc.) as the objective variable, and the timing, content, and state of the notification as the explanatory variables.

[0100] The summary notification unit 218 generates a summary of the conversation content at the end of (before) a session with a user, and notifies the user of the summary. Whether or not the session is about to end can be determined, for example, from lines or the conversation content. The summary notification unit 218 can generate the summary and pass it to the output unit 214 so that the output unit 214 outputs it.

[0101] The personal information extraction unit 219 can extract personal information of the user from the conversation content at the end of (before) the session with the user. The personal information extraction unit 219 can register the extracted personal information in the personal information storage unit 231.

[0102] The personal information extraction unit 219 can analyze the conversation content using natural language processing technology and identify the user's personal information. For example, the personal information extraction unit 219 can use technologies such as name recognition, entity extraction, and relationship extraction to extract information such as the user's name, age, occupation, family composition, hobbies, preferences, living habits, past experiences, and health status from the conversation content. The personal information extraction unit 219 may evaluate the confidence level of the extracted personal information and register only the information whose confidence level exceeds a predetermined threshold in the personal information storage unit 231.

[0103] Through the cooperation between the personal information extraction unit 219 and the personal information storage unit 231, the system can construct a more detailed and accurate user profile each time it has a conversation with the user. The personal information extraction unit 219 can confirm the consistency between the newly extracted personal information and the information already stored in the personal information storage unit 231, and if there are contradictions, it can give priority to updating with newer information or information with a higher confidence level. In addition, the personal information extraction unit 219 can detect the temporal changes in personal information (such as changes like "previously it was 〇〇, but now it is △△") from the context of the conversation and appropriately update the information in the personal information storage unit 231.

[0104] The information stored in the personal information storage unit 231 can be referred to by the conversation generation unit 213 in subsequent sessions. The information stored in the personal information storage unit 231 can be utilized for generating personalized lines that take into account the user's personal background and past conversation content. For example, it is possible to realize a more natural and continuous conversation by incorporating topics related to the user's hobbies and interests or by asking about the progress of issues mentioned in the past. Also, the mental state estimation unit 212 can more accurately estimate the change in the mental state compared to the user's normal state (baseline) by referring to the information stored in the personal information storage unit 231.

[0105] <Operation> FIG. 4 is a diagram for explaining the operation of the management server 2.

[0106] The management server 2 acquires the user's video (S301), estimates the mental state from the video and audio respectively (S302), generates a line according to the mental state (S303), and outputs the line to the user (S304).

[0107] As described above, according to the information processing system of the present embodiment, it is possible to automatically generate a line and advance the conversation while considering the mental state of the user.

[0108] Although the present embodiment has been described above, the above embodiment is for facilitating the understanding of the present invention and is not for limiting and interpreting the present invention. The present invention can be changed and improved without departing from its gist, and the equivalents of the present invention are also included in the present invention.

[0109] For example, the processing by each functional unit of the management server 2 described above may be executed by any functional unit. Also, different functional units that execute a part of the processing of each functional unit described above may be added. Further, the functional units of the management server 2 may be provided in a distributed manner by a plurality of computers.

[0110] Also, the information stored in each storage unit of the management server 2 may be stored by any storage unit. That is, the information stored in the plurality of storage units described above may be stored by one storage unit, or a part of the information stored in a certain storage unit described above may be stored by another storage unit.

[0111] <Modification Example 1> In the above-described embodiment, the configuration in which the information processing system interacts with the user as a whole has been described. In this modification example, a multi-agent system in which a plurality of agents operate in cooperation will be described.

[0112] In Modification Example 1, for example, the management server 2 can include an image acquisition agent, a mental state estimation agent, a dialogue generation agent, and an integrated management agent. Each agent may be implemented as an independent computer system or as an independent software module operating on the same computer system.

[0113] The image acquisition agent is an agent responsible for the functions of the above-described image acquisition unit 211, and acquires an image of the user speaking. The image acquisition agent can manage a plurality of camera devices and has a function of acquiring an image at an optimal angle and resolution. Further, the image acquisition agent can communicate with the user terminal 1 and has a function of acquiring an image from the user terminal 1. The image acquisition agent can perform pre-processing (such as noise removal, resolution adjustment, frame rate adjustment, etc.) on the acquired image data and then transmit it to the mental state estimation agent.

[0114] The mental state estimation agent is an agent responsible for the functions of the above-described mental state estimation unit 212, and can estimate the user's mental state by analyzing the image received from the image acquisition agent. The mental state estimation agent can have a plurality of specialized sub-agents inside. For example, an expression analysis sub-agent, a voice quality analysis sub-agent, a body movement analysis sub-agent, etc. Each sub-agent performs analysis specialized in its respective specialized field and transmits the result to the integrated sub-agent within the mental state estimation agent. The integrated sub-agent integrates the analysis results from each sub-agent, generates a final mental state estimation result, and transmits it to the dialogue generation agent and the integrated management agent.

[0115] The script generation agent is an agent responsible for the functions of the above-mentioned conversation generation unit 213, and can generate a script for obtaining empathy from the user based on the estimated mental state result received from the mental state estimation agent. The script generation agent can also have a plurality of specialized sub-agents internally. For example, an empathy generation sub-agent, a problem-solving sub-agent, a medical advice sub-agent, etc. Each sub-agent generates a script specialized in its respective specialized field and sends the result to the selection sub-agent within the script generation agent. The selection sub-agent selects the script generated by the most appropriate sub-agent according to the user's mental state and the context of the conversation, or combines the scripts generated by multiple sub-agents to generate a final script and sends it to the integrated management agent.

[0116] The integrated management agent is an agent that manages the cooperation between each agent and controls the operation of the entire system. The integrated management agent manages the entire session with the user and can adjust the operation timing of each agent and manage the data transfer between agents. In addition, the integrated management agent can also be responsible for the functions of the above-mentioned output unit 214, conversation acquisition unit 215, determination unit 216, notification setting unit 217, summary notification unit 218, and personal information extraction unit 219.

[0117] The integrated management agent can determine the validity of the script received from the script generation agent and request the script generation agent to regenerate the script if necessary. In addition, at the end of the session, the integrated management agent can set a plan to notify the user to have a conversation next time based on the conversation content, generate a summary of the conversation content and notify the user, or extract the user's personal information from the conversation content and register it in the personal information storage unit.

[0118] According to the multi-agent system of Modification Example 1, by having each agent perform processing specialized in a specific domain, an improvement in the overall performance of the system can be expected. Also, since each agent can be developed and improved independently, the scalability and maintainability of the system are enhanced. Furthermore, it becomes easier to replace only a specific agent as needed.

[0119] Although a configuration in which the mental state estimation agent and the dialogue generation agent have a plurality of sub-agents inside has been described, it is also possible to implement these sub-agents as independent agents. For example, an expression analysis agent, a voice quality analysis agent, a body movement analysis agent, etc. may be implemented as independent agents, and an agent for integrating the analysis results of these agents may be provided separately.

[0120] Furthermore, although the integrated management agent has a configuration that undertakes a plurality of functions, it is also possible to provide a plurality of agents that undertake these functions. For example, an output agent, a dialogue acquisition agent, a determination agent, a notification setting agent, a summary notification agent, a personal information extraction agent, etc. may be implemented as independent agents, and a master agent for adjusting the operations of these agents may be provided separately.

[0121] <Modification Example 2> In Modification Example 2, a multi-agent system is configured by implementing a plurality of agents having different roles and specializations, and having an appropriate agent respond according to the state of the user and the progress of the conversation.

[0122] The management server 2 of Modification Example 2 can further include an agent selection unit, an agent profile storage unit, and an agent cooperation unit.

[0123] The agent profile storage unit stores the profile information of a plurality of agents. The profile information of each agent includes the agent's role, expertise, personality traits, conversation style, conditions of mental state to be applied, physical characteristics, etc. In this modified example, for example, agents with different roles as follows can be prepared.

[0124] (1) Empathetic agent: An agent specialized in empathizing with the user's emotions and providing emotional support. It has a warm way of speaking and a conversation style that affirms and accepts the user's emotions.

[0125] (2) Problem-solving agent: An agent specialized in proposing specific solutions and countermeasures for the problems and issues the user is facing. It has a logical and structured conversation style.

[0126] (3) Medical advice agent: An agent that provides specialized advice and information based on psychiatric knowledge. It provides explanations based on scientific evidence and information on treatment methods.

[0127] (4) Motivational agent: An agent specialized in assisting the user in changing behavior and achieving goals. It has a conversation style that frequently uses encouragement and positive feedback to enhance the user's sense of self-efficacy.

[0128] The agent selection unit selects the most suitable agent for the current situation from among a plurality of agents based on the mental state of the user estimated by the mental state estimation unit 212 and the conversation content acquired by the conversation acquisition unit 215. The agent selection unit can use, for example, the following selection logic.

[0129] (1) Selection based on mental state: When the user shows strong anxiety or sadness, select an empathetic agent; when specifically talking about a problem, select a problem-solving agent; when asking questions about symptoms or treatment, select a medical advice agent.

[0130] (2) Selection based on the conversation phase: Select an appropriate agent according to the progress of the conversation, such as selecting a empathy agent in the initial stage of the conversation, a problem-solving agent when the details of the problem become clear, and a motivation agent at the end of the conversation.

[0131] (3) Selection based on the user's reaction: Analyze the user's reaction to a specific agent (such as changes in mental state and conversation activity), and preferentially select the agent that elicited a positive reaction.

[0132] (4) Selection based on the passage of time: If the same agent has been handling the conversation for a long time, switch to a different agent to give the conversation a new perspective and vitality.

[0133] The agent selection unit can use a composite judgment criterion that combines the above selection logics. In addition, the agent selection unit can use a machine learning algorithm to learn from past conversation data which agent was most effective in what situation, and improve the selection accuracy. The agent selection unit can obtain the profile information of the selected agent from the agent profile storage unit and provide it to the conversation generation unit 213.

[0134] Based on the profile information of the agent provided by the agent selection unit, the conversation generation unit 213 generates a line that matches the role, expertise, and dialogue style of that agent. Specifically, the conversation generation unit 213 gives a prompt including information representing the mental state, conversation content, and agent profile information to a large language model to generate a line that reflects the characteristics of the selected agent. The prompt can include instructions such as "You are a counselor specialized in showing empathy. Based on the following mental state and conversation content, please generate a warm line that accepts the user's feelings."

[0135] The agent cooperation unit manages information sharing and cooperation among multiple agents. The agent cooperation unit can provide functions such as the following, for example.

[0136] (1) Sharing of conversation history: When switching from one agent to another, the context and important points of the conversation up to that point are carried over.

[0137] (2) Cooperation among agents: For complex problems, adjustments are made for multiple agents to cooperate in dealing with them. For example, it enables cooperation such that after the empathy agent provides emotional support, the problem-solving agent proposes specific solutions.

[0138] (3) Natural realization of agent switching: When an agent switches, bridging remarks are generated to maintain the natural flow of conversation. For example, agent switching is naturally expressed in the conversation in the form of "For this problem, I will be replaced by a colleague who can give advice from a more specialized perspective."

[0139] The agent cooperation unit can also simulate conversations among multiple agents using a large language model. For example, when the empathy agent and the problem-solving agent cooperate to deal with a complex problem, the agent cooperation unit gives the large language model a prompt with a setting of "The empathy agent and the problem-solving agent are discussing the user's problem" and can generate a comprehensive response combining the perspectives of both agents.

[0140] The output unit 214 can output remarks in an output format that matches the characteristics of the selected agent. For example, the appearance of a virtual character, voice quality, speech pattern, gestures, etc., unique to each agent are set, and these output characteristics can be switched according to the selected agent. As a result, the user can recognize that they are interacting with different agents both visually and audibly.

[0141] According to Modification Example 2, by having the optimal agent respond according to the user's mental state and the progress of the conversation, more flexible and multifaceted support becomes possible. For example, for a user showing strong emotional fluctuations, a sympathetic agent responds to provide emotional support, and after the emotions have settled, a problem-solving agent proposes specific countermeasures, enabling appropriate support according to the situation. Also, even in cases where complex problems that are difficult to handle with a single agent or multifaceted support are required, by combining the expertise of multiple agents, more comprehensive support becomes possible.

[0142] Moreover, by preparing multiple agents, personalization according to the user's preferences and compatibility becomes possible. For example, it is possible to respond to individual differences such as a certain user responding well to a sympathetic approach and another user preferring a specific problem-solving approach. The agent selection unit can analyze past conversation data with the user and learn the agents that were effective for a specific user, enabling optimized agent selection for each user.

[0143] <Modification Example 3> In Modification Example 3, the long-term change pattern of the user's mental state is tracked, and future state changes are predicted.

[0144] In Modification Example 3, the management server 2 further includes a mental state tracking and prediction unit. The mental state tracking and prediction unit accumulates and analyzes the time-series data of the mental state estimated by the mental state estimation unit 212 to track the long-term change pattern of the user's mental state and predict future state changes. Also, a mental state history storage unit is newly added separately from the personal information storage unit 231. The mental state history storage unit structures and stores the time-series data of the mental state for each user. Specifically, it stores by associating the first and second mental states estimated by the mental state estimation unit 212, the estimation date and time, the situation at the time of estimation (such as a summary of the conversation content), and related biological reaction index values.

[0145] The mental state tracking and prediction unit can perform preprocessing on time-series data. It performs preprocessing such as missing value imputation, outlier handling, and normalization on the time-series data obtained from the mental state history storage unit. As preprocessing methods, a moving average method, a spline interpolation method, a Z-score normalization, etc. can be used.

[0146] Also, the mental state tracking and prediction unit can extract characteristic patterns from time-series data. For pattern extraction, signal processing techniques such as Fourier transform, wavelet transform, empirical mode decomposition (EMD), etc. can be used. Thereby, periodic patterns such as intra-day fluctuations, weekly fluctuations, seasonal fluctuations, and aperiodic patterns related to specific events can be detected.

[0147] Also, the mental state tracking and prediction unit can predict future mental states based on past time-series data of mental states. As prediction models, for example, a recurrent neural network (RNN), a transformer-based time-series prediction model, a state space model (such as a Kalman filter), an ARIMA (autoregressive integrated moving average) model, Prophet, etc. can be adopted.

[0148] Also, the mental state tracking and prediction unit can detect an abnormality when the deviation between the predicted mental state and the actual mental state is large or when a tendency for the mental state to deteriorate rapidly is detected. For anomaly detection, for example, statistical methods (Z-score, modified Z-score, CUSUM method, etc.), density-based methods (LOF, DBSCAN, Isolation Forest, etc.), prediction-based methods (monitoring of prediction errors), etc. can be used.

[0149] Also, the mental state tracking and prediction unit can determine the necessity of intervention based on the detected abnormality or the predicted deterioration of the mental state. The necessity of intervention is comprehensively determined based on the degree of abnormality, the duration, the past mental state change pattern of the user, and medical findings.

[0150] In addition, the mental state tracking and prediction unit can operate in cooperation with the conversation generation unit 213, the notification setting unit 217, the determination unit 216, and the like.

[0151] Based on the future mental state predicted by the mental state tracking and prediction unit and the detected abnormal patterns, the conversation generation unit 213 can generate lines for preventive intervention. For example, when a transition to a depressive state is predicted, lines incorporating a cognitive-behavioral therapy approach for early intervention are generated.

[0152] Based on the change in the mental state predicted by the mental state tracking and prediction unit, the notification setting unit 217 can determine the optimal notification timing. For example, notifications can be set more frequently before the predicted time of mental state deterioration.

[0153] Taking into account the long-term mental state change patterns provided by the mental state tracking and prediction unit, the determination unit 216 can more precisely judge the validity of the lines.

[0154] The mental state tracking and prediction unit acquires the past mental state time series data corresponding to the user ID from the mental state history storage unit, performs preprocessing of the data (missing value completion, outlier processing, normalization), extracts features from the time series data (periodicity, trend, seasonality, etc.), inputs the features into a prediction model (LSTM, Transformer, etc.), predicts the future mental state, applies an anomaly detection algorithm to the predicted mental state, and can calculate the intervention recommendation degree based on the detected anomaly and the predicted degree of mental state deterioration.

[0155] <Disclosure> In addition, the present disclosure also includes the following configurations. [Item 1] An image acquisition unit that acquires an image of the user speaking; A mental state estimation unit that analyzes the image and estimates the mental state of the user; A conversation generation unit that generates a line of dialogue for obtaining empathy from the user based on the presumed mental state, An output unit that outputs the line of dialogue, An information processing system characterized by comprising these. [Item 2] The information processing system according to Item 1, wherein the conversation generation unit generates the line of dialogue by giving a prompt including information representing the presumed mental state and an instruction to generate the line of dialogue based on the information to a large language model. An information processing system characterized by this. [Item 3] The information processing system according to Item 1, wherein it includes a conversation acquisition unit that acquires the conversation content spoken by the user, the mental state estimation unit estimates a second mental state of the user by analyzing the conversation content, in addition to the first mental state analyzed from the image, the conversation generation unit generates the line of dialogue by giving a prompt including information representing the first and second mental states and an instruction to generate the line of dialogue based on the information to a large language model. An information processing system characterized by this. [Item 4] The information processing system according to Item 3, wherein the conversation generation unit generates a question for estimating the second mental state so that the content can be answered more easily the closer it is to the start point of the session with the user, the output unit outputs the generated question to the user, the conversation acquisition unit acquires the conversation content including the answer to the question. An information processing system characterized by this. [Item 5] The information processing system according to Item 3, wherein it includes a determination unit that determines the validity of the line of dialogue based on the second mental state and the conversation content, the conversation generation unit re-creates the line of dialogue according to the validity. An information processing system characterized by [Item 6] The information processing system according to Item 5, wherein the determination unit gives a prompt including an instruction to determine the validity of the script based on the information indicating the second mental state, the conversation content, the script, and the second mental state and the conversation content to a large language model to generate the validity. An information processing system characterized by [Item 7] The information processing system according to Item 3, comprising a notification setting unit that sets a plan to notify the user of a conversation to be conducted next time based on the conversation content when the session with the user ends. An information processing system characterized by [Item 8] The information processing system according to Item 3, comprising a summary notification unit that generates a summary of the conversation content and notifies the user of the summary when the session with the user ends. An information processing system characterized by [Item 9] The information processing system according to Item 3, a personal information extraction unit that extracts the user's personal information from the conversation content when the session with the user ends, a personal information storage unit that stores the personal information, An information processing system characterized by comprising [Item 10] A computer acquires an image of the user speaking, analyzes the image to estimate the mental state of the user, generates a script for obtaining empathy from the user based on the estimated mental state, and outputs the script. An information processing method characterized by

Explanation of Signs

[0156] 1 User terminal 2 Management server

Claims

1. An image acquisition unit that acquires an image of a user speaking; A conversation acquisition unit that acquires the conversation content spoken by the user; A mental state estimation unit that analyzes the image to estimate the first mental state of the user and analyzes the conversation content to estimate the second mental state of the user; A conversation generation unit that generates a line of dialogue for obtaining empathy from the user by giving a prompt including information representing the first and second mental states and an instruction to generate a line of dialogue based on the information to a large language model; An output unit that outputs the line of dialogue; A summary notification unit that generates a summary of the conversation content and notifies the user of the summary at the end of the session with the user; An information processing system, characterized by comprising the above.

2. An image acquisition unit that acquires an image of a user speaking; A conversation acquisition unit that acquires the conversation content spoken by the user; A mental state estimation unit that analyzes the image to estimate the first mental state of the user and analyzes the conversation content to estimate the second mental state of the user; A conversation generation unit that generates a line of dialogue for obtaining empathy from the user by giving a prompt including information representing the first and second mental states and an instruction to generate a line of dialogue based on the information to a large language model; A determination unit that gives a prompt including information indicating the second mental state, the conversation content, the line of dialogue, and an instruction to determine the validity of the line of dialogue based on the second mental state and the conversation content to a large language model to determine the validity of the line of dialogue; An output unit that outputs the line of dialogue; Comprising: The conversation generation unit re-creates the line of dialogue according to the validity; An information processing system, characterized by the above.

3. The information processing system according to claim 1 or 2, wherein the conversation generation unit generates the line of dialogue by giving a prompt including information representing the estimated first and second mental states and an instruction to generate the line of dialogue based on the information to a large language model; An information processing system, characterized by the above.

4. The information processing system according to claim 1 or 2, comprising a conversation acquisition unit that acquires the conversation content spoken by the user, wherein the mental state estimation unit estimates a second mental state of the user by analyzing the conversation content separately from the first mental state analyzed from the image; The conversation generation unit generates the dialogue by providing a prompt including information representing the first and second mental states and an instruction to generate the dialogue based on the information to a large language model. An information processing system characterized by the above.

5. A computer acquires an image of the user speaking, acquires the conversation content spoken by the user, analyzes the image to estimate the first mental state of the user, analyzes the conversation content to estimate the second mental state of the user, generates a dialogue for obtaining empathy from the user by providing a prompt including information representing the first and second mental states and an instruction to generate a dialogue based on the information to a large language model, outputs the dialogue, at the end of the session with the user, generates a summary of the conversation content and notifies the user of the summary. An information processing method characterized by the above.

6. A computer acquires an image of the user speaking, acquires the conversation content spoken by the user, analyzes the image to estimate the first mental state of the user, analyzes the conversation content to estimate the second mental state of the user, generates a dialogue for obtaining empathy from the user by providing a prompt including information representing the first and second mental states and an instruction to generate a dialogue based on the information to a large language model, provides a prompt including information indicating the second mental state, the conversation content, the dialogue, and an instruction to determine the validity of the dialogue based on the second mental state and the conversation content to a large language model to determine the validity of the dialogue, re-creates the dialogue according to the validity, outputs the dialogue. An information processing method characterized by the above.

Citation Information

Patent Citations

  • Information processing system, information processing method, and program

    JP7629254B1

  • Behavior control system

    WO2024214792A1

  • Creation and operation of artificial intelligence-based conversation systems

    JP2022531994A

  • JPP7629254B

Cited By

  • Information processing systems, information processing methods, and programs

    JP7832413B1