Information processing system and information processing method
The information processing system addresses the lack of mental health monitoring in virtual assistants by using voice calls and large-scale language models to assess and support users' mental health continuously, offering timely and automated support.
Patent Information
- Application Number
- JP2025098076
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-11-05
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Conventional virtual assistant technologies lack the capability to continuously monitor and support users' mental health status effectively.
An information processing system that includes a call establishment unit, an avatar dialogue unit, an answer acquisition unit, and a state estimation unit to conduct voice calls, synthesize questions, acquire user responses, and estimate mental health states using large-scale language models, with follow-up support based on estimation results.
The system effectively checks and supports users' mental health by providing continuous monitoring and appropriate interventions, reducing user burden through natural dialogue and automating expert collaboration.
Smart Images

Figure 0007764082000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system and an information processing method. [Background technology]
[0002] A technology has been proposed for triggering a virtual assistant that enables a user to interact with a device using natural language in a spoken format (see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6873038 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional virtual assistant technologies are limited to general user interactions and lack the mechanisms to continuously monitor users' mental health status and provide support at the appropriate time.
[0005] The present invention has been made in view of the above background, and aims to provide a technique that can effectively check and support the mental health state of a user. [Means for solving the problem]
[0006] The main invention of the present invention for solving the above problem is an information processing system comprising a call establishment unit that makes a voice call to a user's communication terminal to establish the call, an avatar dialogue unit that voice-synthesizes questions to ascertain the user's mental health state generated by a large-scale language model and presents the questions during the call, an answer acquisition unit that acquires the user's voice response, and a state estimation unit that estimates the user's mental health state based on the acquired voice response.
[0007] Other problems and solutions disclosed in this application will be made clear in the section on preferred embodiments of the invention and the drawings. [Effects of the Invention]
[0008] According to the present invention, the mental health state of a user can be effectively checked and supported. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of an information processing system. [Figure 2] FIG. 2 illustrates an example of a hardware configuration of a management server 2. [Figure 3] FIG. 2 illustrates an example of the software configuration of a management server 2. [Figure 4] FIG. 2 is a diagram illustrating an example of the software configuration of a user terminal 1. [Figure 5] FIG. 2 is a diagram illustrating an example of the software configuration of the therapist terminal 3. [Figure 6] FIG. 2 is a diagram illustrating a processing flow in the information processing system. DETAILED DESCRIPTION OF THE INVENTION
[0010] <System Overview> An information processing system according to one embodiment of the present invention will be described below. The information processing system of this embodiment is a system that continuously checks and supports a user's mental health status. The system makes a voice call to the user's communication terminal, generates questions to check the user's mental health status using a large-scale language model, and estimates the user's mental health status through dialogue with the user. Furthermore, the system has the function of conducting follow-up at an appropriate time based on the estimation results and consulting with experts as necessary.
[0011] 1 is a diagram showing an example of the overall configuration of an information processing system. The information processing system of this embodiment is configured to include a management server 2. The management server 2 is communicably connected to a user terminal 1 and a therapist terminal 3 via a communication network. The communication network is, for example, the Internet, and is constructed using a public telephone line network, a mobile phone line network, a wireless communication path, Ethernet (registered trademark), etc.
[0012] The user terminal 1 is a computer operated by a user, and may be, for example, a smartphone, a tablet computer, or a personal computer.
[0013] The therapist terminal 3 is a computer operated by a professional therapist. The therapist terminal 3 may be, for example, a smartphone, a tablet computer, or a personal computer.
[0014] The management server 2 may be a general-purpose computer such as a workstation or a personal computer, or may be logically realized by cloud computing.
[0015] <Administration Server 2> FIG. 2 is a diagram illustrating an example of the hardware configuration of the management server 2. Note that the illustrated configuration is an example, and other configurations may also be used. The management server 2 includes a CPU 201, a memory 202, a storage device 203, a communication interface 204, an input device 205, and an output device 206. The storage device 203 stores various data and programs, and is, for example, a hard disk drive, a solid state drive, or a flash memory. The communication interface 204 is an interface for connecting to a communication network, and is, for example, an adapter for connecting to Ethernet (registered trademark), a modem for connecting to a public telephone network, a wireless communication device for wireless communication, or a USB (Universal Serial Bus) connector or an RS232C connector for serial communication. The input device 205 is used to input data, and is, for example, a keyboard, a mouse, a touch panel, a button, a microphone, or the like. The output device 206 is used to output data, and is, for example, a display, a printer, a speaker, or the like. Each functional unit of the management server 2 described below is realized by the CPU 201 reading a program stored in the storage device 203 into the memory 202 and executing it, and each storage unit of the management server 2 is realized as part of the storage area provided by the memory 202 and the storage device 203.
[0016] 3 is a diagram illustrating an example of the software configuration of the management server 2. The management server 2 includes a call establishment unit 211, an avatar dialogue unit 212, a response acquisition unit 213, a state estimation unit 214, a scheduling unit 215, an alert transmission unit 216, a user information storage unit 231, a session history storage unit 232, and a model storage unit 233.
[0017] 4 is a diagram showing an example of the software configuration of the user terminal 1. The user terminal 1 includes a call control unit 111, a voice input / output unit 112, a data transmission / reception unit 113, and a user setting storage unit 131.
[0018] 5 is a diagram showing an example of the software configuration of the therapist terminal 3. The therapist terminal 3 includes an alert receiving unit 311, a display control unit 312, a response transmitting unit 313, and a patient information storage unit 331.
[0019] <Administration Server 2> The functional parts of the management server 2 will be described below.
[0020] The user information storage unit 231 stores basic information about the user, such as the user's identification information, contact information, sleeping time zone information, calendar information, and past mental health evaluation results.
[0021] The session history storage unit 232 stores information about past dialogue sessions with the user, including the date and time of each session, the content of the dialogue, question and answer pairs, estimated mental health status, and the schedule for the next session.
[0022] The model storage unit 233 stores large-scale language models and deep learning models. The model storage unit 233 stores large-scale language models fine-tuned using call scripts and case data sets created by clinical psychologists, deep learning models that estimate sentiment from speech features, and the like.
[0023] The call establishment unit 211 originates a voice call to the user's communication terminal and establishes the call. The call establishment unit 211 references the user's contact information stored in the user information storage unit 231 and originates the voice call to the user terminal 1 at the timing determined by the scheduling unit 215. The call establishment unit 211 receives a response from the user terminal 1 and establishes a call session, creating a state in which voice data can be sent and received.
[0024] The avatar dialogue unit 212 synthesizes voice questions generated by a large-scale language model to check the user's mental health status and presents the synthesized questions during the call. The avatar dialogue unit 212 inputs prompts including the history of previous sessions to the large-scale language model stored in the model storage unit 233 and acquires questions output from the model. For example, the avatar dialogue unit 212 generates a continuous question such as, "We talked about sleep in the last session. How is your sleep going since then?"
[0025] The avatar dialogue unit 212 inputs a prompt including parameters indicating the mental health state estimated by the state estimation unit 214 into a large-scale language model, and generates and presents an empathetic response and a suggested action according to the mental health state. For example, if it is estimated that the user's stress level is high, the avatar dialogue unit 212 generates an empathetic response such as "You look tired. Why don't you take a deep breath?" and a specific suggested action.
[0026] The answer acquisition unit 213 acquires the user's answer voice. The answer acquisition unit 213 receives voice data transmitted from the user terminal 1 during the call session and converts the voice data into text data using voice recognition technology. The answer acquisition unit 213 inputs a character string obtained by transcribing the voice data of the entire call into a large-scale language model, generates a summary of the previous session and proposed follow-up questions for the next session, and outputs them to the avatar dialogue unit 212.
[0027] The state estimation unit 214 estimates the user's mental health state based on the acquired answer voice. The state estimation unit 214 maps the user's answer to a mental health state assessment scale such as the Patient Health Questionnaire-9 (PHQ-9) or the Generalized Anxiety Disorder 7-item scale (GAD-7), and quantifies the mental health state based on the score. For example, if the user answers, "Nothing I do these days is fun," the state estimation unit 214 evaluates this by associating it with the "Lack of interest or pleasure in things" item in the PHQ-9.
[0028] The state estimation unit 214 extracts the voice intensity, speaking rate, and pitch contained in the response voice and inputs them into a deep learning model to estimate sentiment. The state estimation unit 214 comprehensively analyzes both the voice features and the text content to estimate the mental health state. If the estimation result exceeds a predetermined risk threshold, the state estimation unit 214 executes the alert transmission unit 216.
[0029] The scheduling unit 215 determines the timing and frequency of the next call based on the estimation result by the state estimation unit 214 and the session history. The scheduling unit 215 determines the start time of the call taking into consideration the sleeping time zone information and calendar information acquired from the user terminal 1. For example, if the mental health state is deteriorating, the scheduling unit 215 increases the frequency of calls, and if the mental health state is stable, the scheduling unit 215 decreases the frequency.
[0030] The alert sending unit 216 is executed by the state estimation unit 214 and sends an alert to the specialized therapist terminal 3. When the user's mental health state reaches a dangerous level, the alert sending unit 216 sends alert information indicating urgency to the therapist terminal 3.
[0031] <User device 1> The functional units of the user terminal 1 will be described below.
[0032] The user setting storage unit 131 stores personal setting information of the user, such as sleeping hours, calendar information, and whether or not to receive calls.
[0033] The call control unit 111 receives a voice call from the management server 2 and controls the call session. The call control unit 111 notifies the user when an incoming call is received, and starts or rejects the call based on the user's response.
[0034] The audio input / output unit 112 controls audio communication with the user. The audio input / output unit 112 outputs audio data received from the management server 2 from a speaker, and transmits the user's voice acquired from a microphone to the management server 2.
[0035] The data transmission / reception unit 113 transmits and receives data to and from the management server 2. The data transmission / reception unit 113 transmits setting information stored in the user setting storage unit 131 to the management server 2, and receives session results and schedule information.
[0036] <Therapist Terminal 3> The functional units of the therapist terminal 3 will be described below.
[0037] The patient information storage unit 331 stores information about patients for which the patient is responsible, such as basic information about the patient, treatment history, and alert history.
[0038] The alert receiving unit 311 receives alert information from the management server 2. When the alert receiving unit 311 receives an alert with high urgency, it immediately notifies the therapist.
[0039] The display control unit 312 outputs the received alert information and patient information to a display screen. The display control unit 312 visually displays the progress of the patient's mental health status, the details of the most recent session, recommended responses, and the like.
[0040] Based on the therapist's decision, the response sending unit 313 sends response information to the management server 2. The response sending unit 313 sends to the management server 2 information such as whether or not to contact the patient directly and instructions to change the treatment plan.
[0041] FIG. 6 is a diagram illustrating a processing flow in the information processing system.
[0042] The scheduling unit 215 of the management server 2 determines the timing of the next call by referring to the user information storage unit 231 and the session history storage unit 232 (S101). The call establishment unit 211 places a voice call to the user terminal 1 at the determined timing and establishes the call (S102). The avatar dialogue unit 212 generates a question to confirm the mental health state using a large-scale language model, synthesizes the voice, and presents the question during the call (S103). The answer acquisition unit 213 acquires the user's answer voice from the user terminal 1 (S104). The state estimation unit 214 estimates the user's mental health state based on the acquired answer voice (S105). The state estimation unit 214 determines whether the estimation result exceeds a predetermined risk threshold (S106). If the threshold is exceeded, the alert transmission unit 216 sends an alert to the therapist terminal 3 (S107). The scheduling unit 215 determines the timing and frequency of the next call based on the estimation result and the session history (S108), and records the session information in the session history storage unit 232 (S109).
[0043] As described above, the information processing system of this embodiment utilizes a large-scale language model to continuously and effectively check the user's mental health status and provide support at an appropriate time. Furthermore, it is possible to achieve objective state estimation based on specialized evaluation criteria while reducing the burden on the user through natural dialogue via voice calls. Furthermore, by automating collaboration with experts based on the estimation results, it is possible to provide prompt and appropriate medical support.
[0044] Although the present embodiment has been described above, the above embodiment is intended to facilitate understanding of the present invention and is not intended to limit the present invention. The present invention may be modified or improved without departing from the spirit thereof, and equivalents thereof are also included in the present invention.
[0045] For example, the processing by each of the functional units of the management server 2 described above may be performed by any of the functional units. Also, a different functional unit that performs part of the processing by each of the functional units described above may be added. Also, the functional units of the management server 2 may be distributed across multiple computers.
[0046] Furthermore, the information stored in each storage unit of the management server 2 may be stored in any of the storage units. That is, the information stored in the above-mentioned multiple storage units may be stored in one storage unit, or part of the information stored in one of the above-mentioned storage units may be stored in another storage unit.
[0047] <Variation 1> In the above-described embodiment, an example was shown in which a dialogue with a user was performed using a voice call, but a video call may also be used. In this case, the avatar dialogue unit 212 generates a 3D avatar image in addition to voice synthesis and displays it on the screen of the user terminal 1. The state estimation unit 214 extracts visual features such as the user's facial expression, eye movement, and posture through image analysis, and combines them with the audio features to more precisely estimate the mental health state. For example, if the user's facial expression is gloomy or their gaze is downward, it is determined that there is a high possibility that the user is depressed.
[0048] <Variation 2> In the above-described embodiment, an example has been shown in which the management server 2 actively makes a call to the user terminal 1, but the user may also be allowed to make a call to the management server 2 at any time. In this case, an emergency call button is provided on the user terminal 1 so that the user can immediately start a dialogue session with the management server 2 when they sense a mental health problem. The call establishment unit 211 receives an incoming call from the user terminal 1 and establishes a call, and the avatar dialogue unit 212 quickly grasps the user's condition using a question set for emergency situations.
[0049] <Variation 3> Although the above-described embodiment illustrates an example in which a single large-scale language model is used, multiple specialized models may be combined. For example, a model dedicated to depression, a model dedicated to anxiety disorders, a model dedicated to PTSD (Post-Traumatic Stress Disorder), etc. may be prepared, and an appropriate model may be selected based on the initial diagnosis result by the condition estimation unit 214. The avatar dialogue unit 212 generates specialized questions according to the selected model, thereby achieving a more precise condition assessment.
[0050] <Variation 4> In the above-described embodiment, an example was shown in which the PHQ-9 or GAD-7 rating scale was used, but other standardized rating scales may also be used in combination. For example, the state estimation unit 214 performs a multifaceted evaluation by combining rating scales such as the Montgomery-Asberg Depression Rating Scale (MADRS), the Hamilton Anxiety Rating Scale (HAMS), and the Perceived Stress Scale (PSS). The results of each rating scale are integrated to calculate a comprehensive mental health score, thereby achieving a more comprehensive understanding of the state.
[0051] <Variation 5> In the above-described embodiment, an example was shown in which an alert was sent to the therapist terminal 3, but it is also possible to send notifications to emergency contacts such as the user's family and friends. In this case, an emergency contact list is stored in the user information storage unit 231, and the alert sending unit 216 sends notifications in stages according to the seriousness of the situation. A multi-stage alert function is realized, where if the situation is mild, only the therapist is notified, if the situation is moderate, the family is also notified, and if the situation is severe, emergency services are automatically notified as well.
[0052] <Disclosures> The present disclosure also includes the following configurations. [Item 1] a call establishment unit that initiates a voice call to a user's communication terminal and establishes the call; an avatar dialogue unit that synthesizes questions to check the user's mental health status generated by a large-scale language model and presents the questions during the call; an answer acquisition unit that acquires the answer voice of the user; a state estimation unit that estimates the mental health state of the user based on the acquired answer voice; An information processing system comprising: [Item 2] Item 2. The information processing system according to item 1, further comprising a scheduling unit that determines the timing and frequency of the next call based on the estimation result by the state estimation unit and the session history. [Item 3] Item 10. The information processing system according to item 1, wherein the avatar dialogue unit inputs a prompt including a previous session history into the large-scale language model, and acquires and presents the question output from the model. [Item 4] Item 1. The information processing system of claim 1, wherein the avatar dialogue unit inputs a prompt including parameters indicating the mental health state estimated by the state estimation unit into the large-scale language model, and generates and presents an empathetic response and behavioral suggestion according to the mental health state. [Item 5] The information processing system described in item 1, wherein the state estimation unit executes an alert sending unit that sends an alert to a specialized therapist terminal when the estimation result based on the answer voice exceeds a predetermined risk threshold. [Item 6] Item 2. The information processing system according to item 1, wherein the state estimation unit maps the user's answers to a mental health state assessment scale and quantifies the mental health state based on the score. [Item 7] Item 2. The information processing system of item 1, wherein the large-scale language model is fine-tuned using call scripts and case datasets created by clinical psychologists. [Item 8] Item 3. The information processing system according to item 2, wherein the scheduling unit determines the start time of the call taking into consideration sleeping time zone information and calendar information acquired from the user terminal. [Item 9] Item 1. The information processing system according to item 1, wherein the state estimation unit extracts the voice intensity, speaking rate, and pitch contained in the response voice and inputs them into a deep learning model to estimate sentiment. [Item 10] The information processing system described in item 1, wherein the answer acquisition unit inputs a string of characters transcribed from the audio data of the entire call into the large-scale language model, generates a summary of the previous session and proposed follow-up questions for the next session, and outputs them to the avatar dialogue unit. [Item 11] Initiating a voice call to a user's communication terminal and establishing the call; a step of synthesizing a question to check the user's mental health status generated by a large-scale language model and presenting the question to the user during the call; acquiring a response voice of the user; estimating a mental health state of the user based on the acquired answer voice; An information processing method performed by a computer. [Explanation of symbols]
[0053] 1. User terminal 2 Management Server 3 Therapist terminal
Claims
1. a call establishment unit that establishes a voice call by making the voice call to the user's communication terminal; an avatar dialogue unit that synthesizes questions to confirm the user's mental health status, generated using a large-scale language model, and presents the questions during the call; an answer acquisition unit that acquires the answer voice of the user; a state estimation unit that estimates a mental health state of the user based on the acquired answer voice; a scheduling unit that determines the timing and frequency of the next call based on the estimation result by the state estimation unit and the session history; Equipped with The scheduling unit determines the start time of the call in consideration of sleeping time zone information and calendar information acquired from the user terminal.
2. The information processing system according to claim 1 , wherein the avatar dialogue unit inputs a prompt including a history of a previous session to the large-scale language model, and acquires and presents the question output from the model.
3. 2. The information processing system according to claim 1, wherein the avatar dialogue unit inputs a prompt including a parameter indicating the mental health state estimated by the state estimation unit into the large-scale language model, and generates and presents an empathetic response and action suggestion according to the mental health state.
4. The information processing system according to claim 1 , wherein the state estimation unit executes an alert sending unit that sends an alert to a specialized therapist terminal when an estimation result based on the answer voice exceeds a predetermined risk threshold.
5. The information processing system according to claim 1 , wherein the state estimation unit extracts a voice intensity, a speaking rate, and a pitch included in the response voice, and inputs the extracted voice intensity, speaking rate, and pitch into a deep learning model to estimate a sentiment.
6. The information processing system according to claim 1 , wherein the answer acquisition unit inputs a character string obtained by transcribing the voice data of the entire call into the large-scale language model, generates a summary of the previous session and a proposed follow-up question for the next session, and outputs the summary and the proposed follow-up question to the avatar dialogue unit.
7. Initiating a voice call to a user's communication terminal and establishing the call; a step of synthesizing a question to check the user's mental health status generated by a large-scale language model and presenting the question to the user during the call; acquiring a response voice of the user; estimating a mental health state of the user based on the acquired answer voice; determining the timing and frequency of the next call based on the estimation result in the step of estimating the mental health state and the session history; The computer executes In the step of determining the timing and frequency of the next call, the computer determines the start time of the call by taking into account sleeping time zone information and calendar information acquired from the user terminal.
Citation Information
Patent Citations
Action control system
JP2025026419A
System
JP2025048985A
System
JP2025049138A
System
JP2025049486A
System
JP2025055007A