Avatar-based interactive medical interview support program and system

The avatar-based medical interview system addresses the lack of personalized avatar speech and emotion estimation by using health data and dialogue generation models to tailor speech and expressions to a patient's condition, enhancing medical interview efficiency.

JP2026089190AActive Publication Date: 2026-06-01CRYSTAL METHOD CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CRYSTAL METHOD CO LTD
Filing Date
2024-11-20
Publication Date
2026-06-01

AI Technical Summary

Technical Problem

Existing technologies do not effectively generate avatar speech tailored to a patient's medical condition and emotions, nor do they estimate and generate facial expressions based on speech content during medical interviews.

Method used

An avatar-based interactive medical interview system that utilizes a health data generation model and an avatar dialogue generation model to tailor speech and facial expressions to a patient's medical condition and emotions, incorporating speech content acquisition, health data recording, analysis, and avatar dialogue generation.

Benefits of technology

The system generates avatar speech and facial expressions aligned with a patient's medical condition, facilitating easy elicitation and recording of physical and mental state information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026089190000001_ABST
    Figure 2026089190000001_ABST
Patent Text Reader

Abstract

The system generates avatar speech that is tailored to the patient's medical condition and speech content, and further generates facial expressions based on an estimation of the avatar's emotions. [Solution] The system acquires user speech content through dialogue with an avatar, uses a health data generation model trained with input data including speech content and output data including health data related to the user's health status as training data, outputs health data based on the acquired speech content, records the output health data in a personal health record, and uses an avatar dialogue generation model trained with input data including health data and output data including the avatar's speech content as training data to read health data from the recorded personal health record and generate the avatar's speech content based on this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an interactive interview support program and system by an avatar capable of conducting an interview while the avatar interacts with a user as a patient.

Background Art

[0002] In recent years, with the shortage of medical staff in medical institutions, a great burden has been imposed on medical workers. As a result, the number of cases leading to the resignation of medical workers has been increasing. Therefore, in order to reduce the burden on medical workers, studies have been conducted more than ever to replace some of the tasks other than actual medical practices in medical institutions with AI.

[0003] In a medical institution, as one of the tasks that can be replaced by AI, before an actual examination by a doctor, a simple interview is conducted for a user as a patient via an avatar equipped with AI, and health data regarding the user's health condition is obtained.

[0004] For example, Patent Document 1 discloses a configuration for obtaining information from a patient in an interactive form before an examination. This Patent Document 1 discloses a technique that enables a doctor to view interview data previously created by a patient before an examination by receiving interview information created by the patient in advance before the doctor's examination and associating the interview data with the doctor in charge of the examination. A user can input interview information such as subjective symptoms in an interactive form in advance via a tablet terminal or the like.

[0005] Also, Patent Document 2 discloses a configuration in which a virtual doctor (avatar) conducts an interview in steps related to an interactive interview. According to the disclosed technique of Patent Document 2, an interview with a patient is conducted through voice-based interaction using voice synthesis and voice recognition, and an interview form is created based on the response from the patient. Thereby, an interview by a so-called virtual doctor expressed by synthesized voice can be provided to the patient.

Prior Art Documents

Patent Documents

[0006] [Patent Document 1] Patent No. 7450310 [Patent Document 2] Patent No. 7454090 [Overview of the project] [Problems that the invention aims to solve]

[0007] However, Patent Documents 1 and 2 do not specifically disclose any technology for generating avatar speech that is more closely aligned with the patient's medical condition and speech content when dealing with an avatar. In addition, they do not specifically disclose any technology for estimating the avatar's emotions according to the patient's medical condition and speech content, and generating the avatar's facial expressions based on this.

[0008] Therefore, the present invention was devised in view of the above-mentioned problems, and its objective is to provide an avatar-based interactive medical interview support program and system that can generate avatar speech content that is tailored to the patient's medical condition and speech content, and furthermore, estimate the avatar's emotions and generate facial expressions accordingly. [Means for solving the problem]

[0009] The avatar-based interactive medical interview support program according to the present invention is characterized by causing a computer to execute the following steps: a speech content acquisition step of acquiring the content of a user's speech through dialogue with an avatar; a recording step of outputting health data based on the speech content acquired in the speech content acquisition step, using a health data generation model that has been trained using input data including the speech content and output data including a response policy for the user as training data, and recording the outputted health data in a personal health record; an analysis step of reading the health data from the personal health record recorded in the recording step, using an analysis model that has been trained using input data including health data and output data including a response policy for the user as training data, and outputting a response policy based on this; and an avatar dialogue generation step of generating the avatar's speech content based on the response policy output in the analysis step, using an avatar dialogue generation model that has been trained using input data including the response policy and output data including the avatar's speech content as training data.

[0010] The avatar-based interactive medical interview support program according to the present invention is characterized by causing a computer to execute the following steps: a speech content acquisition step of acquiring the content of a user's speech through dialogue with an avatar; a recording step of using a health data generation model that has been trained using input data including the speech content and output data including health data relating to the user's health status as training data, to output health data based on the speech content acquired in the speech content acquisition step and record the output health data in a personal health record; and an avatar dialogue generation step of using an avatar dialogue generation model that has been trained using input data including health data and output data including the avatar's speech content as training data, to read the health data from the personal health record recorded in the recording step, and generate the avatar's speech content based on this.

[0011] The avatar-based interactive medical interview support system according to the present invention is characterized by comprising: speech content acquisition means for acquiring the content of a user's speech through dialogue with an avatar; recording means for outputting health data based on the speech content acquired by the speech content acquisition means, and recording the output health data in a personal health record, using a health data generation model learned using input data including speech content and output data including a response policy for the user as learning data; analysis means for reading the health data from the personal health record recorded in the recording step, and outputting a response policy based on the health data, using an analysis model learned using input data including health data and output data including a response policy for the user as learning data; and avatar dialogue generation means for generating the content of the avatar's speech based on the response policy output in the analysis step, using an avatar dialogue generation model learned using input data including a response policy and output data including the avatar's speech content as learning data.

[0012] The avatar-based interactive medical interview support system according to the present invention is characterized by comprising: speech content acquisition means for acquiring the content of a user's speech through dialogue with an avatar; recording means for outputting health data based on the speech content acquired by the speech content acquisition means, and recording the output health data in a personal health record, using a health data generation model that has been learned using input data including the speech content and output data including the avatar's speech content as learning data; and avatar dialogue generation means for reading the health data from the personal health record recorded by the recording means, and generating the avatar's speech content based on this, using an avatar dialogue generation model that has been learned using input data including health data and output data including the avatar's speech content as learning data. [Effects of the Invention]

[0013] According to the present invention, which has the configuration described above, it is possible to generate avatar speech that is tailored to the user's medical condition and speech content, and furthermore, to estimate the avatar's emotions and generate its facial expressions, making it easy to elicit the user's physical and mental state and store this information in a personal health record. [Brief explanation of the drawing]

[0014] [Figure 1] Figure 1 is a block diagram showing the overall configuration of an interactive medical interview support system to which the present invention is applied. [Figure 2] Figure 2 is a block diagram of an interactive medical interview support system to which the present invention is applied. [Figure 3] Figure 3 is a schematic diagram illustrating an example of the management server's functions. [Figure 4] Figure 4 is a flowchart showing the operating procedure of the interactive medical interview support system. [Figure 5] Figure 5 shows the data flow using each training data set. [Figure 6] Figure 6 shows an example of an avatar dialogue generation model that uses input data including response policies and health data, and output data including the avatar's speech content, as training data. [Figure 7] Figure 7 shows an example where the avatar dialogue generation model takes on the role of an analysis model. [Modes for carrying out the invention]

[0015] The following describes an interactive medical interview support system to which the present invention is applied, with reference to the drawings.

[0016] Figure 1 is a block diagram showing the overall configuration of the interactive medical interview support system 100 to which the present invention is applied.

[0017] The interactive medical interview support system 100 has a management server 1. The management server 1 is connected to one or more user terminals 2, for example, via a communication network 4.

[0018] The interactive consultation support system 100 has a personal health record 3 implemented in or externally connected to the management server 1. In addition to various management information to be managed, the management server 1 may store various data such as information related to transmission and reception data with the user terminal 2, communication setting information, various setting information for performing transmission and reception of various data with the user terminal 2, and various data or information such as agreements. The management server 1 is further connected to a large language model (LLM) 5.

[0019] FIG. 2 shows that the management server 1 may have a configuration such as a client-server system or a cloud system, and includes a housing 10, a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a storage unit 104, and a plurality of each I / F 105 to 107, and each configuration is connected by an internal bus 110. The CPU 101 controls the entire management server 1. The ROM 102 stores the operation code of the CPU 101. The RAM 103 is a work area used during the operation of the CPU 101. Information related to the management server 1, types and conditions of various communication procedures, etc. are stored in the storage unit 104. As the storage unit 104, for example, in addition to an HDD (Hard Disk Drive), a data storage device (not shown) such as an SSD (Solid State Drive), a hard disk, or a semiconductor memory may be used.

[0020] The CPU 101 is realized by executing a program stored in the storage unit 104, etc. with the RAM 103 as a work area. For example, the management server 1 may have a GPU (Graphics Processing Unit) not shown. By having a GPU, high-speed arithmetic processing becomes possible compared to normal.

[0021] I / F105 is an interface for sending and receiving various types of information via the communication network 4. I / F105 sends and receives various types of information with the user terminal 2 via the communication network 4. I / F106 is an interface for sending and receiving information with the input unit 108. For example, a keyboard is used as the input unit 108, and the administrator of the management server 1 inputs various types of information, or control commands for setting various types of information, via the input unit 108. I / F107 is an interface for sending and receiving various types of information with the output unit 109. The output unit 109 outputs various types of information stored in the storage unit 104, the processing status of the management server 1, etc. A display is used as the output unit 109, and for example, a touch panel type may be used.

[0022] As user terminal 2, known electronic devices such as personal computers (PCs), smartphones, tablet devices, and 3D holograms managed for each user can be used. An example of the configuration of user terminal 2 may be the same as the schematic diagram shown in Figure 2, and may include, for example, a housing, a CPU, ROM, RAM, a storage unit, respective I / Fs, an input unit, and an output unit. Each component is connected by an internal bus. Other examples of user terminal 2 include any device capable of displaying an avatar, such as any electronic device capable of displaying an avatar as a video on a screen in two dimensions, any electronic device capable of displaying an avatar as a video in 2.5 dimensions, a character summoning device that displays or projects a preferred avatar character into three-dimensional space, a robot such as a doll, pet, or stuffed animal, or a self-propelled robot that performs tasks such as serving and delivering food.

[0023] Personal Health Record 3 is a server capable of recording all information about a user's individual body. Personal Health Record 3 records health data and physical information specific to each user, linked to their individual health status. Examples of health data recorded in Personal Health Record 3 include all vital data related to health status that can be measured through measuring instruments during health checkups or medical examinations, such as height, weight, body temperature, blood pressure, blood glucose level, heart rate, respiratory rate, and stress level. In addition, it also includes data on lifestyle habits such as the content of meals and the nutrients consumed, working hours and overtime hours, daily exercise amount, and sleep duration. Furthermore, this health data also includes information on health status expressed by the user themselves, or detected through oral or written statements, such as the location and degree of pain in the body, the type of pain, and other locations and types of abnormalities in the body. Health data also includes data on chronic illnesses and disabilities, data on current or past diseases, information on medications currently being taken, information on vaccines previously taken, and information on allergies. Furthermore, this health data also includes information on side effects that occurred in response to certain medications or prescriptions.

[0024] This personal health record 3 may be implemented as a function of management server 1, or it may be configured on an independent server. Since this personal health record 3 records personal information, it may be equipped with all known security technologies to ensure security against external hacking of the information.

[0025] The communication network 4 is the Internet network or the like, through which the management server 1 and user terminal 2 are connected via a communication circuit in the interactive medical interview support system 100. The communication network 4 may be composed of a so-called optical fiber communication network. In addition, the communication network 4 may be implemented using a known public communication network such as a wireless communication network, as well as a wired communication network.

[0026] Figure 3 is a schematic diagram showing an example of the functions of the management server 1. The management server 1 comprises at least a speech content acquisition unit 11 that acquires the content of the user's speech, an avatar dialogue generation unit 12 that generates the content of the avatar's speech, an analysis unit 13 that analyzes health data and generates data related to response policies, a vital data detection unit 17 that detects vital data from the user, an output unit 14 that outputs data to the user terminal 2 via the communication network 4, a control unit 15 that controls the inside of the management server 1, a recording unit 16 that stores various other information, and an expression generation unit 18 that generates new expression data.

[0027] The recording unit 16 stores training data for AI such as the health data generation model, analysis model, and avatar dialogue generation model, which will be described later.

[0028] The avatar dialogue generation unit 12 first generates the avatar's utterances as text data. The avatar dialogue generation unit 12 may also appropriately utilize the large-scale language model 5 when generating this text data of the avatar's utterances.

[0029] The expression generation unit 18 has known image control functions necessary for appropriately generating, modifying, and expressing the avatar's facial expressions on the image.

[0030] The management server 1 will cooperate with the user terminal 2 via the communication network 4 in order to perform operations based on these functions.

[0031] Furthermore, the present invention may also involve implementing each of the components of the management server 1, such as the speech content acquisition unit 11, avatar dialogue generation unit 12, analysis unit 13, vital data detection unit 17, recording unit 16, and expression generation unit 18, on the user terminal 2 side. Alternatively, the AI's learning data itself may be entirely implemented on the user terminal 2 side, thus being realized as a so-called edge AI.

[0032] Next, an example of the operation of the interactive medical interview support system 100 in this embodiment will be described.

[0033] Figure 4 is a flowchart showing the operation procedure of this interactive medical interview support system 100, and Figure 5 shows the data flow using each learning data.

[0034] First, in step S11, the user's speech content is obtained through interaction with the avatar. The avatar dialogue generation unit 12 generates the avatar's speech content, and the user terminal 2 converts this into speech and outputs it. Similarly, the expression generation unit 18 generates the avatar's facial expressions, and the user terminal 2 reflects these on the screen. The user viewing the user terminal 2 can experience the feeling of actually having a conversation with the avatar through the avatar's facial expressions displayed on the screen and the speech content output as audio. For example, by using an image of a nurse as the avatar, the user will feel as if they are receiving a medical consultation from a nurse regarding their health.

[0035] Through this interaction with the avatar, the user provides information about their physical condition. Conversely, by including questions about the user's physical condition, such as "How are you feeling today?", the avatar can create an atmosphere that encourages the user to actively share information about their health. Furthermore, by including more specific questions when the avatar asks about the user's physical condition, the user can provide information such as when the pain or discomfort started, which parts of the body are affected, what kind of pain or discomfort it is, and to what degree, allowing the user to provide necessary information before an actual medical examination by a doctor.

[0036] In addition, through interaction with the avatar, users can also make the avatar speak about their lifestyle, chronic illnesses and disabilities, current or past illnesses, and even information about medications they are currently receiving and allergies.

[0037] User terminal 2 converts the user's spoken content into text data, temporarily stores it, and transmits it to management server 1 via communication network 4. On the other hand, if this system is implemented as edge AI and each component is implemented on the user terminal 2 side, the operation of transmitting the spoken content as text data to management server 1 via communication network 4 can be omitted.

[0038] In the example described above, the user freely speaks through dialogue with the avatar, and the audio is acquired via user terminal 2. However, the explanation is not limited to this. When acquiring information from the user about their medical condition, physical condition, lifestyle, and other chronic illnesses, in addition to audio, the user may be allowed to input information from options displayed on the user interface of user terminal 2 via the touch panel or buttons, and this may be used as the spoken content. Alternatively, the user may be allowed to freely input text data via the touch panel or buttons, and this may also be used as the spoken content.

[0039] Furthermore, in step S11, in addition to acquiring the content of the user's speech, vital data detected from the user may also be acquired. Vital data here includes many data that can be measured from the body by measuring instruments, such as body temperature, height, weight, blood pressure, blood glucose level, heart rate, respiratory rate, pulse rate, and blood oxygen saturation. This vital data may be acquired by directly taking in data detected from actual measuring instruments, or the user may manually input previously measured data into the user terminal 2.

[0040] In step S11, the facial expressions of the user interacting with the avatar may also be captured by a camera provided on the user terminal 2.

[0041] In step S11, the voice data of the user interacting with the avatar may also be detected by a microphone provided on the user terminal 2. The voice data referred to here includes things other than the content of the speech, such as tone of voice, voiceprint, volume, and accent.

[0042] Next, we proceed to step S12, where the health data generation model is used to output health data.

[0043] As shown in Figure 4, the health data generation model is a trained model that has been trained using input data, including utterances, and output data, including health data related to the user's health status, as training data. This health data generation model is formed based on machine learning using artificial intelligence, but is not limited to this, and the output data may also be generated using a large-scale language model 5.

[0044] As a method for generating health data generation models, for example, they may be generated using machine learning modeled after neural networks. The health data generation model may also be a multimodal model. The health data generation model is trained using machine learning modeled after neural networks such as CNN (Convolutional Neural Network), or any other model may be used. Furthermore, as a method for generating health data generation models, they may be generated using methods such as Retrieval-Augmented Generation (RAG), Seq2Seq (Sequence To Sequence) linear discriminant analysis, support vector machines, k-nearest neighbors, random forests, deep learning, etc.

[0045] Through a health data generation model trained on such input and output data, health data can be obtained from the text data of user utterances. For example, if the text data of the utterance is "My stomach area hurts today," then corresponding health data such as "stomach ache," "excessive stomach acid," and "stomach discomfort" can be obtained. Furthermore, even if the text data of the utterance is a vague expression such as "My head feels foggy," health data such as "headache," "autonomic nervous system dysfunction," and "dizziness" can be obtained.

[0046] Furthermore, the input data for the health data generation model may include vital data in addition to speech content. Vital data is data that can be measured from the body using measuring instruments as described above. By training the health data generation model with this vital data together with the health data output, it becomes possible to output the detected vital data as more organized health data.

[0047] Furthermore, the input data for the health data generation model may include not only the spoken content but also the user's facial expressions. By training the model with image data of facial expressions and the health data output from the health data generation model, it becomes possible to output the captured image data of the user's facial expressions as more organized health data. In particular, since a user's health status may change depending on their facial expressions, by training the health data generation model in advance to understand these relationships, it becomes possible to generate health data from facial expressions.

[0048] Furthermore, the input data for the health data generation model may include not only the spoken content but also the audio data. By training the model with the audio data and the health data output from the health data generation model, it becomes possible to output the detected audio data as more organized health data. In particular, since health status may change depending on the tone of voice, etc., by training the health data generation model to learn these relationships and building it in advance, it becomes possible to generate health data from audio data.

[0049] Next, we proceed to step S13, where the health data output from the health data generation model described above is linked to the user and recorded in the personal health record 3 described above.

[0050] Next, we move to step S14, where we use the analysis model to output a strategy for dealing with the user.

[0051] As shown in Figure 4, the analysis model is a model trained using input data, including health data, and output data, including user response strategies, as training data. This analysis model is formed based on machine learning using artificial intelligence, but is not limited to this; a large-scale language model 5 may also be used to generate the output data.

[0052] The analysis model may be generated using machine learning, for example, a neural network model. The analysis model may also be a multimodal model. The analysis model may be trained using machine learning, for example, a neural network model such as a CNN, or any other model may be used. Furthermore, the analysis model may be generated using methods such as search augmentation, seq2seq linear discriminant analysis, support vector machines, k-nearest neighbors, random forests, or deep learning.

[0053] The response strategy data output from this analysis model consists of patterns of data indicating what kind of response should be taken for each user. For example, "User A, who has characteristic XX, should be approached in the following way," "User B, who has characteristic YY, should be shown empathy through dialogue in the following way," "User C, who has characteristic ZZ, should be given advice in the following way," and "User D, who has characteristic YY, should be asked further questions regarding the following." Since this response strategy is generated based on health data, it may include not only physical and mental care and further questions, but also specific medical advice and treatment plans. Incidentally, information about user characteristics in this response strategy data may be additional, and it may even consist only of information indicating what kind of response should be taken, omitting such information.

[0054] Health data can be analyzed through an analytical model trained on such input and output data, and as a result, data regarding response strategies can be obtained for each individual user based on their health data. When actually using this analytical model to generate an exploratory solution, the health data associated with the user being responded to by the avatar is read from Personal Health Record 3 and input into the analytical model. As a result, a response strategy is output as the exploratory solution of the analytical model.

[0055] Next, we move to step S15, where we use the avatar dialogue generation model to generate the avatar's speech.

[0056] As shown in Figure 4, the avatar dialogue generation model is a model trained using input data, which includes the response policy to the user, and output data, which includes the content of the avatar's speech, as training data. This avatar dialogue generation model is formed based on machine learning using artificial intelligence, but is not limited to this, and output data may also be generated using a large-scale language model 5.

[0057] As a method for generating avatar dialogue generation models, for example, they may be generated using machine learning modeled after neural networks. The avatar dialogue generation model may also be a multimodal model. The avatar dialogue generation model is trained using machine learning modeled after neural networks such as CNNs, or any other model may be used. Furthermore, as a method for generating avatar dialogue generation models, they may be generated using methods such as search augmentation, seq2seq linear discriminant analysis, support vector machines, k-nearest neighbors, random forests, deep learning, etc.

[0058] The avatar's utterances output by the avatar dialogue generation model include specific questions about the user's health and medical condition, as well as specific words of encouragement, empathy, sympathy, care, and emotional support. In addition to utterances for medical diagnosis, consultation, and understanding the user's situation, the utterances may also include words that brighten the user's mood or help them relax.

[0059] Through an avatar dialogue generation model trained with such input and output data, it is possible to generate specific avatar utterances based on response strategies.

[0060] In step S15, the avatar's speech content may be generated through the avatar dialogue generation model, and the avatar's facial expressions may also be generated. In this case, an avatar dialogue generation model trained using input data including the response policy and output data including the avatar's speech content and the avatar's emotions as training data is used. The avatar dialogue generation model then outputs the avatar's emotions in addition to the avatar's speech content as described above. By associating each emotion of the avatar with the avatar's facial expressions, it is possible to generate avatar facial expressions on the image that correspond to the outputted emotion of the avatar.

[0061] Finally, the process moves to step S16, where the avatar speaks to the user via user terminal 2 based on the utterances generated through the avatar dialogue generation model. At this time, if the avatar dialogue generation model outputs an emotion for the avatar, the corresponding facial expression of the avatar can be displayed as an image.

[0062] Thus, according to the present invention, the user's utterances are converted into health data through a health data generation model, appropriate response strategies are explored through an analysis model, and specific utterances are generated through an avatar dialogue generation model. As a result, the avatar can make utterances that correspond to the user's health status and condition, which have been analyzed from the user's utterances. Even if the user has anxieties about their health status or suffers from an illness or injury, they can alleviate their anxieties, receive appropriate treatment and responses through conversation with the avatar, and also receive mental support through emotional care.

[0063] In such cases, if the avatar's emotions are further output from the avatar dialogue generation model, displaying the corresponding avatar facial expression as an image will result in an avatar expression that is appropriate to the user's situation, conveying empathy and sympathy to the user and creating the impression of a caring and attentive response.

[0064] In this invention, as shown in Figure 6, the avatar dialogue generation model may utilize a model trained using input data including response policies and health data, and output data including the avatar's speech content, as training data. Based on both health data and response policies, the avatar's speech content can be generated. At this time, by training the model with output data including the avatar's speech content and the avatar's emotions, it is also possible to generate facial expressions based on the avatar's emotions.

[0065] Furthermore, according to the present invention, as shown in Figure 7, the avatar dialogue generation model may also be made to take on the role of an analysis model. In this case, an avatar dialogue generation model trained using input data including health data and output data including the avatar's speech content as training data is used. By reading the health data recorded in the personal health record and training the avatar's speech content accordingly, it is possible to directly generate the avatar's speech content based on the health data. At this time, by training the avatar's speech content and output data including the avatar's emotions, it is also possible to generate facial expressions based on the avatar's emotions. [Explanation of Symbols]

[0066] 1. Management Server 2 User terminals 3. Personal Health Records 4. Communication Network 5. Large-scale language models 10 cabinets 11. Speech content acquisition unit 12 Avatar Dialogue Generation Unit 13 Analysis Department 14 Output section 15 Control Unit 16 Records Section 17. Vital Data Detection Unit 18 Expression generator 100 Interactive Medical Interview Support Systems 101 CPU 102 ROM 103 RAM 104 Preservation Department 105~107 I / F 108 Input section 109 Output section 110 Internal bus

Claims

1. The speech content acquisition step involves acquiring the content of the user's utterances through interaction with an avatar, A recording step is performed using a health data generation model trained with input data including utterances and output data including health data related to the user's health status as training data, to output health data based on the utterances obtained in the utterance acquisition step, and to record the output health data in a personal health record. An analysis step that uses an analytical model trained with input data including health data and output data including user response policies as training data, reads the health data from the personal health record recorded in the recording step, and outputs a response policy based on this, Using an avatar dialogue generation model trained with input data including response policies and output data including the avatar's speech content as training data, the computer is to execute an avatar dialogue generation step that generates the avatar's speech content based on the response policies output in the analysis step. An interactive medical interview support program using avatars, characterized by the following features.

2. The speech content acquisition step involves acquiring the content of the user's utterances through interaction with an avatar, A recording step is performed using a health data generation model trained with input data including utterances and output data including health data related to the user's health status as training data, to output health data based on the utterances obtained in the utterance acquisition step, and to record the output health data in a personal health record. Using an avatar dialogue generation model trained with input data including health data and output data including the avatar's speech content as training data, the computer is instructed to perform an avatar dialogue generation step which involves reading the health data from the personal health record recorded in the recording step and generating the avatar's speech content based on this data. An interactive medical interview support program using avatars, characterized by the following features.

3. In the avatar dialogue generation step described above, an avatar dialogue generation model trained using input data including response policies and output data including the avatar's speech content and the avatar's emotions as training data is used to generate facial expressions corresponding to the avatar's emotions, along with the avatar's speech content, based on the response policies output in the analysis step described above. An avatar-based interactive medical interview support program according to claim 1, characterized by the above.

4. The system further includes a vital data detection step for detecting vital data from the above-mentioned user, In the recording step described above, a health data generation model trained using input data including speech content and vital data, and output data including health data related to the user's health status, as training data is used, and further, the health data is output based on the vital data detected in the vital data detection step described above. An avatar-based interactive medical interview support program according to claim 1 or 2, characterized by the above.

5. The system further includes a facial expression detection step that detects the facial expressions of the user described above. In the recording step described above, a health data generation model trained using input data including speech content and facial expressions, and output data including health data related to the user's health status, as training data is used, and the health data is further output based on the facial expressions detected in the facial expression detection step described above. An avatar-based interactive medical interview support program according to claim 1 or 2, characterized by the above.

6. The system further includes a voice data detection step for detecting the voice data of the above-mentioned user, In the recording step described above, a health data generation model trained using input data including the spoken content and the above-mentioned audio data, and output data including health data related to the user's health status, as training data is used, and further, the health data is output based on the audio data detected in the above-mentioned audio data detection step. An avatar-based interactive medical interview support program according to claim 1 or 2, characterized by the above.

7. A means for acquiring speech content that obtains the content of user speech through interaction with an avatar, A recording means that uses a health data generation model trained using input data including speech content and output data including health data related to the user's health status as training data, outputs health data based on the speech content acquired by the speech content acquisition means, and records the output health data in a personal health record, An analysis means that uses an analysis model trained with input data including health data and output data including a response policy for the user as training data, reads the health data from the personal health record recorded in the recording step above, and outputs a response policy based on this, The system includes an avatar dialogue generation means that uses an avatar dialogue generation model trained using input data including response policies and output data including the avatar's speech content as training data, and generates the avatar's speech content based on the response policies output in the analysis step. An interactive medical interview support system using avatars, characterized by the following features.

8. A means for acquiring speech content that obtains the content of user speech through interaction with an avatar, A recording means that uses a health data generation model trained using input data including speech content and output data including health data related to the user's health status as training data, outputs health data based on the speech content acquired by the speech content acquisition means, and records the output health data in a personal health record, The system includes an avatar dialogue generation means that utilizes an avatar dialogue generation model trained using input data including health data and output data including the avatar's speech content as training data, reads the health data from the personal health record recorded by the recording means, and generates the avatar's speech content based on this. An interactive medical interview support system using avatars, characterized by the following features.