Avatar-based interactive medical interview support program and system

The interactive medical questioning support system uses a health data generation model and an avatar dialogue generation model to align avatar responses, including facial expressions, with a patient's medical condition and emotional state, addressing the limitations of existing systems in generating relevant and emotionally supportive interactions.

JP7676075B1Active Publication Date: 2025-05-14CRYSTAL METHOD CO LTD

Patent Information

Application Number
JP2024202086
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-20
Publication Date
2025-05-14
Estimated Expiration
2044-11-20

AI Technical Summary

Technical Problem

Existing medical questioning support systems using avatars do not effectively generate utterance content aligned with a patient's medical condition and speech content, nor do they estimate and generate facial expressions based on the patient's emotional state.

Method used

An interactive medical questioning support program and system using an avatar that employs a health data generation model and an avatar dialogue generation model. The system acquires user utterance content through dialogue, uses this data as training input, and generates health data and corresponding avatar responses, including facial expressions based on emotional estimation.

Benefits of technology

The system enables the generation of avatar utterances and facial expressions that are in line with a user's medical condition and speech content, allowing for effective health data collection and emotional support, thereby aiding users in understanding their physical and mental state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676075000001_ABST
    Figure 0007676075000001_ABST
Patent Text Reader

Abstract

The system generates speech for the avatar that is tailored to the patient's condition and speech content, and also generates facial expressions based on the avatar's emotions. [Solution] The content of a user's speech is obtained through dialogue with an avatar, and a health data generation model trained using input data including the speech content and output data including health data related to the user's health condition as learning data is utilized to output health data based on the acquired speech content, and the output health data is recorded in a personal health record. An avatar dialogue generation model trained using input data including health data and output data including the speech content of the avatar as learning data is utilized to read out health data from the recorded personal health record, and the speech content of the avatar is generated based on this.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an avatar-based interactive medical interview support program and system that enables an avatar to conduct an interview while having a dialogue with a user as a patient. [Background technology]

[0002] In recent years, the labor shortage at medical institutions has placed a heavy burden on medical workers. As a result, the number of cases of medical workers leaving their jobs is increasing. For this reason, in order to reduce the burden on medical workers, there have been discussions about using AI to replace some of the work at medical institutions other than actual medical procedures.

[0003] One of the tasks that can be replaced by AI in medical institutions is to ask a patient a simple question via an AI-implemented avatar before an actual examination by a doctor, and to obtain health data on the user's health condition.

[0004] For example, Patent Document 1 discloses a configuration for acquiring information from a patient in an interactive format before a medical examination. Patent Document 1 discloses a technology that accepts medical interview information created by a patient before a doctor's examination in advance, and links the medical interview data with the doctor in charge of the examination, thereby enabling the doctor to view the medical interview data created in advance by the patient when the doctor examines the patient. A user can input medical interview information such as subjective symptoms in an interactive format in advance via a tablet terminal or the like.

[0005] Patent Document 2 also discloses a configuration in which a virtual doctor (avatar) conducts a medical interview in a step related to a dialogue-style medical interview. According to the technology disclosed in Patent Document 2, a medical interview is conducted with a patient through a dialogue via voice using voice synthesis and voice recognition, and a medical interview sheet is created based on the response from the patient. This makes it possible to provide the patient with a medical interview by a so-called virtual doctor expressed by synthetic voice. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Patent No. 7450310 [Patent Document 2] Patent No. 7454090 Summary of the Invention [Problem to be solved by the invention]

[0007] However, Patent Documents 1 and 2 do not specifically disclose a technology for generating speech content of an avatar that corresponds to a patient that is more in tune with the patient's condition and speech content. In addition, Patent Documents 1 and 2 do not specifically disclose a technology for estimating the emotion of an avatar according to the patient's condition and speech content and generating the avatar's facial expression based on this.

[0008] Therefore, the present invention has been devised in consideration of the above-mentioned problems, and its purpose is to provide an avatar-based interactive medical interview support program and system that can generate avatar speech content that is tailored to the patient's condition and speech content, and further, can estimate the avatar's emotions and generate its facial expressions. [Means for solving the problem]

[0009] The interactive medical interview support program using an avatar according to the present invention is characterized in that it has a computer execute the following steps: a speech content acquisition step of acquiring utterances made by a user through an interaction with an avatar; a recording step of utilizing a health data generation model trained using input data including the utterance content and output data including health data related to the user's health condition as learning data, outputting health data based on the utterance content acquired by the speech content acquisition step, and recording the output health data in a personal health record; an analysis step of utilizing an analysis model trained using input data including the health data and output data including a response policy for the user as learning data, reading out the health data from the personal health record recorded in the recording step, and outputting a response policy based thereon; and an avatar dialogue generation step of utilizing an avatar dialogue generation model trained using input data including the response policy and output data including the avatar's utterance content as learning data, and generating the avatar's utterance content based on the response policy output in the analysis step.

[0010] The avatar-based interactive medical interview support program of the present invention is characterized in that it has a computer execute the following steps: a speech content acquisition step of acquiring utterances made by a user through an interaction with an avatar; a recording step of utilizing a health data generation model trained using input data including the utterance content and output data including health data related to the user's health condition as learning data, outputting health data based on the utterance content acquired by the speech content acquisition step, and recording the output health data in a personal health record; and an avatar dialogue generation step of utilizing an avatar dialogue generation model trained using input data including health data and output data including the avatar's utterance content as learning data, reading out the health data from the personal health record recorded in the recording step, and generating the avatar's utterance content based on the health data.

[0011] The avatar-based interactive medical interview support system of the present invention is characterized in that it comprises: a speech content acquisition means for acquiring utterances made by a user through an interaction with an avatar; a recording means for utilizing a health data generation model trained using input data including the utterance content and output data including health data related to the user's health condition as learning data, outputting health data based on the utterance content acquired by the speech content acquisition means, and recording the output health data in a personal health record; an analysis means for utilizing an analysis model trained using input data including the health data and output data including a response policy for the user as learning data, reading out the health data from the personal health record recorded in the recording step, and outputting a response policy based thereon; and an avatar dialogue generation means for utilizing an avatar dialogue generation model trained using input data including the response policy and output data including the avatar's utterance content as learning data, and generating the avatar's utterance content based on the response policy output in the analysis step.

[0012] The avatar-based interactive medical interview support system of the present invention is characterized in that it comprises an utterance content acquisition means for acquiring utterances made by a user through an interaction with an avatar, a recording means for utilizing a health data generation model trained using input data including the utterance content and output data including health data related to the user's health condition as learning data, outputting health data based on the utterance content acquired by the utterance content acquisition means, and recording the output health data in a personal health record, and an avatar dialogue generation means for utilizing an avatar dialogue generation model trained using input data including health data and output data including the avatar's utterance content as learning data, reading out the health data from the personal health record recorded in the recording means, and generating the avatar's utterance content based thereon. Effect of the Invention

[0013] According to the present invention having the above-mentioned configuration, it is possible to generate avatar speech content that is tailored to the user's medical condition and speech content, and further to generate the avatar's facial expressions after estimating its emotions, making it possible to easily find out from the user about their own physical and mental condition, which can then be stored in a personal health record. [Brief description of the drawings]

[0014] [Figure 1] FIG. 1 is a block diagram showing the overall configuration of an interactive medical interview support system to which the present invention is applied. [Diagram 2] FIG. 2 is a block diagram of an interactive medical interview support system to which the present invention is applied. [Diagram 3] FIG. 3 is a schematic diagram illustrating an example of the functions of the management server. [Figure 4] FIG. 4 is a flowchart showing the operation procedure of the dialogue-type medical interview support system. [Diagram 5] FIG. 5 is a diagram showing a data flow resulting from the use of each piece of learning data. [Figure 6] FIG. 6 is a diagram showing an example of an avatar dialogue generation model in which input data including a response policy and health data, and output data including avatar utterance content are used as learning data. [Figure 7] FIG. 7 is a diagram showing an example in which the avatar dialogue generation model serves as an analysis model. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, an interactive medical interview support system to which the present invention is applied will be described with reference to the drawings.

[0016] FIG. 1 is a block diagram showing the overall configuration of an interactive medical interview support system 100 to which the present invention is applied.

[0017] The dialogue-type medical interview support system 100 includes a management server 1. The management server 1 is connected to one or more user terminals 2 via a communication network 4, for example.

[0018] The interactive medical interview support system 100 has a personal health record 3 implemented in the management server 1 or connected externally. The management server 1 records various management information to be managed, and may also store various data or information such as information on data transmitted to and received from the user terminal 2, communication setting information, various setting information for transmitting and receiving various data to and from the user terminal 2, and rules. A large scale language model (LLM) 5 is further connected to the management server 1.

[0019] In FIG. 2, the management server 1 may be configured as, for example, a client server system or a cloud system, and includes a housing 10, a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a storage unit 104, and a plurality of I / Fs 105 to 107, and each component is connected by an internal bus 110. The CPU 101 controls the entire management server 1. The ROM 102 stores the operation code of the CPU 101. The RAM 103 is a working area used when the CPU 101 operates. The storage unit 104 stores information about the management server 1, types and conditions of various communication procedures, and the like. As the storage unit 104, for example, a data storage device (not shown) such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), a hard disk, or a semiconductor memory may be used.

[0020] The CPU 101 is realized by executing a program stored in the storage unit 104 or the like, using the RAM 103 as a working area. For example, the management server 1 may have a GPU (Graphics Processing Unit) not shown. By having a GPU, faster calculation processing than usual becomes possible.

[0021] The I / F 105 is an interface for transmitting and receiving various information connected via the communication network 4. The I / F 105 transmits and receives various information to and from the user terminal 2 via the communication network 4. The I / F 106 is an interface for transmitting and receiving information to and from the input unit 108. For example, a keyboard is used as the input unit 108, and the administrator of the management server 1 inputs control commands for transmitting and receiving various information or setting various information via the input unit 108. The I / F 107 is an interface for transmitting and receiving various information to and from the output unit 109. The output unit 109 outputs various information stored in the storage unit 104, the processing status of the management server 1, etc. A display is used as the output unit 109, and may be, for example, a touch panel type.

[0022] As the user terminal 2, for example, a known electronic device such as a personal computer (PC), a smartphone, a tablet terminal, or a 3D hologram, which is managed for each user, is used. An example of the configuration of the user terminal 2 may be the same as the schematic diagram shown in FIG. 2, and may include, for example, a housing, a CPU, a ROM, a RAM, a storage unit, each I / F, an input unit, and an output unit. Each configuration is connected by an internal bus. Note that, as another example of the user terminal 2, it may be configured with any device that can display an avatar, any electronic device that can project an avatar on a screen in two dimensions as a moving image, any electronic device that can project an avatar in 2.5 dimensions as a moving image, a device such as a character summoning device that projects or displays a favorite avatar character in three-dimensional space, or a so-called doll-type, pet-type, stuffed animal-type, or other robot, or may be configured with a self-propelled robot that serves and distributes food.

[0023] The personal health record 3 is a server capable of recording all information related to the individual body of a user. The personal health record 3 records health data related to the health condition and body of each user linked to each individual user. Examples of health data recorded in the personal health record 3 include height, weight, body temperature, blood pressure, blood sugar level, heart rate, respiratory rate, stress level, and other vital data related to the health condition that can be measured using measuring instruments in so-called health checkups and medical checkups. In addition, information on the contents of meals and nutrients taken, working hours, overtime hours, daily exercise, sleep hours, and other data on lifestyle habits are also included. Furthermore, information on the health condition expressed by the user himself or detected by oral or written recording is also included in this health data, such as the part of the body where there is pain and its degree, the type of pain, and other parts of the body where there is an abnormality and their types. The health data also includes data on chronic diseases and disabilities, data on current or past diseases, information on medicines currently being taken, information on vaccines taken in the past, and even information on allergies. Furthermore, this health data also includes information on side effects caused by certain medicines or prescriptions.

[0024] This personal health record 3 may be embodied as one function of the management server 1, or may be configured as an independent server. Since this personal health record 3 records personal information, it may be equipped with any well-known technology for ensuring security to prevent information hacking from outside.

[0025] The communication network 4 is an internet network or the like to which the management server 1 and the user terminal 2 are connected via communication circuits in the interactive medical interview support system 100. The communication network 4 may be configured as a so-called optical fiber communication network. The communication network 4 may be realized by a known public communication network such as a wireless communication network, in addition to a wired communication network.

[0026] 3 is a schematic diagram showing an example of the functions of the management server 1. The management server 1 includes at least an utterance content acquisition unit 11 that acquires the contents of a user's utterance, an avatar dialogue generation unit 12 that generates the contents of an avatar's utterance, an analysis unit 13 that analyzes health data and generates data related to a response policy, a vital data detection unit 17 that detects vital data from the user, an output unit 14 that outputs data to the user terminal 2 via the communication network 4, a control unit 15 that controls the management server 1, a recording unit 16 that stores various other information, and an expression generation unit 18 that generates new expression data.

[0027] The recording unit 16 records AI learning data such as a health data generation model, an analysis model, and an avatar dialogue generation model, which will be described later.

[0028] The avatar dialogue generation unit 12 first generates the avatar's speech content as text data. The avatar dialogue generation unit 12 may appropriately use the large-scale language model 5 to generate the text data of the avatar's speech content.

[0029] The expression generating unit 18 has known image control functions required for appropriately generating, changing and expressing the facial expression of the avatar on the image.

[0030] The management server 1 cooperates with the user terminal 2 via the communication network 4 in carrying out operations based on these functions.

[0031] Furthermore, in the present invention, each component of the management server 1, such as the speech content acquisition unit 11, the avatar dialogue generation unit 12, the analysis unit 13, the vital data detection unit 17, the recording unit 16, and the expression generation unit 18, may be implemented on the user terminal 2 side. Furthermore, all of the AI ​​learning data itself may be implemented on the user terminal 2 side, and may be embodied as a so-called edge AI.

[0032] Next, an example of the operation of the dialogue-type medical interview support system 100 in this embodiment will be described.

[0033] FIG. 4 is a flow chart showing the operation procedure of this dialogue-type medical interview support system 100, and FIG. 5 shows a data flow resulting from the use of each learning data.

[0034] First, in step S11, the contents of a user's utterance are acquired through a dialogue with an avatar. The avatar dialogue generation unit 12 generates the contents of the avatar's utterance, and the user terminal 2 converts this into voice and outputs it. Similarly, the expression generation unit 18 generates facial expressions of the avatar, and the user terminal 2 reflects this on the screen. A user viewing the user terminal 2 can experience the sensation of actually conversing with the avatar through the facial expressions of the avatar displayed on the screen and the contents of the utterance output as voice. For example, by making the avatar an image of a nurse, the user feels as if the nurse is asking about his or her physical condition.

[0035] Through dialogue with the avatar, the user utters information about his or her own physical condition. Conversely, by including in the dialogue content of the avatar a question about the user's physical condition, such as "How are you feeling today?", an atmosphere can be created in which the user can proactively talk about his or her own physical condition. Furthermore, by including more specific questions when the avatar asks about the user's physical condition, the user can obtain from the dialogue content such information as when the pain or discomfort started, which part of the body is in, what kind of pain or discomfort there is, and the degree of the pain, and the like, and the user can utter the necessary questions before the doctor actually examines the user.

[0036] In addition, through dialogue with the avatar, the user can also speak about lifestyle habits, chronic illnesses and disabilities, current or past illnesses, and even information about medications currently being administered and allergies.

[0037] The user terminal 2 converts the user's speech content into text data, temporarily stores it, and transmits it to the management server 1 via the communication network 4. On the other hand, if this system is embodied as an edge AI and each component is implemented on the user terminal 2 side, the operation of transmitting the text data of the speech content to the management server 1 via the communication network 4 may be omitted.

[0038] In the above example, the case where the user freely speaks through dialogue with the avatar and the voice is acquired through the user terminal 2 has been described as an example, but the present invention is not limited to this. When acquiring information such as the user's medical condition, physical condition, lifestyle habits, and other chronic illnesses from the user, the user may, in addition to voice, input information via a touch panel, buttons, etc. from options displayed on the user interface of the user terminal 2 and use this as the content of the speech, or may freely input text via a touch panel, buttons, etc. as text data and use this as the content of the speech.

[0039] Furthermore, in step S11, in addition to acquiring the contents of the user's speech, detected vital data from the user may be acquired. The vital data here includes a variety of data that can be measured from the body using measuring instruments, such as body temperature, height, weight, blood pressure, blood glucose level, heart rate, respiratory rate, pulse rate, blood oxygen concentration, etc. This vital data may be acquired by directly importing data detected from an actual measuring instrument, or by having the user manually input previously measured data into the user terminal 2.

[0040] In addition, in this step S11, the facial expression of the user who is interacting with the avatar may be captured by a camera provided in the user terminal 2.

[0041] Also, in step S11, voice data of the user interacting with the avatar may be detected by a microphone provided in the user terminal 2. The voice data referred to here is not the content of the speech but includes the tone of voice, voiceprint, volume of voice, accent, etc.

[0042] Next, the process proceeds to step S12, where the health data generation model is used to output health data.

[0043] The health data generation model is a trained model trained using input data including speech content and output data including health data related to the user's health condition as training data, as shown in Fig. 4. This health data generation model is formed based on machine learning by artificial intelligence, but is not limited to this, and the output data may be formed using a large-scale language model 5.

[0044] The health data generation model may be generated, for example, using machine learning modeled on a neural network. The health data generation model may be a model using multimodality. The health data generation model may be trained using machine learning modeled on a neural network such as a CNN (Convolution Neural Network), or any other model may be used. The health data generation model may be generated, for example, using Retrieval-Augmented Generation (RAG), Sequence To Sequence (seq2seq) linear discrimination, support vector machine, k-nearest neighbor method, random forest, deep learning, or the like.

[0045] Through a health data generation model that has learned such input data and output data, health data can be obtained from text data of user utterances. For example, if the text data of the utterance is "Today I have a pain near my stomach", corresponding health data such as "stomachache", "hyperacidity", and "stomach discomfort" can be obtained. Even if the text data of the utterance is an ambiguous expression such as "I feel dizzy", health data such as "headache", "autonomic imbalance", and "dizziness" can be obtained.

[0046] Furthermore, the input data of the health data generation model may include vital data in addition to the speech content. Vital data is data that can be measured from the body using a measuring device as described above, and by learning this data together with the health data that is the output of the health data generation model, it is possible to output the detected vital data as more organized health data.

[0047] Furthermore, the input data of the health data generation model may include the user's facial expression in addition to the speech content. By learning the facial expression image data and the health data output from the health data generation model as a set, it is possible to output the captured image data of the user's facial expression as more organized health data. In particular, since the health condition of the user may change depending on the facial expression, etc., by learning these relationships and constructing the health data generation model in advance, it is possible to generate health data from facial expressions.

[0048] Furthermore, the input data of the health data generation model may include voice data in addition to the contents of the speech. By learning the voice data and the health data, which is the output of the health data generation model, as a set, it is possible to output the detected voice data as more organized health data. In particular, since the health condition may change depending on the tone of the voice, etc., by learning these relationships and constructing the health data generation model in advance, it is possible to generate health data from the voice data.

[0049] Next, the process proceeds to step S13, where the health data output from the above-mentioned health data generation model is linked to the user and recorded in the above-mentioned personal health record 3.

[0050] Next, the process proceeds to step S14, where the analysis model is utilized to output a response policy for the user.

[0051] The analytical model is a model trained using input data including health data and output data including a response policy for the user as training data, as shown in Fig. 4. This analytical model is formed based on machine learning by artificial intelligence, but is not limited to this, and the output data may be formed using a large-scale language model 5.

[0052] The analytical model may be generated, for example, by machine learning using a neural network as a model. The analytical model may be a multimodal model. The analytical model may be trained using machine learning using a neural network such as CNN, or any other model may be used. The analytical model may be generated, for example, by search extension generation, seq2seq linear discrimination, support vector machine, k-nearest neighbor method, random forest, deep learning, or the like.

[0053] The data on the response policy output from this analysis model is composed of data patterns on how to respond to users, such as "User A with XX characteristics should be approached in the following way", "User B with △△ characteristics should be sympathized with in the following way through dialogue", "User C with □□ characteristics should be given advice in the following way", "User D with ×× characteristics should be asked further questions about ~", etc. This response policy is generated based on health data, so in addition to physical and mental care and further questions, it may also include specific medical advice and treatment policies. Incidentally, in this data on the response policy, the information on the user's characteristics may be additional, or the information may be omitted and the data may simply be composed of information on what kind of response should be made.

[0054] The health data can be analyzed through an analytical model that has been trained on such input data and output data, and as a result, data on a response policy can be obtained from the health data for each individual user. When actually using this analytical model to generate a search solution, health data linked to the user who is the counterpart of the avatar is read from the personal health record 3 and input to the analytical model. As a result, a response policy is output as a search solution of the analytical model.

[0055] Next, the process proceeds to step S15, where the avatar dialogue generation model is used to generate the avatar's utterance contents.

[0056] The avatar dialogue generation model is a model trained using input data including a response policy for a user and output data including avatar utterances as training data, as shown in Fig. 4. This avatar dialogue generation model is formed based on machine learning using artificial intelligence, but is not limited to this, and the output data may be formed using a large-scale language model 5.

[0057] The avatar dialogue generation model may be generated, for example, by using machine learning modeled on a neural network. The avatar dialogue generation model may be a model using multimodality. The avatar dialogue generation model may be trained, for example, by using machine learning modeled on a neural network such as CNN, or any other model may be used. The avatar dialogue generation model may be generated, for example, by using search extension generation, seq2seq linear discrimination, support vector machine, k-nearest neighbor method, random forest, deep learning, or the like.

[0058] The avatar's speech content output from the avatar dialogue generation model includes specific questions about the user's health condition or illness, specific words of encouragement, empathy, sympathy, compassion, and emotional care for the user, etc. In addition to the speech content for medical diagnosis, interviews, and grasping the user's condition, the speech content may also include words that cheer up the user's mood and relax them.

[0059] Using an avatar dialogue generation model trained on such input and output data, it is possible to generate specific avatar speech content based on a response policy.

[0060] In step S15, the avatar's utterance contents may be generated through the avatar dialogue generation model, and the avatar's facial expressions may also be generated. In this case, an avatar dialogue generation model is used that is trained using input data including the response policy and output data including the avatar's utterance contents and the avatar's emotions as training data. The avatar dialogue generation model outputs the avatar's emotions in addition to the avatar's utterance contents. By linking each of the avatar's emotions to the avatar's facial expressions, it is possible to generate an avatar's facial expression on an image according to the output avatar's emotions.

[0061] Finally, the process proceeds to step S16, where the avatar speaks to the user through the user terminal 2 based on the speech content generated through the avatar dialogue generation model. At this time, if the emotion of the avatar is output from the avatar dialogue generation model, the corresponding facial expression of the avatar can be displayed as an image.

[0062] In this way, according to the present invention, the content of the user's utterance is converted into health data through the health data generation model, an appropriate response policy is searched for through the analysis model, and specific content of the utterance is generated through the avatar dialogue generation model. This allows the avatar to speak according to the health condition and state analyzed from the content of the user's utterance. Even if the user is worried about their health condition or suffers from an illness or injury, they can relieve their anxiety through conversation with the avatar, receive appropriate treatment or response, and even receive mental support by receiving mental care.

[0063] In such a case, if the avatar's emotions are further output from the avatar dialogue generation model, the corresponding facial expression of the avatar can be displayed as an image, so that the avatar's facial expression corresponds to the user's situation, thereby conveying empathy and sympathy to the user and making the impression of a caring response.

[0064] In the present invention, as shown in Fig. 6, the avatar dialogue generation model may be a model trained using input data including a response policy and health data, and output data including the avatar's utterance as training data. The avatar's utterance can be generated based on both the health data and the response policy. In this case, by training the avatar's utterance and output data including the avatar's emotions, facial expressions based on the avatar's emotions can also be generated.

[0065] Furthermore, according to the present invention, as shown in Fig. 7, the avatar dialogue generation model may be made to play the role of an analysis model. In such a case, an avatar dialogue generation model trained using input data including health data and output data including avatar utterances as training data is utilized. By reading out health data recorded in a personal health record and training the avatar utterances corresponding to the health data, it is possible to generate the avatar utterances directly based on the health data. At this time, by training the avatar utterances and output data including the avatar's emotions, it is also possible to generate facial expressions based on the avatar's emotions. [Explanation of symbols]

[0066] 1 Management Server 2. User terminal 3. Personal Health Record 4. Communication Network 5. Large-scale language models 10. Chassis 11 Speech content acquisition unit 12 Avatar dialogue generation unit 13 Analysis Department 14 Output section 15 Control section 16 Recording Section 17 Vital Data Detector 18 Expression generator 100 Interactive medical interview support system 101 CPU 102 ROM 103 RAM 104 Preservation Department 105~107 Interface 108 Input section 109 Output section 110 Internal Bus

Claims

1. a speech content acquisition step of acquiring a speech content by a user through a dialogue with an avatar; a recording step of outputting health data based on the utterance content acquired in the utterance content acquisition step by using a health data generation model trained using input data including the utterance content and output data including health data related to the user's health condition as training data, and recording the output health data in a personal health record; an analysis step of reading the health data from the personal health record recorded in the recording step, using an analysis model trained using input data including health data and output data including a treatment policy for the user as training data, and outputting a treatment policy based on the health data; and executing an avatar dialogue generation step of generating the avatar's utterance contents based on the response policy outputted in the analysis step by using an avatar dialogue generation model trained using input data including the response policy and output data including the avatar's utterance contents as training data. An avatar-based interactive medical interview support program that features:

2. a speech content acquisition step of acquiring a speech content by a user through a dialogue with an avatar; a recording step of outputting health data based on the utterance content acquired in the utterance content acquisition step by using a health data generation model trained using input data including the utterance content and output data including health data related to the user's health condition as training data, and recording the output health data in a personal health record; and executing an avatar dialogue generating step of using an avatar dialogue generation model trained using input data including health data and output data including the speech content of the avatar as training data, reading out the health data from the personal health record recorded in the recording step, and generating the speech content of the avatar based on the health data. An avatar-based interactive medical interview support program that features:

3. In the avatar dialogue generation step, an avatar dialogue generation model trained using input data including a response policy and output data including the avatar's speech content and the avatar's emotion as training data is utilized, and a facial expression according to the avatar's emotion is generated together with the avatar's speech content based on the response policy output in the analysis step.

2. The avatar-based interactive medical interview support program according to claim 1 .

4. The method further includes a vital data detection step of detecting vital data from the user, In the recording step, a health data generation model trained using input data including the speech content and the vital data, and output data including health data related to the user's health condition as training data is utilized, and the health data is output based on the vital data detected in the vital data detection step.

3. The avatar-based interactive medical interview support program according to claim 1 or 2.

5. The method further includes a facial expression detection step of detecting a facial expression of the user, In the recording step, a health data generation model trained using input data including the speech content and the facial expression and output data including health data related to the user's health condition as training data is utilized, and the health data is output based on the facial expression detected in the facial expression detection step.

3. The avatar-based interactive medical interview support program according to claim 1 or 2.

6. The method further includes a voice data detection step of detecting voice data of the user, In the recording step, a health data generation model trained using input data including the speech content and the voice data, and output data including health data related to the user's health condition as training data is utilized, and the health data is output based on the voice data detected in the voice data detection step.

3. The avatar-based interactive medical interview support program according to claim 1 or 2.

7. A speech content acquisition means for acquiring a speech content by a user through a dialogue with an avatar; a recording means for utilizing a health data generation model trained using input data including the utterance content and output data including health data related to the user's health condition as training data, outputting health data based on the utterance content acquired by the utterance content acquisition means, and recording the output health data in a personal health record; an analysis means for utilizing an analysis model trained using input data including health data and output data including a treatment policy for the user as training data, reading out the health data from the personal health record recorded in the recording step, and outputting a treatment policy based on the health data; and an avatar dialogue generation means for generating the avatar's utterance contents based on the response policy output in the analysis step by utilizing an avatar dialogue generation model trained using input data including the response policy and output data including the avatar's utterance contents as training data. This is an avatar-based interactive medical interview support system.

8. A speech content acquisition means for acquiring a speech content by a user through a dialogue with an avatar; a recording means for utilizing a health data generation model trained using input data including the utterance content and output data including health data related to the user's health condition as training data, outputting health data based on the utterance content acquired by the utterance content acquisition means, and recording the output health data in a personal health record; and an avatar dialogue generation means for utilizing an avatar dialogue generation model trained using input data including health data and output data including the speech content of the avatar as training data, reading out the health data from the personal health record recorded in the recording means, and generating the speech content of the avatar based on the health data. This is an avatar-based interactive medical interview support system.

Citation Information

Patent Citations

  • Health examination method, health examination system, health examination device and storage medium to be used therefor

    JP2002230159A

  • Healthcare apparatus and program for driving the same to function

    JP2008011865A

  • Information processing method, diagnostic support device, and computer program

    JP2022037803A

  • Medical interview support system

    JP7450310B1

  • Medical Support Devices

    JP7454090B1

Cited By

  • Information processing system and information processing method

    JP7764082B1