Medical interview apparatus and control program for medical interview apparatus

The medical interview device enhances non-verbal communication skills by dynamically changing patient facial expressions based on examiner questions, addressing the limitations of conventional simulators and reducing costs.

JP2026021985APending Publication Date: 2026-02-12平田 創一郎
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123290
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Conventional medical examination simulators fail to provide realistic non-verbal communication skills, as they display a fixed patient facial expression, making it difficult for examiners to practice appropriate responses and build trust with patients.

Method used

A medical interview device that utilizes machine learning to dynamically change patient facial expressions based on the content of the examiner's questions, incorporating voice recognition and image generation to simulate realistic patient responses and evaluate the examiner's performance.

Benefits of technology

Enables examiners to fully acquire non-verbal communication skills necessary for medical interviews, reducing the need for trained patients and financial burden while providing consistent responses and evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021985000001_ABST
    Figure 2026021985000001_ABST
Patent Text Reader

Abstract

To provide a medical interview device and a control program of the medical interview device, allowing sufficient learning of a skill necessary for a medical interview.SOLUTION: The medical interview apparatus 1 is a medical interview apparatus for providing a patient's answer to an examiner's question. The medical interview apparatus 1 includes the voice receiving unit 152 that receives the input of the voice asked by the examiner, the question specifying unit 153 that specifies the question asked by the examiner based on the voice received by the voice receiving unit 152 by using the first learned model created by the machine learning, and the answer output unit 154 that outputs the answer of the patient to the question specified by the question specifying unit 153 and the face image of the patient. The facial expression of the face image output by the answer output unit 154 changes according to the content of the question asked by the examiner.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a medical interview device and a control program for the medical interview device. More particularly, the present invention relates to a medical interview device and a control program for the medical interview device that enable a user to fully acquire the skills required for a medical interview. [Background technology]

[0002] In Japan, there is a common examination that medical and dental students take before they begin clinical training. In Japan, the law allows only those who hold national medical or dental licenses to practice medicine exclusively. This is called a monopoly on medical or dental practice. However, as an exception to the monopoly on medical or dental practice, students who pass this common examination are allowed to practice medicine or dentistry as clinical training. Passing the common examination is also one of the qualifications for taking the national medical and dental examinations.

[0003] The common examination consists of CBT (Computer Based Testing), which tests knowledge, and OSCE (Objective Structured Clinical Examination), which tests skills and attitudes. One of the tasks in the OSCE is a medical interview.

[0004] A medical interview is what was previously called a medical interview. The medical interview has three objectives: to build a relationship of trust (rapport) between the patient and the medical professional, to gather information for the evaluation of the patient's health problems, and to educate the patient and motivate them to treatment. The OSCE evaluates the skills needed to achieve these three objectives, as well as the patient's attitude toward the patient. As mentioned above, the medical interview is necessary for candidates taking the national medical or dental examination to complete the OSCE tasks, and it is also an essential skill for doctors and dentists after obtaining national qualifications.

[0005] Generally, when students or examiners such as doctors or dentists practice to acquire medical interview skills, they need someone to play the role of a patient. The patient role can be played not only by students or teachers, but also by a person who has undergone specific training called a simulated patient. In particular, in OSCEs, it is necessary to accurately evaluate the examinee's medical interview, so the patient role is played by a standardized simulated patient who has been trained to standardize the patient's responses.

[0006] As mentioned above, because it is necessary to undergo specific training to become a simulated patient, it was difficult to secure simulated patients when practicing medical interviews. Furthermore, if it was necessary to pay the simulated patients, the financial burden on the examiner was significant. Furthermore, even with trained simulated patients, it was difficult for multiple simulated patients to give uniform answers based on the scenario to the examiner's questions.

[0007] Therefore, in order to solve the problems of securing mock patients and the financial burden on the examiner, and to always obtain uniform answers to the examiner's questions, examination simulators that simulate medical interviews have been proposed.

[0008] For example, the medical examination simulator disclosed in Patent Document 1 below includes a case database storing healthy subject data representing the condition of healthy subjects and case data representing patient cases, a presentation unit that presents to the subject, based on the case data, selection items for the subject related to the examination and examination information obtained from the patient for the selection items, using at least either image information or audio information, and a control unit that controls the operation of the presentation unit. This medical examination simulator displays selection items related to the subject's actions together with an image of the patient on a head-mounted display (HMD) worn by the subject of the medical examination simulator, and accepts the selection of an action related to the examination from the displayed selection items. [Prior art documents] [Patent documents]

[0009] [Patent Document 1] Japanese Patent Application Publication No. 2023-183472 Summary of the Invention [Problem to be solved by the invention]

[0010] One of the skills required for a medical interview is non-verbal communication, such as nodding and responding. Non-verbal communication skills are necessary to obtain appropriate information from patients who are in a state of nervousness when meeting the examiner for the first time, and to build a relationship of trust with the patient.

[0011] However, in conventional medical examination simulators such as the medical examination simulator of Patent Document 1, the facial expression of the patient displayed on a display unit such as a head-mounted display is always the same. When a medical interview simulation is performed using a conventional medical examination simulator, the examiner is unable to practice displaying an attitude corresponding to the patient's facial expression, and is therefore unable to fully acquire non-verbal communication skills. This makes it difficult for the examiner to fully acquire the skills necessary for a medical interview.

[0012] The present invention has been made to solve the above-mentioned problems, and its object is to provide a medical interview device and a control program for the medical interview device that enable a person to fully acquire the skills necessary for medical interviews. [Means for solving the problem]

[0013] A medical interview device according to one aspect of the present invention is a medical interview device that provides a patient's answer to a question asked by an examiner, and includes: a voice receiving means that receives input of the voice of the question asked by the examiner; a question identification means that uses a first trained model created by machine learning to identify the question asked by the examiner based on the voice input received by the voice receiving means; and an answer output means that outputs the patient's answer to the question identified by the question identification means and an image of the patient's face, and the expression of the facial image output by the answer output means changes depending on the content of the question asked by the examiner.

[0014] Preferably, the above-mentioned medical interview device further comprises a storage means for storing question-and-answer information including each of a plurality of mutually different example questions, each of a plurality of example answers to each of the plurality of example questions, and each of a plurality of example facial images of the patient when each of the plurality of example answers is output, and the answer output means outputs the example answers and example facial images associated with the questions identified by the question identification means in the question-and-answer information.

[0015] In the above-mentioned medical interview device, preferably, the storage means stores a plurality of pieces of question-and-answer information corresponding to each of the plurality of scenarios, the example answers associated with one example question common to each of the plurality of pieces of question-and-answer information among the plurality of example questions are different for each of the plurality of pieces of question-and-answer information, and the answer output means outputs the example answers and example facial images associated with the questions identified by the question identification means in the question-and-answer information corresponding to a specific scenario among the plurality of scenarios.

[0016] In the above-mentioned medical interview device, preferably, the answer output means uses a second trained model created by machine learning to input the question identified by the question identification means and outputs the answer and face image of the patient.

[0017] The above-mentioned medical interview device preferably further comprises a moving image creating means for creating a composite moving image by combining a moving image of the patient's face with the voice of the patient's answer, and the answer output means outputs the composite moving image.

[0018] The medical interview device preferably further comprises an imaging means for imaging the examiner asking questions, and an evaluation means for evaluating the medical interview conducted by the examiner based on the voice input received by the voice receiving means and the photographed object captured by the imaging means.

[0019] In the above-mentioned medical interview device, preferably, the facial expression of the facial image output by the answer output means changes at least between a smiling expression and a distressed expression according to the content of the question asked by the examiner.

[0020] In the above-mentioned medical interview device, preferably, when the voice input received by the voice receiving means is interrupted for a predetermined period of time, the question identifying means identifies the question asked by the examiner by treating the voice input received by the voice receiving means before the interruption as one question.

[0021] A control program for a medical interview device according to another aspect of the present invention is a control program for a medical interview device that provides a patient's answer to a question asked by an examiner, and causes a computer to execute a voice receiving step of receiving input of the voice of the question asked by the examiner, a question identification step of identifying the question asked by the examiner based on the voice input received in the voice receiving step using a first trained model created by machine learning, and an answer output step of outputting the patient's answer to the question identified in the question identification step and an image of the patient's face, and the expression of the facial image output in the answer output step changes depending on the content of the question asked by the examiner. [Effects of the Invention]

[0022] According to the present invention, it is possible to provide a medical interview device and a control program for the medical interview device that enable a person to fully acquire the skills necessary for medical interviews. [Brief explanation of the drawings]

[0023] [Figure 1] 1A and 1B are diagrams showing the appearance and basic usage of a medical interview device 1 according to an embodiment of the present invention. [Figure 2] 1 is a block diagram showing the hardware configuration of a medical interview device 1 according to an embodiment of the present invention. [Figure 3] 1 is a block diagram showing the functional configuration of a medical interview device 1 according to an embodiment of the present invention. [Figure 4] 1 is a diagram schematically illustrating a neural network 500 of each of models 121, 122, and 123 in one embodiment of the present invention. [Figure 5] FIG. 10 is a diagram showing an example of learning data SD1 for creating a model 121. [Figure 6]FIG. 2 is a diagram showing a schematic diagram of question and answer information 124 stored in the medical interview device 1. [Figure 7] 10A to 10C are diagrams illustrating variations of example patient facial images. [Figure 8] FIG. 10 is a diagram showing a modified example of question and answer information 124 stored in the medical interview device 1. [Figure 9] FIG. 10 is a diagram showing an example of learning data SD2 for creating a model 122. [Figure 10] FIG. 10 is a diagram showing an example of learning data SD3 for creating a model 123. [Figure 11] FIG. 1 is a diagram illustrating non-verbal communication. [Figure 12] 3 is a flowchart showing the operation of the medical interview device 1 according to the embodiment of the present invention. [Figure 13] FIG. 1 is a diagram conceptually showing the configuration of a first modified example of the medical interview device 1 according to an embodiment of the present invention. [Figure 14] FIG. 10 is a diagram conceptually showing the configuration of a second modified example of the medical interview device 1 according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0024] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0025] (Configuration of medical interview device)

[0026] FIG. 1 shows the appearance and basic usage of a medical interview device 1 according to an embodiment of the present invention.

[0027] Referring to FIG. 1, the medical interview device 1 (an example of a medical interview device) in this embodiment is a device used for practicing medical interviews, particularly for initial dental interviews. The medical interview device 1 provides the patient PT's answers to questions from the examiner DC. The patient PT is a virtual patient, not a real person. The medical interview device 1 receives voice input from the examiner DC via a microphone 104. This voice is a question posed to the patient PT. The medical interview device 1 outputs the patient PT's answers to the questions asked by the examiner DC and an image of the patient PT's face. The patient PT's answers are output as voice via a speaker 105, or displayed as text on a display unit 103 together with an image of the patient PT's face.

[0028] The facial expression of the patient PT's facial image changes depending on the content of the question posed by the examiner DC. A change in the facial expression of the patient PT means that the components of the face change in addition to the movement of the patient PT's mouth when answering. Specifically, this means changes in the movement of the eyebrows, the distance between the eyebrows, the size of the eyes, or the vertical position of the corners of the mouth. These changes in components cause the facial expression of the patient PT to change between a normal expression, a smiling expression, a pained expression, or a sad expression. Generally, a happy expression involves upturned corners of the mouth, drooping corners of the eyes, and upturned cheeks. A pained expression involves downturned corners of the mouth, closed eyes, close eyebrows, and an open mouth. It is preferable that the facial expression of the patient PT change at least between a smiling expression and a pained expression. The patient PT's attitude (body movements, gestures, voice volume, etc.) may also change depending on the content of the question posed by the examiner DC. In particular, as a change in the gestures of the patient PT, the patient PT may nod or press a painful area with his / her hand in response to the content of the question asked by the examiner DC.

[0029] The medical interview device 1 may also evaluate the medical interview conducted by the examiner. In this case, the medical interview device 1 captures an image of the examiner asking questions to the patient PT via the camera 106. The medical interview device 1 evaluates the medical interview conducted by the examiner DC based on the voice of the examiner DC whose input has been accepted and the image of the examiner DC.

[0030] It is preferable that the medical interview device 1 is not connected to other devices via the Internet or a network (it is a stand-alone device), which reduces the risk of other devices acquiring various data held by the medical interview device 1 and prevents the leakage of medical interview scenarios, evaluation criteria, etc.

[0031] FIG. 2 is a block diagram showing the hardware configuration of the medical interview device 1 according to one embodiment of the present invention.

[0032] 2, the medical interview device 1 is composed of a computer such as a PC (Personal Computer) or a smartphone, and includes a control unit 101, an operation unit 102, a display unit 103, a microphone 104, a speaker 105, a camera 106, a network interface 107, and a storage device 108. The control unit 101, the operation unit 102, the display unit 103, the microphone 104, the speaker 105, the camera 106, the network interface 107, and the storage device 108 are all interconnected via a bus or the like.

[0033] The control unit 101 includes a CPU (Central Processing Unit) 111, a ROM (Read Only Memory) 112, and a RAM (Random Access Memory) 113. The CPU 111 controls the entire device. The CPU 111 executes a control program stored in the ROM 112. This allows the various functions of the medical interview device 1 to be realized.

[0034] The ROM 112 is, for example, a flash ROM. The ROM 112 stores various programs executed by the CPU 111 and various fixed data. The ROM 112 may be non-rewritable.

[0035] The RAM 113 is the main memory of the CPU 111. The RAM 113 is used to temporarily store data and the like required when the CPU 111 executes a control program.

[0036] The operation unit 102 is made up of a keyboard, a mouse, etc., and accepts various inputs. However, if the medical interview device 1 is made up of a smartphone or the like, a software keyboard may be used as the keyboard, and the mouse may be omitted.

[0037] The display unit 103 is made up of a monitor or the like and displays various information.

[0038] The microphone 104 collects sound and converts it into an electrical signal.

[0039] The speaker 105 converts the electrical signal into sound and outputs it.

[0040] The camera 106 records an image of the object.

[0041] The network interface 107 communicates with other devices according to instructions from the control unit 101 using a communication protocol such as TCP / IP.

[0042] The storage device 108 (an example of a storage means) is composed of, for example, a hard disk drive (HDD) or a solid state drive (SSD), and stores programs and various data. The storage device 108 stores models 121 (an example of a first trained model), 122 (an example of a second trained model), and 123, question and answer information 124, image data 125, and standard data 126. The standard data 126 is standard data of a patient's facial image and voice that is output at the start and end of a medical interview.

[0043] Each of the models 121, 122, and 123 is a trained model created by machine learning. In this embodiment, each of the models 121, 122, and 123 is created by supervised learning and includes a neural network and parameters.

[0044] FIG. 3 is a block diagram showing the functional configuration of the medical interview device 1 according to one embodiment of the present invention.

[0045] 2 and 3, the medical interview device 1 includes an operation reception unit 151, a voice reception unit 152 (an example of a voice reception means), a question identification unit 153 (an example of a question identification means), an answer output unit 154 (an example of an answer output means), a video creation unit 155 (an example of a video creation means), a photographing unit 156 (an example of a photographing means), and an evaluation unit 157 (an example of an evaluation means).

[0046] The operation receiving unit 151 receives various operations via the operation unit 102 or the like, such as an operation to start a medical interview, an operation to end a medical interview, or an operation to select a specific scenario from among a plurality of scenarios.

[0047] The voice receiving unit 152 receives voice input of a question uttered by the examiner via the microphone 104 .

[0048] The question identification unit 153 uses the model 121 to identify the question asked by the examiner based on the voice input received by the voice receiving unit 152 .

[0049] The answer output unit 154 outputs the patient's answer to the question identified by the question identification unit 153 and an image of the patient's face. The answer output unit 154 may input the question identified by the question identification unit 153 to the model 122 and obtain the patient's answer to the question and an image of the patient's face from the model 122.

[0050] The moving image creating unit 155 creates a composite moving image by combining a moving image of the patient's face with the voice of the patient's answer.

[0051] The photographing unit 156 uses the camera 106 to photograph the examiner asking the question.

[0052] The evaluation unit 157 evaluates the medical interview conducted by the examiner based on the voice received by the voice receiving unit 152 and the photographed object captured by the photographing unit 156. The evaluation unit 157 may input the voice received by the voice receiving unit 152 and the photographed object captured by the photographing unit 156 to the model 123, and obtain an evaluation of the medical interview from the model 123.

[0053] (Model and Q&A information)

[0054] FIG. 4 is a diagram schematically illustrating a neural network 500 for each of the models 121, 122, and 123 in one embodiment of the present invention.

[0055] Referring to Figure 4, the neural network 500 of each of the models 121, 122, and 123 (Figure 2) is a so-called hierarchical neural network, in which a large number of artificial neurons (represented by circles in Figure 4) are connected to form a hierarchy. The hierarchical neural network includes input artificial neurons, processing artificial neurons, and output artificial neurons. The neural network 500 is preferably a convolutional neural network such as an "Efficient Net."

[0056] Problem data 510 is the target of processing by the neural network 500. The problem data 510 is acquired by input artificial neurons in the input layer 501. The input artificial neurons are arranged in parallel to form the input layer 501. The problem data 510 is distributed to the processing artificial neurons.

[0057] The processing artificial neurons are connected to the input artificial neurons. The processing artificial neurons are arranged in parallel to form the hidden layer 502. The hidden layer 502 may have multiple layers. A neural network with three or more layers and hidden layers 502 is called a deep neural network.

[0058] The neural network may be a so-called convolutional neural network, which is a deep neural network consisting of alternating convolutional layers and pooling layers.

[0059] The output artificial neuron outputs training data 511 to the outside. The output artificial neuron constitutes the output layer 503. The neural network 500 is trained so that training data 511 is output when problem data 510 is input (specifically, the parameters of the models 121, 122, and 123 are adjusted).

[0060] FIG. 5 is a diagram showing an example of the training data SD1 for creating the model 121.

[0061] Referring to FIG. 5, model 121 is a trained model created by machine learning using training data SD1. The training data SD1 includes a plurality of pairs of an examiner's voice example and a question example. The examiner's voice example is problem data, and the question example is training data. Each of the examiner's voice example and the question example may be text data or voice data. Through machine learning using training data SD1, model 121, when input with the examiner's voice, outputs a question that corresponds to the voice.

[0062] In the training data SD1, one example question is associated with multiple similar voice examples of the examinee. Specifically, one example question, "What's wrong?", is associated with three similar voice examples, "What's wrong?", "How are you feeling?", and "What are your symptoms?"

[0063] When a doctor asks a patient a question, there are often many different ways to express the question, even if the content of the question is the same. Even if the way the doctor's question is expressed differs from the example questions contained in the question-and-answer data described below, by using the learning data SD1, the doctor's question can be converted into the example questions contained in the question-and-answer data.

[0064] Fig. 6 is a diagram showing a schematic diagram of the question and answer information 124 stored in the medical interview device 1. Fig. 7 is a diagram showing variations of example patient face images.

[0065] 6 and 7, the question and answer information 124 includes a plurality of example questions each different from one another, a plurality of example answers to each of the plurality of example questions, and a plurality of example facial images of the patient when each of the plurality of example answers is output. Each of the plurality of example questions included in the question and answer information 124 is the same as each of the plurality of example questions included in the training data SD1 (FIG. 5). The example facial images of the patient may be image data itself or information on the storage location of the image data (such as a uniform resource locator (URL)). In this embodiment, the example facial images of the patient are stored in the storage device 108 (FIG. 2) as image data 125 (FIG. 2). Each of the example questions and example answers may be text data or audio data. The example facial images of the patient may be a still image or a video. The question identification unit 153 (FIG. 3) identifies the question asked by the examiner from the plurality of example questions included in the question and answer information 124.

[0066] Specifically, the example question "What's wrong?", the example answer "My upper right molar is hurting," and the example patient face image "Image 2 (painful)" (Fig. 7(b)) are all associated with each other. The example question "What kind of illness do you think it is?", the example answer "I think it's a cavity," and the example patient face image "Image 1 (normal)" (Fig. 7(a)) are all associated with each other. The example question "Have you ever been to the dentist?", the example answer "Yes," and the example patient face image "Image 3 (smiling)" (Fig. 7(c)) are all associated with each other. "Image 1 (normal)" is an image of a normal facial expression, without any emotions such as joy, anger, sadness, or happiness. "Image 2 (painful)" is an image of a painful facial expression. "Image 3 (smiling)" is an image of a smiling facial expression.

[0067] In this way, in the question and answer information, a plurality of example questions are associated with a plurality of different example facial images, so that the facial expression of the patient's facial image changes depending on the content of the question asked by the examiner.

[0068] In "Image 3 (Smiling)" shown in Figure 7(c), the patient is smiling and touching his chin with his hand. In this way, the patient's facial expression and the patient's movements may change depending on the questions asked by the examiner. This makes the patient's attitude during the medical interview more natural, making the medical interview practice more effective.

[0069] FIG. 8 is a diagram showing a modified example of the question and answer information 124 stored in the medical interview device 1. In FIG.

[0070] 8, the storage device 108 (FIG. 2) may store a plurality of pieces of question-and-answer information 1241, 1242, 1243, and 1244, each corresponding to a plurality of different scenarios. A scenario is a patient's symptoms, attributes, medical history, or overall condition assumed in a medical interview practice. The scenarios assumed in each of the plurality of pieces of question-and-answer information 1241, 1242, 1243, and 1244 are different from each other. Therefore, when one example question (e.g., the example question "What's wrong?") common to each of the plurality of pieces of question-and-answer information 1241, 1242, 1243, and 1244 is extracted, the example answers associated with the example question are different from each other in each of the plurality of pieces of question-and-answer information 1241, 1242, 1243, and 1244.

[0071] At the start of the interview, one piece of question-and-answer information 124 from the plurality of pieces of question-and-answer information 1241, 1242, 1243, and 1244 may be selected by the administrator of the medical interview device 1 or may be selected by the medical interview device 1 itself. The medical interview device 1 may output example answers and example facial images associated with questions identified in the question-and-answer information 124 corresponding to a particular scenario selected from the plurality of scenarios.

[0072] FIG. 9 is a diagram showing an example of the training data SD2 for creating the model 122.

[0073] 9, the medical interview device 1 may store a model 122 instead of storing the question-and-answer information 124. The model 122 is a trained model created by machine learning using the question-and-answer information 124 as training data SD2. In this case, the example questions in the question-and-answer information 124 become problem data, and the example answers and example facial images in the question-and-answer information 124 become training data. Through machine learning using the question-and-answer information 124, when a question identified based on the examiner's voice is input, the model 122 outputs the patient's answer to the question and a facial image of the patient.

[0074] When a question identified based on the examiner's voice is input, the model 122 may create and output a composite video by combining a video of the patient's face with the patient's voice response. In this case, the response output unit 154 (FIG. 3) may output this composite video. The model 122 may be a system in which a virtual human equipped with artificial intelligence communicates using facial expressions and gestures (a so-called AI (Artificial Intelligence) talking head). In this case, the model 122 may be made up of multiple different learning models depending on scenarios such as the patient's condition.

[0075] FIG. 10 is a diagram showing an example of the training data SD3 for creating the model 123.

[0076] Referring to Fig. 10, in addition to providing the patient's answers to the examiner's questions, the medical interview device 1 may also evaluate the medical interview conducted by the examiner. The evaluation of the medical interview may be performed using a model 123 based on the voice of the examiner asking the question and a video of the examiner asking the question. The voice of the examiner asking the question is acquired through a microphone 104 (Fig. 2). The video of the examiner asking the question is captured using a camera 106 (Fig. 2).

[0077] Model 123 is a trained model created by machine learning using training data SD3. Training data SD3 includes multiple pairs of footage of the examiner (this footage includes audio of the medical professional asking questions) and an evaluation of the medical interview. The footage of the examiner asking the questions serves as problem data, and the evaluation of the medical interview serves as training data. Through machine learning using training data SD3, model 123 outputs an evaluation of the medical interview when the audio of the examiner asking the questions and the footage of the examiner asking the questions are input.

[0078] As the photographed images of the examiner asking the questions that serve as problem data, a plurality of photographed images may be prepared, each of which differs from the other in the content of the questions asked by the examiner, the way of speaking, the attitude of the examiner asking the questions, etc. Specifically, a plurality of photographed images may be prepared, each of which differs from the other in the content of the questions such as general environmental information (patient's name, sex, age, chief complaint, etc.), history of present illness (site, past symptoms, medical treatment behavior, medication, etc.), and medical history (previous illnesses, progress, etc.), the way of speaking (whether or not technical terms are used, honorific language, volume, speed, tone, intonation of voice), or the attitude of the examiner asking the questions (eye contact, facial expression, gestures), etc.

[0079] The evaluation of the medical interviews that serve as training data is preferably conducted in terms of the content of the questions and the attitude of the examiner. The evaluation in terms of the content of the questions may be conducted in terms of whether or not the necessary questions were asked in order to collect information for evaluating the patient's health problems. The evaluation of the examiner's attitude may be conducted in terms of whether or not desirable actions, such as eye contact or mirroring, were performed in order to build a trusting relationship (rapport) between the patient and the medical professional.

[0080] (The importance of nonverbal communication in medical interviews)

[0081] FIG. 11 is a diagram illustrating non-verbal communication.

[0082] Referring to Figure 11, nonverbal communication is a way of conveying information and emotions without using words. Nonverbal communication includes elements such as body language, facial expressions, eye contact, tone and pitch of voice, personal space, grooming and clothing, and touch. Body language refers to body movements, posture, or gestures. Facial expressions are used to show emotions and reactions through facial expressions. Eye contact is used to convey intentions and interest through gaze. Tone and pitch of voice are used to express emotions and attitudes through changes in voice pitch, speed, and intensity. Personal space refers to the sense of distance and space between oneself and others. Grooming and clothing refer to clothing and appearance as a form of self-expression. Touch refers to conveying emotions and intentions through physical contact such as handshakes and hugs.

[0083] Mirroring is one of the nonverbal communication skills required in medical interviews. Mirroring involves imitating the facial expressions and movements of the other person as if they were a mirror. By mastering mirroring, the examiner can more easily build rapport with the patient, which is the purpose of the medical interview mentioned above, and can promote patient education and motivation for treatment.

[0084] Therefore, when the examiner DC practices a medical interview using the medical interview device 1, whether or not the examiner DC is appropriately mirroring the changes in the facial expression of the patient PT's facial image displayed on the display unit 103 can be used as a criterion for evaluating the examiner DC's attitude during the medical interview.

[0085] From the viewpoint of mirroring, as shown in FIG. 11(a), when the facial image of the patient PT displayed on the display unit 103 is "Image 1 (Normal)" (FIG. 7(a)), it is preferable that the examiner DC show a normal facial expression without any emotions, just like the patient PT. As shown in FIG. 11(b), when the facial image of the patient PT displayed on the display unit 103 is "Image 2 (Distressed)" (FIG. 7(b)), it is preferable that the examiner DC show a distressed facial expression just like the patient PT. As shown in FIG. 11(c), when the facial image of the patient PT displayed on the display unit 103 is "Image 3 (Smiling)" (FIG. 7(c)), it is preferable that the examiner DC show a smile just like the patient PT and touch their chin with their hand. In addition, when the patient PT makes a gesture such as nodding or pressing a painful area with their hand, it is preferable that the examiner DC also make a gesture such as nodding or pressing a painful area with their hand.

[0086] (Operation of medical interview device)

[0087] FIG. 12 is a flowchart showing the operation of the medical interview device 1 according to one embodiment of the present invention.

[0088] 2 and 12, the medical interview device 1 determines whether an input to start the medical interview (a predetermined operation such as selecting a scenario, or a voice input indicating that the medical interview is to begin) has been received (S101). The medical interview device 1 repeats the process of step S101 until it determines that an input to start the medical interview has been received.

[0089] In step S101, if it is determined that an input to start the medical interview has been received (YES in S101), the medical interview device 1 outputs the standard data 126 (patient's facial image and voice data) for the start of the medical interview via the display unit 103 and speaker 105 (S103). Next, the medical interview device 1 starts recording the voice of the examiner asking questions and photographing the examiner asking questions (S105), and determines whether voice input has been received (S107).

[0090] If it is determined in step S107 that voice input has not been accepted (NO in S107), the medical interview device 1 proceeds to the processing of step S119.

[0091] If it is determined in step S107 that voice input has been received (YES in S107), the medical interview device 1 attempts to identify the question asked by the examiner based on the received voice input using the model 121 (S109). In step S109, if the received voice input is interrupted for a predetermined period of time, the medical interview device 1 may identify the question asked by the examiner by treating the voice input received before the interruption as one question. The medical interview device 1 may wait silently without outputting an answer until the received voice input is interrupted for the predetermined period of time. Following the processing of step S109, the medical interview device 1 determines whether the question asked by the examiner has been identified (S111).

[0092] If it is determined in step S111 that the question asked by the examiner has been identified (YES in S111), the medical interview device 1 uses the model 122 or the question-and-answer information 124 to acquire the patient's answer to the identified question in the selected scenario and a facial image (S113). If the question-and-answer information 124 is used in step S113, the medical interview device 1 outputs example answers and facial images associated with the question in the question-and-answer information 124. Following the processing of step S113, the medical interview device 1 outputs the acquired patient's answer and facial image (S115), and proceeds to the processing of step S119.

[0093] If it is determined in step S111 that the question asked by the examiner cannot be identified (NO in S111), the medical interview device 1 outputs the patient's response indicating that it does not understand the question and an image of their face (S117), and proceeds to the processing of step S119.

[0094] In step S119, the medical interview device 1 determines whether or not an input to end the medical interview (a predetermined operation or a voice input indicating that the medical interview is to be ended) has been received (S119).

[0095] In step S119, if it is determined that the input to end the medical interview has not been accepted (NO in S119), the medical interview device 1 proceeds to the processing of step S107.

[0096] If it is determined in step S119 that an input to end the medical interview has been received (YES in S119), the medical interview device 1 ends recording the voice of the examiner asking questions and photographing the examiner asking questions (S121), and outputs standard data 126 (patient facial image and voice data) for use at the end of the medical interview (S123). Next, the medical interview device 1 uses the model 123 to evaluate the medical interview based on the recorded voice and photographed images (S125). The medical interview device 1 then outputs the evaluation to the display unit 103 or the like (S127), and ends the process.

[0097] (Effects of the embodiment)

[0098] The medical interview device 1 in the above-described embodiment outputs the patient's answers to the examiner's questions and a facial image of the patient. The facial expression of the output facial image changes depending on the content of the question asked by the examiner. This allows the examiner to practice displaying an attitude that corresponds to the patient's facial expression, and allows the examiner to fully acquire non-verbal communication skills. As a result, the examiner can fully acquire the skills necessary for a medical interview.

[0099] In addition, when the medical interview device 1 evaluates the medical interview conducted by the examiner, there is no need for a person to evaluate the medical interview, which reduces the cost required for the medical interview.

[0100] (Variation)

[0101] FIG. 13 is a diagram conceptually showing the configuration of a first modified example of the medical interview device 1 according to an embodiment of the present invention.

[0102] Referring to FIG. 13, the medical interview device 1 may store at least a portion of the models 121, 122, and 133 in an external server and acquire necessary data from the external server, such as using a technology using RAG (Retrieval-Augmented Generation).

[0103] Specifically, the medical interview device 1 in the first modified example is connected to each of the external servers 11, 12, and 13 via a network or the Internet 20. The medical interview device 1 and each of the external servers 11, 12, and 13 can communicate with each other. Each of the models 121, 122, and 123 is stored in each of the external servers 11, 12, and 13, instead of being stored in the storage device 108 (FIG. 2).

[0104] The question identification unit 153 transmits the voice received by the voice receiving unit 152 to the external server 11, and receives the question asked by the examiner from the external server 11. The external server 11 inputs the received voice into the model 121, obtains the question asked by the examiner from the model 121, and transmits it to the medical interview device 1.

[0105] The answer output unit 154 transmits the question identified by the question identification unit 153 to the external server 12, and receives the patient's answer to the question and an image of the patient's face from the external server 12. The external server 12 inputs the received question into the model 122, obtains the patient's answer to the question and an image of the patient's face from the model 122, and transmits them to the medical interview device 1.

[0106] The evaluation unit 157 transmits the voice received by the voice receiving unit 152 and the photographed object captured by the photographing unit 156 to the external server 13, and receives an evaluation of the medical interview from the external server 13. The external server 13 inputs the voice received by the voice receiving unit 152 and the photographed object captured by the photographing unit 156 to the model 123, obtains an evaluation of the medical interview from the model 123, and transmits it to the medical interview device 1.

[0107] FIG. 14 is a diagram conceptually showing the configuration of a second modified example of the medical interview device 1 according to an embodiment of the present invention.

[0108] Referring to FIG. 14, the medical interview device 1 may not include the evaluation unit 157 (FIG. 3), and the evaluator ET may evaluate the medical interview instead of the medical interview device 1. Specifically, the medical interview device 1 in the second modification may transmit the voice received by the voice receiving unit 152 and the photographed object captured by the photographing unit 156 to an external terminal 14 via a network or the Internet 20, where the evaluator ET can view them. The external terminal 14 is composed of a CPU, ROM, RAM, display unit, operation unit, speaker, and network interface. The evaluator ET views the photographed object of the examiner DC displayed on the display unit of the external terminal 14 and listens to the voice of the examiner DC output from the speaker of the external terminal 14, thereby viewing and evaluating the medical interview conducted by the examiner DC. The evaluator ET may also directly view and evaluate the medical interview conducted by the examiner DC without using the external terminal 14.

[0109] (others)

[0110] At the start of the medical interview, the medical interview device 1 may output standard audio and images in which the patient introduces himself or herself. The medical interview device 1 may output different answers depending on the order of questions asked by the examiner. When a closed question is received from the examiner, the medical interview device 1 may output an answer of "yes," "no," or a specific option. When a question about the examiner's name is received from the examiner, the medical interview device 1 may prompt the examiner to ask an additional question about the patient's first name by outputting only the last name as an answer. When a question about how to write the patient's first name in kanji is received from the medical interview device 1, the medical interview device 1 may explain the patient's first name in kanji. If the examiner's voice is too quiet to identify the question asked by the examiner, the medical interview device 1 may output an answer urging the examiner to ask the question again because the question was inaudible. When a predetermined question is received, the medical interview device 1 may output information necessary to verify the validity of the answer and evaluation, such as the type of learning model used. Furthermore, to verify the identity of the examinee, the medical interview device 1 may output predetermined questions, and when a correct answer is obtained from the examinee, the medical interview device 1 may output an evaluation of the medical interview of that examinee.

[0111] Furthermore, machine learning of the model 123 may be performed so that the evaluation of the medical interview reflects factors such as whether technical terms are used without explanation, whether both the patient's first and last name are asked, whether the examiner introduces themselves or greets the patient at the start of the medical interview, and whether the questions flow in a systematic manner.

[0112] The medical interview device of the present invention is not limited to a device used for practicing initial dental interviews, but may be used for practicing medical interviews in general, including medical interviews and subsequent medical interviews. The medical interview device of the present invention may also be used for practicing medical interviews for common examinations conducted by students of medical or dental schools.

[0113] The processes in the above-described embodiments and variations may be performed by software or hardware circuits. A program for executing the processes in the above-described embodiments may be provided, or the program may be recorded on a recording medium such as a CD-ROM, flexible disk, hard disk, ROM, RAM, or memory card and provided to the user. The program is executed by a computer such as a CPU. The program may also be downloaded to a device via a communication line such as the Internet.

[0114] The above-described embodiments and modifications can be combined as appropriate.

[0115] The above-described embodiments and modifications should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the above description, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]

[0116] 1 Medical interview device (example of medical interview device) 10 Camera 11,12,13 External Server 14 External Terminal 20. Internet 101 Control section 102 Operation section 103 Display section 104 Microphone 105 speakers 106 Camera 107 Network Interface 108 Storage device (an example of storage means) 111 CPU(Central Processing Unit) 112 ROM (Read Only Memory) 113 RAM (Random Access Memory) 121,122,123 model (example of the first and second trained models) 124 Question and answer information 125 Image Data 126 Standard Data 151 Operation reception unit 152 voice reception unit (an example of a voice reception means) 153 Question identification unit (an example of a question identification means) 154 Answer output unit (an example of an answer output means) 155 Video creation unit (an example of a video creation means) 156 Photography unit (an example of photography means) 157 Evaluation section (example of evaluation means) 500 Neural Networks 501 Input Layer 502 Middle Class 503 Output Layer 510 Problem Data 511 Teacher Data 1241,1242,1243,1244 Question and answer information ET Evaluator DC Examiner SD1, SD2, SD3 training data

Claims

1. 1. A medical interview device that provides a patient's answers to questions from a medical examiner, comprising: a voice receiving means for receiving a voice input of a question posed by the examiner; a question identification means for identifying a question asked by the examiner based on the voice input received by the voice receiving means, using a first trained model created by machine learning; an answer output means for outputting an answer of the patient to the question identified by the question identification means and a facial image of the patient; A medical interview device, wherein the facial expression of the facial image output by the answer output means changes depending on the content of the question asked by the examiner.

2. The system further comprises a storage means for storing question and answer information including a plurality of example questions different from one another, a plurality of example answers to the plurality of example questions, and a plurality of example facial images of the patient when outputting each of the plurality of example answers, 2. The medical interview device according to claim 1, wherein said answer output means outputs example answers and example face images associated with the questions identified by said question identification means in said question-and-answer information.

3. the storage means stores a plurality of pieces of question and answer information corresponding to a plurality of scenarios, an example answer associated with one example question common to each of the plurality of pieces of question and answer information among the plurality of example questions is different from another example answer in each of the plurality of pieces of question and answer information; 3. The medical interview device according to claim 2, wherein the answer output means outputs example answers and example facial images associated with a question identified by the question identification means in the question-and-answer information corresponding to a specific scenario among the plurality of scenarios.

4. 2. The medical interview device according to claim 1, wherein the answer output means uses a second trained model created by machine learning to input the question identified by the question identification means and outputs the answer and a facial image of the patient.

5. a moving image creating means for creating a composite moving image by combining a moving image of the patient's face with a voice of the patient's answer; 2. The medical interview device according to claim 1, wherein said answer output means outputs said composite video.

6. an imaging means for imaging the examiner asking questions; 2. The medical interview device according to claim 1, further comprising an evaluation means for evaluating the medical interview conducted by the examiner based on the voice input received by the voice receiving means and the photographed object captured by the photographing means.

7. 2. The medical interview device according to claim 1, wherein the facial expression of the facial image output by said answer output means changes at least between a smiling expression and a distressed expression depending on the content of the question asked by the examiner.

8. 2. The medical interview device according to claim 1, wherein, when the voice input received by the voice receiving means is interrupted for a predetermined period of time, the question identifying means identifies the question asked by the examiner by treating the voice input received by the voice receiving means before the interruption as one question.

9. 1. A control program for a medical interview device that provides patient responses to examiner questions, comprising: a voice receiving step of receiving a voice input of a question from the examiner; a question identification step of identifying a question asked by the examiner based on the voice input received in the voice receiving step by using a first trained model created by machine learning; an answer output step of outputting the answer of the patient to the question identified in the question identification step and a facial image of the patient; A control program for a medical interview device, wherein the facial expression of the facial image output in the answer output step changes depending on the content of the question asked by the examiner.

Citation Information

Patent Citations

  • Medical examination simulator

    JP2023183472A