Conversation-based mental disorder screening method and its device
A conversation-based method and apparatus enable remote mental disorder screening by analyzing responses to stimuli, enhancing accessibility and accuracy.
Patent Information
- Application Number
- JP2024103211
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-15
- Filing Date
- 2024-06-26
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2041-07-13
AI Technical Summary
Current mental disorder examinations require individuals to visit a specific location at a designated time, causing inconvenience and limiting accessibility.
A conversation-based method and apparatus that outputs stimuli such as stories, words, sounds, pictures, movements, or directions, receives responses, and analyzes correct answer rates or voice features to screen for mental disorders using a user terminal and analysis server, enabling remote screening.
Allows for convenient mental disorder screening at home or other locations without time constraints, improving accuracy through conversation content and voice data analysis.
Smart Images

Figure 0007701762000001 
Figure 0007701762000002 
Figure 0007701762000003
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a method and apparatus for screening psychiatric disorders through conversation.
Background Art
[0002] Current examinations for mental disorders (such as dementia, attention deficit disorder, learning disorder, schizophrenia, mood disorder, addiction, etc.) are conducted by experts in a specific space and at a specific time. Therefore, those who want to undergo an examination for the presence or absence of a mental disorder have the inconvenience of having to make a reservation for the examination and then visit a specific place such as a hospital at the reserved time.
Summary of the Invention
Problems to be Solved by the Invention
[0003] The technical problem to be solved by the embodiments of the present invention is to provide a method and an apparatus capable of easily screening psychiatric disorders in a conversation-based manner at home or other places rather than in a hospital without being restricted by time and space.
Means for Solving the Problems
[0004] An example of the conversation-based mental disorder screening method according to an embodiment of the present invention for achieving the above technical problem includes a step of outputting a stimulus including at least one or more of a story, a word, a sound, a picture, a movement, a color, and a direction, a step of receiving a response to the stimulus from an examination subject, and a step of comparing a correct answer rate of the response or a voice feature included in the response with a correct answer rate or a voice feature of a normal group and a disease group, or analyzing conversation content to screen for the presence or absence of a mental disorder.
[0005] An example of a conversation-based mental disorder screening device according to an embodiment of the present invention for achieving the above technical problem includes a data output unit that outputs a stimulus including at least one or more of a story, a word, a sound, a picture, a movement, a color, and a direction, a voice input unit that receives a response to the stimulus from an inspection target, and a voice / conversation analysis unit that compares the correct answer rate of the response or the voice features included in the response with the correct answer rate or voice features of a normal group and a disease group, or analyzes the conversation content to determine the presence or absence of a mental disorder.
[0006] An example of a recording medium according to an embodiment of the present invention for achieving the above technical problem is a computer-readable recording medium that stores computer-readable instructions, and when the instructions are executed by at least one processor, the at least one processor performs steps, and the steps include a step of outputting a stimulus including at least one or more of a story, a word, a sound, a picture, a movement, a color, and a direction, a step of receiving a response to the stimulus from an inspection target, and a step of comparing the correct answer rate of the response or the voice features included in the response with the correct answer rate or voice features of a normal group and a disease group, or analyzing the conversation content to screen for the presence or absence of a mental disorder.
Advantages of the Invention
[0007] According to an embodiment of the present invention, an inspection target can diagnose the presence or absence of a conversation-based mental disorder in a comfortable space such as home without having to visit a hospital deliberately. In addition, the accuracy of mental disorder screening can be improved by utilizing both conversation content and voice data.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0009] Hereinafter, with reference to the accompanying drawings, a conversation-based mental disorder screening method and its device according to an embodiment of the present invention will be described in detail.
[0010] FIG. 1 is a diagram showing an example of a schematic structure of an entire system implementing a mental disorder screening method according to an embodiment of the present invention.
[0011] Referring to FIG. 1, the system for mental disorder screening mainly includes a user terminal 100 and an analysis server 140. The user terminal 100 and the analysis server 140 can be connected via a communication network 150 such as a wired network or a wireless network.
[0012] The user terminal 100 includes a voice input device 110, a data output device 120, and a communication unit 130. Here, the voice input device 110 means a device capable of inputting sound such as a microphone, the data output device 120 means a speaker for outputting sound or a display device for outputting video, and the communication unit 130 means various communication modules capable of transmitting and receiving data to and from the analysis server 140. The user terminal 100 can further include various configurations necessary for the implementation of the present embodiment, such as a processor, a memory, etc. However, for the sake of convenience of explanation, the present embodiment shows mainly the configurations necessary for mental disorder screening.
[0013] In this embodiment, the user terminal 100 can be any terminal equipped with a voice input device 110, a data output device 120, and a communication unit 130. Therefore, it can be implemented on a general computer, a tablet PC, a smartphone, a smart refrigerator, a smart TV, an AI speaker, various IoT (Internet of Things) devices, etc., and is not limited to a specific device.
[0014] The analysis server 140 is a device that analyzes the data received from the user terminal 100 to determine the presence or absence of a mental disorder (such as dementia, etc.) to be examined. The analysis server 140 is not limited to the term "server" and can be implemented not only on a server but also on a general computer, a cloud system, etc.
[0015] This embodiment shows a structure in which the user terminal 100 and the analysis server 140 are connected via a communication network 150 for mental disorder screening, but it is not necessarily limited to this. For example, part or all of the functions performed by the analysis server 140 can be performed by the user terminal 100. When all the functions of the analysis server 140 are performed by the user terminal 100, the analysis server 140 may be omitted. That is, mental disorder screening can be performed by the user terminal 100 without the analysis server 140, and the result can be displayed. However, hereinafter, for the convenience of explanation, the structure of this embodiment in which the user terminal 100 and the analysis server 140 are connected via the communication network 150 will be mainly described.
[0016] FIG. 2 is a flowchart showing an example of a mental disorder screening method according to an embodiment of the present invention.
[0017] Referring to FIGS. 1 and 2, the user terminal 100 outputs a stimulus via the data output device 120 (S200). Here, the stimulus means a story, sound, image, etc. that can stimulate the vision or hearing of the subject to be examined. For example, the stimulus can be composed of at least one or more of a story, a word, a sound, a picture, a movement, a color, and a direction. FIGS. 5 and 6 show an example of a story stimulus.
[0018] In one embodiment, the user terminal 100 may receive the stimulation content in real time from the analysis server 140 and output it via the data output device 120, or may output the stimulation content pre-stored in the user terminal.
[0019] In another embodiment, interference stimuli may be arranged during or within the stimuli output via the data output device 120. Here, the interference stimuli are stimuli for disturbing the test subject in order to widen the response difference between the normal group and the disease group and improve the accuracy of mental disease screening. For example, in the case of the story stimuli as shown in FIG. 5, familiar words and situations appearing in the familiar story can be changed to unfamiliar words and situations for output, and the words and situations generated here correspond to the interference stimuli. The position and content of the interference stimuli for each stimulus may be predetermined. As another example, before the test subject answers, interference stimuli such as another story, sound, picture, etc. may be output before asking the test subject a question so that the memory of the previous stimulus can be erased, or a pause period of a certain time may be given between stimuli.
[0020] In another embodiment, the user terminal 100 can output by changing the content of the stimulus output in response to the response of the test subject in real time. The stimulus content changed in real time may be selected by the user terminal 100, or the response of the test subject may be provided to the analysis server 140 in real time and the stimulus changed by the analysis server 140 in real time may be received and output. For example, in the story stimuli of FIGS. 5 and 6, when the correct answer rate of the response of the test subject to the question is above a certain level, a story stimulus with more difficult content may be output, or other types of stimuli such as sound and image may be output. Depending on the correct answer rate of the test subject, the type of the next output stimulus etc. may be predetermined.
[0021] The user terminal 100 receives the response of the inspection target to the stimulus via the voice input device 110 (S210). For example, when the stimulus is a story stimulus as shown in FIG. 5, the user terminal 100 may output the story by sound via the data output device 120, visually display it on the screen, or output it simultaneously by sound and on the screen. After outputting the stimulus, the user terminal 100 can ask questions as shown in FIG. 6 and input the response of the inspection target via the voice input device 110.
[0022] The user terminal 100 can repeatedly perform the process of outputting the stimulus and receiving the response of the inspection target to the stimulus while changing the stimulus. For example, after the user terminal 100 outputs the first stimulus, it can receive the response of the inspection target to it, and after outputting the second stimulus, it can receive the response of the inspection target to it. That is, since the user terminal 100 can perform the inspection in the form of continuing a kind of conversation with the inspection target, for this reason, conventional AI speakers and various smart devices can be utilized in this embodiment.
[0023] The user terminal 100 transmits the response input from the inspection target to the analysis server 140, and the analysis server 140 analyzes the response of the inspection target to screen out mental disorders (S220, S230). Specifically, the analysis server 140 can analyze the response content of the inspection target to determine whether the correct answer rate and / or voice characteristics of the response of the inspection target to the stimulus are closer to the normal group or the disease group, so as to judge the presence or absence of mental disorders. An example of screening out mental disorders based on the correct answer rate of the inspection target is shown in FIG. 3, and an example of screening out mental disorders based on the voice characteristics of the inspection target is shown in FIG. 4. As another example, the analysis server 140 can also consider the number of words included in the response of the inspection target, the completion degree of the article, the presence or absence of the use of low-frequency words, the understanding degree of polysemous or ambiguous articles, the usage frequency of words expressing emotions, etc. in conjunction with the judgment of the presence or absence of mental disorders.
[0024] When the analysis of mental disorders is completed, the analysis server 140 can provide the mental disorder screening result to the user terminal 100 or a predetermined terminal (for example, the terminal of the protector or medical staff of the inspection target, etc.).
[0025] Figure 3 is a diagram showing an example of a method for grasping mental disorder screening according to an embodiment of the present invention based on the correct answer rate of a test subject.
[0026] Referring to Figure 3, after the analysis server 140 grasps the correct answer rate 300 of the test subject for the stimulus, it compares it with the correct answer rate 310 of the normal group and the correct answer rate 320 of the disease group. In the case of the stimuli of the stories in Figures 5 and 6, the analysis server 140 analyzes the responses of the test subject using various conventional speech recognition technologies, grasps the answers of the test subject for each question content, and then can determine whether the answer is correct. When a plurality of stimuli are output to the test subject, the analysis server 140 can grasp whether each stimulus is correct or not.
[0027] For example, when the correct answer rate 310 of the normal group and the correct answer rate 320 of the disease group of dementia diseases are defined for the stimuli of the stories in Figures 5 and 6, the analysis server 140 can grasp which group among the normal group and the disease group the correct answer rate 300 of the test subject is closer to and thereby grasp the presence or absence of dementia diseases. That is, when the correct answer rate of the normal group is 70% and the correct answer rate of the disease group is 30%, if the correct answer rate of the test subject is 20%, since the correct answer rate of the test subject is lower than that of the disease group, the analysis server 140 can determine that there is a dementia disease. As another example, if the correct answer rate of the test subject is between 30% and 70%, the analysis server 140 can determine that there may be a dementia disease because the correct answer rate of the test subject is closer to the disease group between the normal group and the disease group. In this case, it is also possible to calculate and provide the possibility of the presence of dementia diseases as a probability according to the relative distance between the correct answer rate of the test subject and the correct answer rates of the normal group and the disease group.
[0028] Figure 4 is a diagram showing an example of a method for grasping mental disorder screening according to an embodiment of the present invention based on the voice characteristics of a test subject.
[0029] Referring to FIG. 4, the analysis server 140 can analyze the voice feature 400 to be inspected and compare it with the voice feature 410 of the normal group and the voice feature 420 of the disease group. As examples of voice features, the analysis server 140 can analyze formants, Mel-Frequency Cepstral Coefficients (MFCCs), pitch, voice length, voice quality, etc.
[0030] The analysis server 140 has previously grasped and stored the voice feature 410 for the normal group and the voice feature 420 for the disease group, and can grasp whether the voice feature 400 to be inspected is closer to the normal group or the disease group, and select whether there is a mental disorder in the inspection target. For example, if the voice features 410 of the normal group and the voice features 420 of the disease group for dementia are defined based on the stimuli of the stories in FIGS. 5 and 6, the analysis server 140 can analyze how close the voice feature 400 to be inspected is to either the normal group or the disease group, and grasp the possibility of dementia disease. For example, when the voice feature 400 to be inspected is 80% similar to the voice feature 420 of the disease group, the analysis server 140 can output that the possibility of dementia disease is 80%.
[0031] The comparison of voice features can be performed in various ways. For example, the analysis server 140 has predetermined the values of voice features (e.g., formants, MFCCs, pitch, etc.) extracted from the response of the inspection target, grasps the values of predetermined voice features from the response of the inspection target, and creates a vector using these values as variables. The values of voice features of the normal group and the disease group are also previously created as vectors. The analysis server 140 can grasp the similarity (e.g., Euclidean distance, etc.) between the vector of the voice feature of the inspection target and each vector of the normal group and the disease group, and grasp which one it is more similar to.
[0032] In another embodiment, the analysis server 140 can analyze the conversation content and use it for screening for mental disorders. For example, the analysis server 140 can grasp the number of words, the degree of sentence completion, the use of low-frequency words (i.e., difficult words), the degree of understanding of polysemous or ambiguous sentences, the frequency of use of words expressing emotions, etc. from the responses of the test subject, and then compare them with the predetermined reference values for mental disorder screening to determine the presence or absence of mental disorders. For example, if the number of words is below a predetermined value, the analysis server 140 can determine that it is a mental disorder; if the frequency of use of low-frequency words is above a certain level, it can determine that it is not a mental disorder; if the frequency of use of words expressing emotions is above a certain level, it can determine that it is not a mental disorder. Alternatively, after grasping the degree of sentence completion, the degree of understanding of polysemous or ambiguous sentences, etc. for the responses of the test subject using artificial intelligence or various conventional text analysis techniques, the presence or absence of mental disorders can also be screened based on this.
[0033] In another embodiment, the analysis server 140 can analyze the conversation content using artificial intelligence. For example, the artificial intelligence model can be trained to classify the normal group and the disease group through conversations with people belonging to the normal group and conversations with people belonging to the disease group. The analysis server 140 can use the artificially intelligent model trained in this way to determine the presence or absence of mental disorders through conversations with the test subject. The artificial intelligence model can be composed of a model capable of having conversations with users, such as an AI speaker. For example, the analysis server 140 outputs daily conversations such as weather and date as stimuli through the artificial intelligence model, and can grasp which class of the normal group and the disease group the test subject belongs to using the conversation content grasped through the process of receiving the responses of the test subject to this.
[0034] The analysis server 140 can improve the accuracy of mental disorder screening by considering at least one or more of the correct answer rate 300 of the test subject in FIG. 3, the voice characteristics 400 of the test subject in FIG. 4, and the conversation content.
[0035] FIG. 5 and FIG. 6 are diagrams showing an example of a story stimulus among the stimuli used for mental disorder screening according to an embodiment of the present invention.
[0036] Referring to FIG. 5, the story stimulus includes a predetermined amount of story. For example, the story stimulus can be a well-known story such as the story of Hunbu. When the user terminal 100 outputs the story stimulus for mental disorder screening, it may output it as it is, or replace specific words with other predetermined words and then output. For example, by replacing the familiar word "rice ladle" in the familiar story of Hunbu being slapped on the cheek by a rice ladle with the unfamiliar "basin", the interference with dementia patients can be maximized.
[0037] Referring to FIG. 6, the user terminal 100 can receive a response to the story stimulus from the inspection target through questions.
[0038] FIG. 7 is a diagram showing the configuration of an example of a mental disorder screening device according to an embodiment of the present invention.
[0039] Referring to FIG. 7, the mental disorder screening device 600 includes a data output unit 610, an audio input unit 620, and an audio / conversation analysis unit 630.
[0040] The mental disorder screening device 600 may be implemented by the user terminal 100 and the analysis server 140 connected by the communication network 150 as shown in FIG. 1, or may be implemented only by the user terminal 140. For example, the data output unit 610, the audio input unit 620, and the audio / conversation analysis unit 630 can be implemented as an application and installed on a smartphone, an AI speaker, etc. to perform. Alternatively, the data output unit 610 and the audio input unit 620 are implemented by an application, installed and executed on a smartphone, an AI speaker, etc., and the audio / conversation analysis unit 630 can be implemented on the analysis server 140. The mental disorder screening device 600 can be implemented in various forms according to the embodiment.
[0041] The data output unit 610 outputs stimuli. For example, when the data output unit 610 is implemented in an AI speaker, a smartphone, etc., the data output unit 610 can output stimuli via the AI speaker. According to an embodiment, after receiving a stimulus from an external analysis server 140, the data output unit 610 can output it via the AI speaker. Alternatively, the data output unit 610 may output daily question contents such as date, weather, family relationships, etc. as stimuli.
[0042] The voice input unit 620 receives the response of the inspection target to the stimulus. For example, when the voice input unit 620 is implemented in an AI speaker, the answer of the inspection target to the stimulus can be input via the AI speaker.
[0043] The voice / conversation analysis unit 630 analyzes the response of the inspection target input via the voice input unit 620 to screen for the presence or absence of mental disorders. For example, as shown in FIGS. 5 and 6, when receiving the response of the inspection target to the story stimulus, the voice / conversation analysis unit 630 can analyze the correct answer rate of the inspection target and analyze the voice characteristics of the inspection target. The voice / conversation analysis unit 630 may be implemented in the user terminal 100 according to an embodiment, or may be implemented in the analysis server 140 of FIG. 1. When the voice / conversation analysis unit 630 is implemented in the analysis server 140 of FIG. 1, the voice input unit 620 can transmit the response of the inspection target to the analysis server 140.
[0044] In another embodiment, the voice / conversation analysis unit 630 analyzes the conversation content of the subject to be examined input via the voice input unit 620 to screen for the presence or absence of mental disorders. For example, the voice / conversation analysis unit 630 can analyze the number of words included in the response of the subject to be examined, the presence or absence of the use of low-frequency words, the frequency of use of words expressing emotions, etc., and then compare with a predetermined standard to screen for the presence or absence of mental disorders. Alternatively, the voice / conversation analysis unit 630 can use an artificial intelligence model trained using the conversation content of people belonging to the normal group and the disease group as learning data. In this case, the voice / conversation analysis unit 630 can grasp whether the conversation content with the subject to be examined belongs to the normal group or the disease group via the artificial intelligence model. The voice / conversation analysis unit 630 can determine the presence or absence of mental disorders using at least one or more of the correct answer rate, voice characteristics, and conversation content.
[0045] In addition, the present invention can be implemented as computer-readable code on a computer-readable recording medium. The computer-readable recording medium includes any type of recording device in which data readable by a computer system is stored. Examples of the computer-readable recording medium include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc. Further, the computer-readable recording medium can be distributed to a computer system connected to a network and the computer-readable code can be stored and executed in a distributed manner.
[0046] So far, the preferred embodiments of the present invention have been mainly looked at. Those having ordinary knowledge in the technical field to which the present invention pertains will understand that the present invention can be implemented in a modified form without departing from the essential characteristics of the present invention. Therefore, the disclosed embodiments should be considered from an explanatory perspective rather than a limiting perspective. The scope of the present invention is shown not in the above description but in the claims, and all differences within the equivalent scope should be construed as being included in the present invention.
Claims
1. outputting a first stimulus including at least one of a story, a word, a sound, a picture, a movement, a color, and a direction; receiving a first response to the first stimulus from the test subject, and a data output unit outputting a second stimulus that has been modified based on the first response; receiving a second response to the modified second stimulus from the test subject, and calculating the likelihood of a mental disorder of the test subject based on at least one of the following: a correct answer rate of the first response and the second response, a voice feature included in the first response and the second response, or an analysis of the conversation content of the first response and the second response; The step of outputting the modified second stimulus includes: A conversation-based mental disorder screening method performed by a mental disorder screening device, comprising a step of changing the type of the first stimulus or changing the content of the first stimulus when the voice characteristics of the first response to the first stimulus are closer to the voice characteristics of a normal group than to the voice characteristics of a disease group.
2. A conversation-based mental disorder screening method performed by the mental disorder screening device described in claim 1, characterized in that the step of outputting the first stimulus includes a step of replacing words contained in the story stimulus with predetermined other words and outputting it in order to increase the difference between a normal group and a disease group for the stimulus.
3. A conversation-based mental disorder screening method performed by the mental disorder screening device described in claim 1, characterized in that the step of outputting the first stimulus includes a step of determining the stimulus to be output according to the response content of the test subject.
4. A conversation-based mental disorder screening method performed by the mental disorder screening device described in claim 1, characterized in that the step of outputting the first stimulus includes a step of controlling the output interval between stimuli or placing an interfering stimulus in the stimulus, the interfering stimulus being a stimulus for confusing the test subject to widen the response difference between a normal group and a disease group and improve the accuracy of mental disorder screening.
5. 2. A conversation-based mental disorder screening method performed by a mental disorder screening device according to claim 1, wherein the speech features include at least one of formants, MFCC, pitch, duration, and sound quality.
6. a data output unit that outputs a first stimulus including at least one of a story, a word, a sound, a picture, a movement, a color, and a direction, and outputs a second stimulus that is changed based on a first response to the first stimulus; a voice input unit that receives a first response to the first stimulus and a second response to the modified second stimulus from a test subject; a voice analysis unit that calculates the possibility of a mental disorder of the test subject based on at least one of an analysis of the accuracy rate of the first response and the second response, a voice feature included in the first response and the second response, or an analysis of the conversation content of the first response and the second response; A mental disorder screening device, characterized in that the data output unit changes the type of the first stimulus or the content of the first stimulus when the voice characteristics of the first response to the first stimulus are closer to the voice characteristics of a normal group than to the voice characteristics of a disease group.
7. 1. A computer-readable medium storing computer-readable instructions that, when executed by at least one processor, cause the at least one processor to perform steps, the steps including: outputting a first stimulus including at least one of a story, a word, a sound, a picture, a movement, a color, and a direction; receiving a first response to the first stimulus from a test subject and outputting a second stimulus that is modified based on the first response; receiving a second response to the modified second stimulus from the test subject, and calculating the likelihood of a mental disorder of the test subject based on at least one of the following: a correct answer rate of the first response and the second response, a voice feature included in the first response and the second response, or an analysis of the conversation content of the first response and the second response; The step of outputting the second stimulus includes: A recording medium, comprising a step of changing the type of the first stimulus or changing the content of the first stimulus when the voice characteristics of the first response to the first stimulus are closer to the voice characteristics of a normal group than to the voice characteristics of a disease group.
Citation Information
Patent Citations
Dementia diagnostic support system
JP2007282992A
Dementia testing system
JP2018015139A
Cognitive function evaluation apparatus, cognitive function evaluation method, and program
JP2018050847A
JPP6712028B
System and method for assessing physiological state
WO2019081915A1