Method, program, information processing apparatus, and information processing system
The system enhances cognitive impairment evaluation by parallel execution of models to analyze spoken voice, improving accuracy and reducing the burden on medical staff and patients.
Patent Information
- Application Number
- JP2025016540
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-02-04
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-02-04
AI Technical Summary
Existing methods for evaluating cognitive impairment using voice data have limitations in accuracy and efficiency, particularly in early detection and diagnosis.
A method involving a computer system that executes a diagnosis support model, an MMSE estimation model, and a feature quantity explanation model in parallel and independently to analyze a user's spoken voice, providing comprehensive cognitive function evaluation.
Improves the accuracy and efficiency of cognitive function evaluation by reducing the burden on medical staff and patients, while maintaining high estimation accuracy through parallel and independent model processing.
Smart Images

Figure 0007713184000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method, a program, an information processing apparatus, and an information processing system.
Background Art
[0002] In recent years, with the progress of an aging society, early detection of cognitive impairment has become an important social issue. In order to address such issues, there have been efforts to utilize voice data to early estimate the risk of cognitive impairment.
[0003] In Japanese Patent Application Laid-Open No. 2011-255106 (Patent Document 1) below, based on a plurality of learning data including a plurality of types of prosodic feature amounts extracted from voice data and an HDS-R score obtained for the speaker of the voice data, a combination of prosodic feature amounts having the highest correlation with the HDS-R score is selected from the plurality of types of prosodic feature amounts by a feature amount selection unit 22. A weighting determination unit 24 determines a weighting for each of the selected combinations of prosodic feature amounts based on the selected combinations of prosodic feature amounts and the HDS-R score of each of the plurality of learning data. A feature amount extraction unit 28 extracts a plurality of types of prosodic feature amounts from the input voice data. It is disclosed that a risk degree calculation unit 30 calculates the risk degree of cognitive impairment based on the selected combination of the extracted prosodic feature amounts and the weighting determined by the weighting determination unit 24.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] By the technique shown in Patent Document 1, it is possible to estimate the risk of cognitive impairment using prosodic features and the like extracted from voice data. On the other hand, there is still room for improvement in the method for evaluating cognitive function.
[0006] An object of the present disclosure is to provide a method for improving the evaluation accuracy of cognitive function by utilizing the speech voice of a user.
Means for Solving the Problems
[0007] One embodiment shown in the present disclosure is a method executed by a computer including a processor and a memory, the method including: a step in which the processor receives an input of speech voice of a user; a step in which, based on the information of the speech voice, a diagnosis support model, an MMSE estimation model, and a feature quantity explanation model are executed in parallel and independently; and a step in which execution results of the diagnosis support model, the MMSE estimation model, and the feature quantity explanation model are output.
Effects of the Invention
[0008] According to the present disclosure, it becomes possible to further improve the evaluation accuracy of cognitive function by utilizing the speech voice of a user.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0010] The following is an embodiment of the present disclosure. The description of the embodiment of the present disclosure will be made with reference to the drawings. Also, in the description of the embodiment of the present disclosure, the same parts are denoted by the same reference numerals. Their names and functions are also the same. Therefore, detailed descriptions thereof will not be repeated.
[0011] <Overview of the First Embodiment> The diagnostic system 1 shown in one embodiment of the present disclosure supports the diagnosis of cognitive functions using the user's spoken voice. Note that the "diagnosis" in this specification does not mean a medical act on a human being performed by a medical professional, but means a process in which the system evaluates and judges an object, and is positioned as a means for presenting the result.
[0012] <1.1 Configuration Diagram of the Entire System> FIG. 1 is a block diagram showing an example of the overall configuration of the diagnostic system 1 of the present embodiment. As shown in FIG. 1, the diagnostic system 1 includes a first device 10 and a server 20. These devices are communicably connected to each other by a network 80.
[0013] The example of FIG. 1 shows the first device 10. The first device 10 is a terminal device that can use the diagnostic system 1. The first device 10 is a terminal device for receiving the spoken voice from the user. The first device 10 is realized by, for example, a smartphone, a tablet, or the like. In the present embodiment, the user is a user who uses the diagnostic system 1 through the first device 10. In this embodiment, the first device 10 and the server 20 are configured as separate devices, but the present invention is not limited to this. For example, a configuration in which all processing is completed on a single terminal device (edge device) can also be adopted as appropriate. In this case, the first device 10 is implemented in a form that includes the functions of the server 20, and the diagnostic process related to the recognition function can be executed without performing communication via the network 80, etc., and it can be flexibly adapted according to the user's usage environment.
[0014] The first device 10 accesses the server 20, for example, in response to an input from the user. The first device 10 requests predetermined information from the server 20, for example, and receives a response from the server 20. Then, the first device 10 outputs a diagnostic result based on the information received from the server 20.
[0015] The first device 10 provides an environment for the user to operate the diagnostic system 1 by executing a program. The first device 10 reads and executes a program to establish a communication connection between the first device 10 and the server 20. Then, the first device 10 transmits and receives data related to the diagnostic system 1 between the first device 10 and the server 20.
[0016] The first device 10 includes a communication IF (Interface) 12, an input device 13, an output device 14, a memory 15, a storage unit 16, and a processor 19.
[0017] The communication IF 12 is an interface for inputting and outputting signals so that the first device 10 can communicate with an external device.
[0018] The input device 13 is a device for receiving an input operation from the first device 10. Devices for receiving an input operation include, for example, pointing devices such as a touch panel, a touch pad, a mouse, and a keyboard.
[0019] The output device 14 is a device (such as a display, a speaker, etc.) for presenting information to the first device 10.
[0020] Memory 15 is for temporarily storing programs and data to be processed by programs and the like. Memory 15 is, for example, a volatile memory such as DRAM (Dynamic Random Access Memory).
[0021] Storage unit 16 is for storing data. Storage unit 16 includes, for example, a flash memory, an HDD (Hard Disk Drive), and the like.
[0022] Processor 19 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.
[0023] Server 20 is a device for managing information related to diagnostic system 1. For example, server 20 manages information necessary for diagnostic results, diagnostic records, etc. in diagnostic system 1. Server 20 appropriately transmits data necessary for diagnostic system 1 to first device 10 to cause first device 10 to output diagnostic results and the like.
[0024] Server 20 includes a communication IF 22, an input / output IF 23, a memory 25, a storage 26, and a processor 29.
[0025] Communication IF 22 is an interface for inputting and outputting signals so that server 20 can communicate with an external device.
[0026] Input / output IF 23 functions as an input device for receiving an input operation from first device 10. Also, input / output IF 23 functions as an interface with an output device for presenting information to first device 10.
[0027] The memory 25 is for temporarily storing programs and data processed by programs and the like. The memory 25 is, for example, a volatile memory such as DRAM (Dynamic Random Access Memory).
[0028] The storage 26 is for storing data. The storage 26 includes, for example, a flash memory, an HDD (Hard Disk Drive), and the like.
[0029] The processor 29 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.
[0030] <1.2 Functional Configuration of the First Device 10> FIG. 2 is a diagram showing the functional configuration of the first device 10. The first device 10 includes an antenna 111, a first wireless communication unit 121, a processor 19, an operation reception unit 130, a memory 15, a storage unit 16, a display 132, an audio processing unit 140, a microphone 141, and a speaker 142.
[0031] The antenna 111 radiates the signal emitted by the first device 10 into space as radio waves. Also, the antenna 111 receives radio waves from space and supplies the received signal to the first wireless communication unit 121.
[0032] The first wireless communication unit 121 performs modulation / demodulation processing and the like for transmitting and receiving signals via an antenna or the like for the first device 10 to communicate with other communication devices. The first wireless communication unit 121 is a communication module for wireless communication including a tuner, a high-frequency circuit, and the like, performs modulation / demodulation and frequency conversion of the wireless signals transmitted and received by the first device 10, and supplies the received signal to the processor 19.
[0033] The processor 19 controls the operation of the first device 10 by reading and executing the program stored in the storage unit 16. The processor 19 is realized, for example, by an application processor.
[0034] The operation reception unit 130 has a mechanism for receiving input operations from the user. The operation reception unit 130 is realized as a pointing device such as a mouse, a touch pad, a touch panel, a keyboard, a controller, a photographing means for sensing the movement of the user's body as an input operation, etc. The operation reception unit 130 receives, as input operations, movements of body parts such as hands and facial expressions of the user by sensing, for example, the movement of the user's body. The operation reception unit 130 determines the type of operation such as whether the user's operation is a flick operation, a tap operation, a drag operation, etc., based on the coordinates at which an input operation is received by the user touching a finger on a touch panel or the like.
[0035] The storage unit 16 is composed of a flash memory, a RAM (Random Access Memory), etc., and stores programs used by the first device 10 and various data received by the first device 10 from the server 20.
[0036] The storage unit 16 stores voice data information 161. In this embodiment, although the user's spoken voice is acquired and analyzed in real time in principle, it is also possible to hold the acquired voice data and transmit it later as necessary. With such a configuration, the diagnostic system 1 can utilize voice data while flexibly corresponding to the usage environment, communication status, etc.
[0037] The display 132 displays data such as text, voice, images, videos, etc. according to the control of the processor 19. The display 132 is realized by a display device such as an LCD (Liquid Crystal Display), an organic EL (Electro Luminescence), etc.
[0038] The voice processing unit 140 performs modulation and demodulation of voice signals. The voice processing unit 140 modulates the signal given from the microphone 141 and gives the modulated signal to the processor 19. Also, the voice processing unit 140 gives the voice signal to the speaker 142.
[0039] The microphone 141 receives voice input and provides a voice signal corresponding to the voice input to the voice processing unit 140. The microphone 141 receives, for example, the input of the user's spoken voice.
[0040] The speaker 142 converts the voice signal provided from the voice processing unit 140 into voice and outputs the voice to the outside of the first device 10.
[0041] When the processor 19 operates according to a program, it functions as an input operation reception unit 191, a transmission / reception unit 192, a data processing unit 193, and a notification control unit 194. The input operation reception unit 191 performs processing to receive a user's input operation on an input device such as an operation reception unit. When the operation reception unit 130 is a touch device, for example, the input operation reception unit 191 determines the type of operation, such as whether the user's operation is a flick operation or a drag operation, based on the information of the coordinates where the user touches the touch device with a finger or the like. The transmission / reception unit 192 performs processing for the first device 10 to transmit and receive data according to a communication protocol with an external device such as a server. The data processing unit 193 performs processing to perform an operation on the data received by the first device 10 according to a program and output the operation result to a memory or the like. The notification control unit 194 performs processing such as displaying a display image on the display 132, outputting voice to the speaker 142, and generating vibration by a vibrator or the like as processing for presenting information to the user.
[0042] <1.3 Functional Configuration of Server 20> FIG. 3 is a diagram showing the functional configuration of the server 20. As shown in FIG. 3, the server 20 functions as a communication unit 201, a storage unit 202, and a control unit 203.
[0043] The communication unit 201 performs processing for the server 20 to communicate with an external device.
[0044] The storage unit 202 stores a user database 2021, an audio database 2022, a diagnostic record database 2023, a diagnostic result database 2024, a diagnostic support model 2025, an MMSE estimation model 2026, a feature quantity explanation model 2027, and the like.
[0045] Details of the user database 2021, the audio database 2022, the diagnostic record database 2023, and the diagnostic result database 2024 will be described later.
[0046] The diagnostic support model 2025 is a model learned based on predetermined diagnostic criteria with the user's spoken voice as input. Note that the user's spoken voice may be a voice of a predetermined time length (for example, 60 seconds). As an example, this predetermined diagnostic criterion is a general term for judgment indicators defined based on known scientific bases such as a diagnostic protocol for cognitive impairment and clinical guidelines formulated in advance, including the knowledge of experts such as doctors and speech therapists. The diagnostic support model 2025 uses acoustic feature quantities (speech rate, voice frequency band distribution, number of speech interruptions, etc.) extracted from the user's spoken voice as input data. In addition, the diagnostic support model 2025 uses the diagnostic results of dementia by experts (e.g., presence or absence of cognitive impairment, degree of suspicion, etc.) as correct labels, and uses combinations of acoustic feature quantities (input data) as teacher data. That is, the diagnostic support model 2025 collects a large number of acoustic feature quantities and the corresponding diagnostic results (correct labels) of experts, and is constructed by supervised learning. In this learning process, the parameters are optimized based on the error between the output result of the model and the correct label, and finally, it can output the presence or absence of cognitive impairment, the degree of suspicion, etc., and the progress of symptoms.
[0047] Furthermore, by fully utilizing the output results of the diagnostic support model 2025, the diagnostic system 1 can replace part of the conventional diagnostic process by experts such as doctors, thus greatly reducing the workload of medical staff. In particular, in the medical and nursing care fields with a large number of patients and the elderly, there is a problem that the time for interviews by experts is limited. However, by using this model, it becomes possible to automatically present an initial evaluation regarding cognitive functions from the user's spoken voice, so that the examination time and human resources of experts can be effectively utilized. Therefore, this embodiment improves the diagnostic work efficiency of medical staff and contributes to reducing the burden on the entire medical field.
[0048] The MMSE estimation model 2026 is a model that estimates the score of the MMSE test (Mini-Mental State Examination) for evaluating cognitive functions using the user's spoken voice as input. Note that the user's spoken voice may be a voice with a predetermined time length. Specifically, the MMSE estimation model 2026 uses, as input data, the acoustic feature quantities extracted from the user's spoken voice, and uses the MMSE scores actually measured by experts as correct labels, and uses these combinations as teacher data. That is, the MMSE estimation model 2026 collects a large number of pairs of acoustic feature quantities and the corresponding MMSE scores (correct labels) of experts, and constructs this model by supervised learning. In this learning process, the parameters are optimized based on the error between the prediction result of the model and the correct label, and finally an estimated score is calculated in the range of 0 to 30 points. The result of the calculated MMSE score is displayed in, for example, a gauge format. Thereby, the user can intuitively grasp the rough position of his / her own estimated score.
[0049] Note that the MMSE test is a useful indicator for determining the severity of cognitive impairment to some extent, but it may not always match the definitive diagnosis made by a doctor. Usually, doctors summarize the results of interviews, MMSE scores, and detailed examinations to make a diagnosis, so there may be a discrepancy with the MMSE score. Also, the determination of the severity at the borderline level such as MCI (mild cognitive impairment) is considered very difficult. However, the MMSE score is widely used as a de facto standard indicator for determining severity, and corresponding numerical values for MCI (mild cognitive impairment), mild dementia, etc. have been defined. Therefore, the estimated value of the MMSE score is a reference value because it is an existing indicator that doctors trust.
[0050] In addition, since the MMSE test is conducted in a written test format, it generally takes about 10 to 15 minutes, and it may take about 20 minutes including preparation. There is also concern that the MMSE test may impose a mental burden on the patient or damage their dignity during the process. Also, there are situations where the severity of the patient cannot be fully grasped only from the diagnosis result, and there is often a need to finally refer to the MMSE score as well.
[0051] By referring to the MMSE score obtained by utilizing the MMSE estimation model 2026, the user can easily estimate the user's cognitive function level, severity, etc. Also, since the estimation result by this model can partially replace the actual MMSE test process, there is an advantage that it can reduce the test time, load, etc. performed by experts. In particular, in the medical and nursing care fields with a large number of patients and elderly people, it is possible to contribute to reducing the burden on medical staff because it is not necessary to allocate the resources of medical staff every time the MMSE test is conducted.
[0052] On the one hand, there is also an opinion that the accuracy based solely on the MMSE examination is not necessarily high, and a definitive diagnosis by a doctor is more reliable. For example, in terms of the MMSE score, data such as a score of 23 or less indicates suspected dementia (sensitivity 81%, specificity 89%), and a score of 27 or less indicates suspected mild cognitive impairment (MCI) (sensitivity 45% - 60%, specificity 65% - 90%) have been reported (Reference: https: / / www.jpn-geriat-soc.or.jp / tool / tool_02.html). Therefore, there are cases where it is difficult to make a definitive diagnosis based only on the MMSE examination results. However, as described above, the MMSE score is a useful indicator for roughly grasping the severity, and its reliability has been established from years of usage in actual medical practice. Therefore, in this embodiment, while omitting the conventional MMSE examination (written test), an estimated value of the MMSE score can be obtained from the user's spoken voice (for example, 60 seconds of spoken voice data). As a result, it is possible to reduce the burden on the patients, doctors, medical staff, etc. described above.
[0053] And since the diagnosis support model 2025 and the MMSE estimation model 2026 are constructed using their respective learning methods, etc., by using both models together, it is possible to present both the accuracy of the definitive diagnosis and the estimated value of the MMSE score. For example, referring to the estimation result of the diagnosis support model 2025, information close to the final definitive diagnosis by a doctor can be obtained, while by using the MMSE estimation model 2026, it is possible to grasp the score as a standard representing the severity of cognitive function as a dedicated indicator. That is, if the results of the diagnosis support model and the MMSE estimation model can be obtained in parallel, a composite value can be created with a small amount of time and effort, and it is expected to construct a practical diagnosis flow for both doctors and users. In this way, the diagnosis system 1 presents the results of different learning models together, making it easier for users, doctors, medical staff, etc. to obtain comprehensive information, increasing the accuracy of diagnosis, and making it easier to comprehensively understand the degree of cognitive function decline, etc.
[0054] The feature description model 2027 takes the user's spoken voice as input and visualizes the analysis results of the user's spoken voice in a quantified form. The feature description model 2027 is a model that scores indicators related to multiple voices within a certain range. Indicators related to multiple voices are, for example, intonation of the voice, length / frequency of silence, stability of voice volume, stuttering / trembling of the voice, degree of match with the speaking pattern of an average healthy person, etc. For example, the feature description model 2027 may score multiple voice indicators in the range from 1 to 5 points. Also, the feature description model 2027 may consider weighting another indicator within a certain range, etc. The examples of the feature description model 2027 shown above are merely examples and are not limited thereto. Note that the user's spoken voice may be a voice of a predetermined time length.
[0055] Specifically, the feature description model 2027 uses, as input data, acoustic feature quantities extracted from the user's spoken voice, and uses the indicator scores set by experts such as speech therapists as correct labels, and uses these as teacher data. That is, the feature description model 2027 collects a large number of pairs of acoustic feature quantities and corresponding indicator scores (correct labels), and constructs this model through supervised learning. In this learning process, parameters are optimized based on the error between the output result of the model and the correct label, and finally, for example, the estimation results of each evaluation indicator defined in advance in the range from 1 to 5 points are calculated. Each evaluation indicator is, for example, the length of the silent interval during speech, etc. The calculated evaluation indicators are visually displayed, for example, in the form of a bar graph. As a result, the user can intuitively grasp the problems and features in their own speech, which is also useful for doctors and experts in evaluating speech.
[0056] These diagnostic support models 2025, MMSE estimation models 2026, and feature description models 2027 are stored in the memory unit 202 and executed in parallel and independently by the processor 29 of the control unit 203, and the execution results are output. The server 20, for example, does not operate multiple models in one server. Instead, three small servers (FaaS) etc. start up in parallel and independently, each performs processing, and is discarded as soon as the processing is completed. With this parallel and independent processing method, each model can operate simultaneously without competing for the resources allocated to each other, and it is difficult for the processing content, processing timing, etc. between the models to interfere, so accurate estimation can be performed independently. Specifically, by allocating dedicated processes etc. for each model, resources such as memory and CPU can be divided and efficiently managed, enabling a system design where even a model with a high load is less likely to inhibit the processing of other models. Also, since the operations proceed in parallel between the models, the overall processing time is shortened. Moreover, since each task performs its own operation and is not affected by the results of others, it is possible to maintain the estimation accuracy. By adopting such parallel and independent processing, models with different purposes such as diagnostic support, MMSE estimation, and feature description can operate simultaneously while securing their respective computing resources and perform estimation independently.
[0057] Also, by performing parallel and independent processing, the load of each model can be efficiently distributed, minimizing unnecessary resource competition, waiting states, etc., so that energy consumption can be significantly reduced. That is, performing parallel and independent processing of multiple models can ensure sufficient processing power even in edge terminals with power consumption constraints and enable stable operation.
[0058] In this embodiment, a configuration in which a plurality of learned models operate in parallel and independently is adopted. However, in realizing such a configuration, the following problems may occur, so it is necessary to fully consider solutions to them.
[0059] First, in the case where all models simultaneously refer to the same audio data, access conflicts, unnecessary duplication, etc. may occur, leading to potential processing stalls. To address this issue, it is effective to buffer or duplicate the audio data in advance so that each model can handle independent inputs, and to adopt a design that centrally manages the feature extraction process as a common phase without complicating the exclusive control.
[0060] Furthermore, when the diagnostic support model 2025, the MMSE estimation model 2026, and the feature description model 2027 each require different acoustic features, there is a concern that redundant format conversion, extraction processing, etc. may overlap due to parallel execution, and the synchronization control may also become complicated. Therefore, it is desirable to sort out the definition of the necessary features in advance, collectively extract the parts that can be shared, and minimize unnecessary calculations by using parallelization, caching, etc. when individual processing is inevitably required.
[0061] Regarding the results of parallel execution, since the processing speeds of each model are different, it is assumed that only some models will complete first while other models are still continuing the processing. The diagnostic system 1 may introduce an asynchronous event-driven design considering such time differences and adopt a mechanism to integrate the completion notifications of each model in one place, so that the processing can proceed while suppressing result inconsistencies. Furthermore, the server 20 may be designed to ensure operational flexibility, such as presenting the results of the models that have completed first step by step and outputting the diagnostic results at the timing when all models have completed.
[0062] Also, in the diagnostic system 1, the result formats may be different for each model, for example, the diagnostic support model 2025 outputs text related to cognitive functions, the MMSE estimation model 2026 outputs numerical scores (e.g., text and gauge displays), and the feature description model 2027 outputs digitized results (e.g., graphs). To handle such format diversity, the server 20 can adopt a design where each model outputs based on a common schema (e.g., JSON), enabling unified display policies and reduced conversion costs.
[0063] Furthermore, if the diagnostic system 1 adopts a configuration that forcibly makes all models independent for the realization of parallelization, there is a risk that the consistency of the results will be impaired in situations where order dependence is required. Therefore, the server 20 may adopt a configuration that pre-distinguishes parts where parallel execution is not a problem and parts where order dependence should be maintained, and sets a control flow that clarifies the dependence relationships, such as inserting waiting control as necessary in preprocessing, aggregation stages, etc.
[0064] The server 20 can receive the user's spoken voice and, by simultaneously executing these models, obtain information for outputting a text display of the diagnostic support result, a gauge display of the MMSE estimation score, and a graph display of the voice analysis result. Then, the server 20 outputs such information. As a result, the user can obtain, in a form close to a definitive diagnosis by a doctor or the like, the diagnostic result, the MMSE estimation score, and furthermore, information on the feature quantity analysis of the voice over a plurality of indicators all at once.
[0065] The control unit 203 is realized by the processor 29 reading the program stored in the storage unit 202 and executing the instructions included in the program. The control unit 203 exhibits the functions shown as the reception control module 2031 and the transmission control module 2032 by operating according to the program.
[0066] The reception control module 2031 controls the process in which the server 20 receives a signal from an external device according to the communication protocol.
[0067] The transmission control module 2032 controls the process in which the server 20 transmits a signal to an external device according to the communication protocol.
[0068] <2 Data Structure> FIG. 4, FIG. 5, FIG. 6, and FIG. 7 are diagrams showing the data structure of the database stored in the server 20. Note that FIG. 4, FIG. 5, FIG. 6, and FIG. 7 are examples and do not exclude data not described.
[0069] Figure 4 is a diagram showing the data structure of the user database 2021. The user database 2021 is a database for managing information about users in the diagnostic system 1. In this embodiment, as an example, it has a configuration including the following items.
[0070] Each record in the user database 2021 in Figure 4 includes the item "User ID", the item "Name", the item "Age", and the item "Gender".
[0071] The item "User ID" indicates the identification information issued by the server 20 in the diagnostic system 1. Specifically, it is a unique ID required when managing users. For the item "User ID", IDs such as "U001" and "U002" are assigned, for example.
[0072] The item "Name" indicates the name of the user. For the item "Name", names such as "Taro Tanaka" and "Hanako Yamada" are used, for example.
[0073] The item "Age" indicates the age of the user. For the item "Age", numerical values such as "75 years old" and "80 years old" are recorded, for example.
[0074] The item "Gender" indicates the gender of the user. For the item "Gender", information such as "Male" and "Female" is recorded, for example.
[0075] The existence of the user database 2021 enables reference to past search results, answer histories, etc. for each user, improving the accuracy and convenience of the diagnostic system.
[0076] Figure 5 is a diagram showing the data structure of the voice database 2022. The voice database 2022 is a database for managing voice information in the information graph of the diagnostic system 1. In this embodiment, as an example, it has a configuration including the following items.
[0077] Each record in the voice database 2022 includes an item "voice data ID", an item "user ID", an item "voice file name", and an item "voice duration".
[0078] The item "voice data ID" indicates an ID for uniquely identifying voice data. The voice data includes the user's spoken voice. The item "user ID" is specifically an identifier required when managing voice data. For example, IDs such as "V001" and "V002" are assigned.
[0079] The item "user ID" is an ID indicating the user who registered the voice data. The item "user ID" corresponds to the "user ID" in the user database 2021.
[0080] The item "voice file name" indicates the file name of the stored voice data. The item "voice file name" is recorded in a format such as "Voice_U001_20231001.wav".
[0081] The item "voice duration" is information indicating the playback time of the recorded voice data. The item "voice duration" is, for example, the time required to play the data in forms such as 60 seconds, 30 seconds, and 120 seconds.
[0082] Due to the existence of the voice database 2022, the voice data recorded by the user can be centrally managed, and the recording date and time corresponding to each data can be easily referred to.
[0083] Figure 6 is a diagram showing the data structure of the diagnosis record database 2023. The diagnosis record database 2023 is a database for managing diagnosis information for each user in the diagnosis system 1, and can comprehensively grasp the history of each diagnosis. In this embodiment, as an example, it has a configuration including the following items.
[0084] Each record in the diagnostic record database 2023 includes the item "diagnostic record ID", the item "user ID", the item "number of diagnoses", the item "diagnosis date and time", the item "voice data ID", and the item "diagnosis result ID".
[0085] The item "diagnostic record ID" indicates an ID for uniquely identifying diagnostic data. The item "diagnostic record ID" is used to identify each diagnostic record in the diagnostic system 1. IDs such as "D001", "D002", etc. are assigned to the item "diagnostic record ID".
[0086] The item "user ID" is an ID indicating the user who received the diagnosis. The item "user ID" corresponds to the "user ID" in the user database 2021, enabling the tracking and management of multiple diagnosis results associated with the same user.
[0087] The item "number of diagnoses" represents the cumulative number of diagnoses for the same user. The item "number of diagnoses" is recorded in a form such as "1" for the first diagnosis and "2" for the second diagnosis. The existence of the item "number of diagnoses" makes it easier to manage the follow-up observation of the user and the changes in symptoms.
[0088] The item "diagnosis date and time" indicates the date and time when the diagnosis related to cognitive function was performed. The item "diagnosis date and time" is recorded including "year, month, day" and "time", such as "2023 / 10 / 01 10:00", so that the implementation timing of each diagnosis can be clearly grasped.
[0089] The item "voice data ID" is an ID indicating the voice data related to the diagnosis. The item "voice data ID" corresponds to the "voice data ID" in the voice database 2022, enabling the identification of which voice file should be referred to. This allows for the association with the voice data obtained during the diagnosis.
[0090] The item "Diagnosis Result ID" indicates an ID for managing the diagnosis results associated with a diagnosis. The item "Diagnosis Result ID" corresponds to the "Diagnosis Result ID" stored in the diagnosis result database 2024, enabling mutual reference between the diagnosis records and the diagnosis results. For the item "Diagnosis Result ID", IDs such as "T001" and "T002" are assigned, for example.
[0091] By providing the diagnosis record database 2023 in this way, it is possible to centrally manage the multiple diagnosis histories received by the user, as well as the associated voice data, diagnosis results, etc.
[0092] Figure 7 is a diagram showing the data structure of the diagnosis result database 2024. The diagnosis result database 2024 is a database for managing diagnosis results, voice analysis results, etc. in the diagnosis system 1. In this embodiment, as an example, it has a configuration including the following items.
[0093] Each record in the diagnosis result database 2024 includes the item "Diagnosis Result ID", the item "User ID", the item "Diagnosis Result", the item "MMSE Score", and the item "Voice Analysis Result".
[0094] As described above, the item "Diagnosis Result ID" is an ID for uniquely identifying a diagnosis result. The item "Diagnosis Result ID" corresponds to the item "Diagnosis Result ID" in the diagnosis record database 2023 and plays a role in linking the diagnosis result database 2024 and the diagnosis record database 2023. By assigning IDs such as "T001", "T002", etc. to the item "Diagnosis Result ID", it is possible to refer to which diagnosis result corresponds to which diagnosis record.
[0095] The item "User ID" is an ID indicating the user for whom the diagnosis result is targeted. The item "User ID" corresponds to the "User ID" in the user database 2021 and is associated with the diagnosis results received by each user.
[0096] The item "Diagnosis Result" indicates a comprehensive judgment, comprehensive findings, etc. regarding the user's cognitive function. Specifically, based on the evaluation results estimated by the diagnostic support model 2025, etc., comments such as "Possibility of Cognitive Function Decline" and "Good Cognitive Function" are recorded in the item "Diagnosis Result" so that the user's state can be clearly displayed. Cognitive function is generally closely related to the progression of dementia. Therefore, when the cognitive function of the user (diagnostician) is declining, it may be necessary to consider the possibility of dementia, MCI (mild cognitive impairment), etc.
[0097] The item "MMSE Score" indicates the result output by the MMSE estimation model 2026. The item "MMSE Score" calculates the estimated score in the range of 0 to 30 points. Numerical values such as "24" and "28" are stored in the item "MMSE Score", and based on this score, it becomes possible to estimate the level regarding the user's cognitive function.
[0098] The item "Voice Analysis Result" includes information such as the result of analyzing the user's spoken voice by the feature quantity description model 2027 and quantifying each of a plurality of voice indicators. The plurality of voice indicators are, for example, intonation of the voice, length / frequency of silence, stability of voice volume, stuttering / trembling of the voice, degree of coincidence with the speaking pattern of an average healthy person, etc. The item "Voice Analysis Result" records the score in a form such as "Silent Interval: 2 (Somewhat Long)", for example, and enables the analysis result to be quantified and displayed as a graph (e.g., bar graph). With such information, speech therapists, doctors, users, etc. can grasp problems from the voice aspect and understand the user's cognitive function evaluation.
[0099] <Operation of the First Embodiment> Next, the operation of each device constituting the diagnostic system 1 will be described. The following description will be given by taking the case of receiving voice from the user as an example.
[0100] FIG. 8 is a flowchart showing the processing flow of the cognitive function diagnosis support method according to the present embodiment. In the present embodiment, a configuration is adopted in which databases such as a user database 2021, an audio database 2022, a diagnosis record database 2023, and a diagnosis result database 2024 are referred to as necessary to perform efficient data management.
[0101] In step S8021, first, the first device 10 records the user's uttered voice and transmits the corresponding audio file to the server 20. The first device 10 records, for example, a voice of 60 seconds, which is a predetermined voice length. Note that although 60 seconds is exemplified in the present embodiment, it is not limited thereto. The server 20 registers the received audio file in the audio database 2022 and manages it in association with an audio data ID (for example, "V001"), a user ID (for example, "U001"), an audio file name, an audio time length, and the like. Thereby, the server 20 can uniformly handle the audio input from the user.
[0102] In step S8022, the server 20 refers to the audio data registered in step S8021 and performs determination by independently and parallelly executing a diagnosis support model, an MMSE estimation model, and a feature quantity explanation model. By combining these models, an advantage of being able to comprehensively evaluate the cognitive function of the user can be obtained.
[0103] In step S8023, the server 20 aggregates the results obtained in step S8022 and registers and stores them in the diagnosis database 2023 and the diagnosis result database 2024. In the diagnosis database 2023, a user ID, a diagnosis count, a diagnosis date and time, and an audio data ID are stored using a diagnosis record ID (for example, "D001") as a key. In the diagnosis result database 2024, a diagnosis result ID (for example, "T001") is assigned, and comments on the diagnosis result, an MMSE score, an audio analysis result, and the like are recorded.
[0104] In step S8024, the control unit 203 (processor 29) of the server 20 outputs the diagnosis result stored in step S8023 to the first device 10 and presents the details of the diagnosis result. The method of presenting the diagnosis result is, for example, presenting the evaluation regarding the cognitive function in words, presenting the numerical value of the MMSE score in a gauge format, and presenting the voice analysis result in a graph format or the like.
[0105] In this way, through a series of procedures from the voice input in steps S8021 to S8024, the determination using multiple models, and the output of the final result, the diagnosis system 1 of this embodiment analyzes the user's spoken voice and realizes a multi-faceted diagnosis method. Furthermore, by linking various databases, a comprehensive evaluation including the user's profile, past diagnosis history, etc. becomes possible, and it is expected that the accuracy and reliability of the diagnosis can be further improved.
[0106] FIG. 9 is a diagram regarding the diagnosis result. FIG. 9 shows an example of a diagnosis result screen when presenting the analysis result of the spoken voice in this embodiment. The voice information input from the user is analyzed by a diagnosis support model, an MMSE estimation model, etc., and the result is displayed in an easy-to-see form.
[0107] S9001 is an area for displaying the result of the diagnosis support model 2025 in words. S9001 is displayed, for example, as "Good cognitive function".
[0108] S9002 is an area for displaying the MMSE score output by the MMSE estimation model 2026 in a gauge format and in words. In S9002, the MMSE score is, for example, noted together with the score and the estimation error in the form of "28 points ±1" so that it is easy to make a judgment based on the numerical value. Also, by visually showing the score in a gauge format, the user can grasp the approximate current cognitive function level at a glance.
[0109] S9003 is a button for referring to detailed analysis results. When the user presses S9003, it may be configured to transition to a screen where, in addition to the feature amounts calculated from the voice data, differences and trends in the current and past scores can be viewed. The diagnostic system 1 outputs scores and the like obtained by quantifying each of a plurality of voice indexes calculated by the feature amount explanation model 2027. The plurality of voice indexes are, for example, intonation of the voice, length / frequency of silence, stability of voice volume, stuttering / trembling of the voice, degree of match with the speech pattern of an average healthy person, and the like. Furthermore, the diagnostic system 1 also has a function of visualizing those scores in a graph or the like. The diagnostic system 1 can display information such as silent interval: 2 (somewhat long) on the graph, for example, enabling the user, an expert, etc. to intuitively confirm each index. By viewing the detailed analysis results, the user can objectively recognize the current state regarding cognitive function. As a result, the user can grasp the trend of changes regarding cognitive function, improvement measures, etc., and it becomes easier to work on improving the state of cognitive function.
[0110] FIG. 10 shows a case different from the screen example of FIG. 9 and is a screen example for displaying analysis results obtained by a similar mechanism. In this screen example, since a situation suggesting a decline in the user's cognitive function is assumed, the presented text, the numerical value of the MMSE score, etc. are different from those in the case of FIG. 9.
[0111] In S10001, there is an area for displaying the result of the diagnostic support model 2025 in words. S9001 is displayed, for example, as "Possibility of cognitive function decline". By taking heed of the warning indicated by this message, the user can objectively grasp their own state, and the possibility of discovering their own symptoms early increases. As a result, the user, an expert, etc. can take necessary measures at an appropriate timing and can easily proceed quickly with measures against cognitive function decline.
[0112] S10002 is an area that displays the MMSE score output by the MMSE estimation model 2026 in a gauge format and in words. In S9002, the MMSE score combines the score and the estimation error in a form such as "24 points ± 1" to facilitate making a judgment based on the numerical value. Also, by visually showing the score in a gauge format, the user can grasp the current approximate cognitive function level at a glance.
[0113] As described above, some embodiments of the present disclosure have been explained. However, these embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are to be included in the scope and gist of the invention, and are also to be included in the invention described in the claims and the equivalent scope thereof.
Explanation of Reference Numerals
[0114] 1: Diagnostic system 10: First device 12: Communication IF 13: Input device 14: Output device 15: Memory 16: Storage unit 19: Processor 20: Server 22: Communication IF 23: Input / output IF 25: Memory 26: Storage 29: Processor 80: Network 111: Antenna 121: First wireless communication unit 130: Operation reception unit 132: Display 140: Audio processing unit 141: Microphone 142: Speaker 161: Audio data information 191: Input operation reception unit 192: Transmission / reception unit 193: Data processing unit 194: Notification control unit 201: Communication unit 202: Memory unit 203: Control unit 2021: User database 2022: Voice database 2023: Diagnostic record database 2024: Diagnostic result database 2025: Diagnostic support model 2026: MMSE estimation model 2027: Feature description model 2031: Reception control module 2032: Transmission control module
Claims
1. A method executed by a computer comprising a processor and a memory, the method comprising: the processor receiving an input of a user's spoken voice; based on the information of the spoken voice, executing a diagnosis support model, an MMSE estimation model, and a feature description model in parallel and independently; outputting execution results of the diagnosis support model, the MMSE estimation model, and the feature description model; performing, wherein the diagnosis support model has learned a correspondence relationship between acoustic feature quantities extracted from the spoken voice and a diagnosis result regarding cognitive impairment, and outputs a diagnosis result of cognitive impairment in words when the acoustic feature quantities are input; the MMSE estimation model has learned a correspondence relationship between the acoustic feature quantities and an MMSE test score, and outputs the MMSE test score numerically when the acoustic feature quantities are input; the feature description model has learned a correspondence relationship between the acoustic feature quantities and scores regarding a plurality of voice indexes, and outputs a quantified analysis result of the spoken voice when the acoustic feature quantities are input. A method.
2. The method according to claim 1, wherein the information of the spoken voice is voice having a predetermined time length.
3. A program for causing a computer comprising a processor and a memory to execute, the program causing the processor of the computer to receive an input of a user's spoken voice; based on the information of the spoken voice, execute a diagnosis support model, an MMSE estimation model, and a feature description model in parallel and independently; output execution results of the diagnosis support model, the MMSE estimation model, and the feature description model; perform, wherein the diagnosis support model has learned a correspondence relationship between acoustic feature quantities extracted from the spoken voice and a diagnosis result regarding cognitive impairment, and outputs a diagnosis result of cognitive impairment in words when the acoustic feature quantities are input; the MMSE estimation model has learned a correspondence relationship between the acoustic feature quantities and an MMSE test score, and outputs the MMSE test score numerically when the acoustic feature quantities are input; the feature description model has learned a correspondence relationship between the acoustic feature quantities and scores regarding a plurality of voice indexes, and outputs a quantified analysis result of the spoken voice when the acoustic feature quantities are input. A program.
4. An information processing apparatus comprising a control unit and a storage unit, wherein the control unit means for receiving an input of a user's spoken voice means for executing a diagnostic support model, an MMSE estimation model, and a feature quantity explanation model in parallel and independently based on the information of the spoken voice, means for outputting the execution results of the diagnostic support model, the MMSE estimation model, and the feature quantity explanation model, executes, the diagnostic support model has learned the correspondence between the acoustic feature quantities extracted from the spoken voice and the diagnostic results regarding cognitive impairment, and outputs the diagnostic results of cognitive impairment in words when the acoustic feature quantities are input, the MMSE estimation model has learned the correspondence between the acoustic feature quantities and the scores of the MMSE test, and outputs the MMSE test scores numerically when the acoustic feature quantities are input, the feature quantity explanation model has learned the correspondence between the acoustic feature quantities and the scores regarding a plurality of voice indices, and outputs the analysis results of the spoken voice in a quantified manner when the acoustic feature quantities are input, information processing apparatus.
5. An information processing system comprising an information processing apparatus, means for receiving an input of a user's spoken voice, means for executing a diagnostic support model, an MMSE estimation model, and a feature quantity explanation model in parallel and independently based on the information of the spoken voice, means for outputting the execution results of the diagnostic support model, the MMSE estimation model, and the feature quantity explanation model, has, the diagnostic support model has learned the correspondence between the acoustic feature quantities extracted from the spoken voice and the diagnostic results regarding cognitive impairment, and outputs the diagnostic results of cognitive impairment in words when the acoustic feature quantities are input, the MMSE estimation model has learned the correspondence between the acoustic feature quantities and the scores of the MMSE test, and outputs the MMSE test scores numerically when the acoustic feature quantities are input, the feature quantity explanation model has learned the correspondence between the acoustic feature quantities and the scores regarding a plurality of voice indices, and outputs the analysis results of the spoken voice in a quantified manner when the acoustic feature quantities are input, information processing system.
Citation Information
Patent Citations
Cognitive function evaluation apparatus and cognitive function evaluation system
JP2019083902A
System and method of presenting risk of dementia
JP2021099608A
Cognitive function determination apparatus, cognitive function determination system, and computer program
JP2021108843A
Acoustic feature amount extraction method in disease estimation program, disease estimation program using the acoustic feature amount, and device
JP2023015420A
Computer program, information processing device, and information processing method
JP2023098155A