A system and method for detecting cognitive decline using speech analysis.

The system uses a trained ensemble classifier to analyze speech samples for early detection of cognitive decline, enhancing the ability to identify and treat MCI and potential progression to Alzheimer's disease.

JP7843138B2Active Publication Date: 2026-04-09JANSSEN PHARMA NV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-04-14
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Current methods are inadequate for early detection of cognitive decline, particularly mild cognitive impairment (MCI), which can progress to Alzheimer's disease, lacking effective screening tools for timely intervention.

Method used

A system and method utilizing speech analysis through a trained ensemble classifier to identify cognitive decline by extracting features from speech samples, normalizing the data, and analyzing it with multiple component classifiers to generate an ensemble output indicating cognitive impairment.

Benefits of technology

Enables early detection and prediction of cognitive decline, allowing for timely intervention and treatment, improving patient outcomes by identifying at-risk individuals and providing targeted therapeutic interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007843138000003
    Figure 0007843138000003
  • Figure 0007843138000004
    Figure 0007843138000004
  • Figure 0007843138000005
    Figure 0007843138000005
Patent Text Reader

Abstract

A system and method for detecting cognitive decline in a subject using a classification system for detecting cognitive decline in a subject based on speech samples. The classification system is trained using speech data corresponding to audio recordings of speech from normal and cognitively declining patients to generate an ensemble classifier including a plurality of component classifiers and an ensemble module. Each of the plurality of component classifiers is a machine learning classifier configured to generate a component output that identifies the sample data as corresponding to a normal patient or a cognitively declining patient. The machine learning classifiers are generated based on a subset of available features. The ensemble module receives the component outputs from all of the component classifiers and generates an ensemble output that identifies the sample data as corresponding to a normal patient or a cognitively declining patient based on the component outputs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Claim of Priority) This application claims the priority of U.S. Provisional Patent Application No. 62 / 834,170, entitled "System and Method for Predicting Cognitive Decline", filed on April 15, 2019, the entire contents of which are incorporated herein by reference.

Background Art

[0002] Mild cognitive impairment (MCI) causes a measurable decline in cognitive abilities, including memory and thinking skills, that is slight but noticeable. The changes caused by MCI are not severe enough to affect daily life, and people with MCI may not meet the diagnostic criteria for dementia. However, those with MCI have a high risk of eventually developing Alzheimer's disease (AD) or other types of dementia. Early therapeutic intervention can lead to better prospects for the outcome.

[0003] Episodic memory is the memory of events or "episodes". Episodic memory includes a prospective (newly encountered information) or retrospective (past events) component. Decline in verbal episodic memory occurs earliest in patients with preclinical / pre-dementia AD and predicts disease progression. Assessment of decline in verbal episodic memory in MCI represents early cognitive changes and can be used as a screening tool for timely detection and initiation of treatment for early / preclinical AD.

Summary of the Invention

Means for Solving the Problems

[0004] An exemplary embodiment of the present invention relates to a method for detecting cognitive decline in a subject. The method includes obtaining subject baseline speech data corresponding to multiple audio recordings of the subject's speech in response to a first set of commands given to the subject, and obtaining subject test speech data corresponding to further audio recordings of the subject's speech in response to a second set of commands given to the subject. Furthermore, the method includes extracting multiple features from the subject baseline speech data and subject test speech data, and generating subject test data by normalizing the subject test speech data using the subject baseline speech data. The method further includes analyzing the subject test data using a trained ensemble classifier. The trained ensemble classifier includes multiple component classifiers and an ensemble module. Each of the multiple component classifiers is configured to produce a component output that identifies the subject test data as corresponding to a normal patient or a patient with cognitive decline. Each component classifier is configured to analyze a subset of features selected from multiple features. The ensemble module receives the component outputs from the component classifiers and is configured to produce an ensemble output that identifies the subject test data as corresponding to a normal patient or a patient with cognitive decline based on the component outputs. The multi-component classifier is trained using baseline training speech data corresponding to prior audio recordings of speeches from groups of normal and cognitively impaired patients responding to a first set of commands, and test training speech data corresponding to prior audio recordings of speeches from groups of normal and cognitively impaired patients responding to a second set of commands.

[0005] A device for detecting cognitive decline in a subject is provided. The device comprises an audio output device configured to generate audio output, an audio input device configured to receive an audio signal and generate data corresponding to the recording of the audio signal, and a display. Furthermore, the device comprises a processor and a non-temporary computer-readable storage medium containing an instruction set that can be executed by the processor. The command set can operate to direct the voice output equipment to provide a first command set to the subject audibly multiple times, receive subject baseline speech data from the voice input equipment corresponding to multiple audio recordings of the subject's speech in response to the first command set, direct the voice output equipment to provide a second command set to the subject audibly, receive subject test speech data from the voice input equipment corresponding to further audio recordings of the subject's speech in response to the second command set, extract multiple features from the subject baseline speech data and subject test speech data, generate subject test data by normalizing the subject test speech data using the subject baseline speech data, analyze the subject test data using a trained ensemble classifier to generate an output indicating whether the subject is likely to have cognitive impairment, and direct the display to provide a visual representation of the output to the user. The device further comprises memory configured to store the trained ensemble classifier. The trained ensemble classifier comprises multiple component classifiers and an ensemble module. Each of the multiple component classifiers is configured to generate a component output that identifies the subject test data as corresponding to a normal patient or a patient with cognitive impairment. Each component classifier is configured to analyze a subset of features selected from multiple features. The ensemble module receives component outputs from the component classifiers and generates an ensemble output that identifies subject test data as corresponding to normal patients or patients with cognitive impairment based on the component outputs.The multi-component classifier is trained using baseline training speech data corresponding to prior audio recordings of speeches from groups of normal and cognitively impaired patients responding to a first set of commands, and test training speech data corresponding to prior audio recordings of speeches from groups of normal and cognitively impaired patients responding to a second set of commands.

[0006] In another exemplary embodiment, a computer-implemented method for training a classification system is provided. The classification system is configured to detect cognitive impairment in a subject based on a sample of the subject's speech. The method includes obtaining training baseline speech data and training test speech data from groups of normal and cognitively impaired patients. The training baseline speech data corresponds to audio recordings of speech from groups of normal and cognitively impaired patients in response to a first set of instructions, and the training test speech data corresponds to audio recordings of speech from groups of normal and cognitively impaired patients in response to a second set of instructions. Furthermore, the method includes extracting a plurality of features from (i) the training baseline speech data and (ii) the training test speech data. The method further includes generating an ensemble classifier comprising a plurality of component classifiers and an ensemble module. Each of the plurality of component classifiers is configured to produce a component output that identifies sample data as corresponding to a normal patient or a cognitively impaired patient. Each component classifier is configured to analyze a subset of features selected from a plurality of features. The ensemble module is configured to receive component outputs from a component classifier and generate an ensemble output that identifies sample data as corresponding to normal patients or patients with cognitive impairment based on the component outputs. This method further includes generating a training dataset by normalizing training test speech data using training baseline speech data, and training the ensemble classifier using the training dataset.

[0007] A system for training the classification system is also provided. The classification system is configured to detect cognitive impairment in subjects based on the subjects' speech samples. The system includes a database configured to store training baseline speech data and training test speech data from groups of normal and cognitively impaired patients. The training baseline speech data corresponds to audio recordings of speech from groups of normal and cognitively impaired patients in response to a first set of instructions, and the training test speech data corresponds to audio recordings of speech from groups of normal and cognitively impaired patients in response to a second set of instructions. The system further includes a computing unit operablely connected to communicate with the database. The computing unit includes a processor and a non-temporary computer-readable storage medium containing an instruction set that can be executed by the processor. The instruction set can operate to retrieve the training baseline speech data and training test speech data from the database, extract multiple features from (i) the training baseline speech data and (ii) the training test speech data, and generate an ensemble classifier comprising multiple component classifiers and an ensemble module. Each of the multiple component classifiers is configured to produce a component output that identifies sample data as corresponding to a normal patient or a cognitively impaired patient. Each component classifier is configured to analyze a subset of features selected from multiple features. The ensemble module receives component outputs from the component classifiers and generates an ensemble output that identifies sample data as corresponding to normal patients or cognitively impaired patients based on the component outputs. The instruction set can further operate to generate a training dataset by normalizing training test speech data using training baseline speech data, and to train the ensemble classifier using the training dataset. The system further includes memory configured to store the trained ensemble classifier.

[0008] These and other aspects of the present invention will become apparent to those skilled in the art after reading the following “Modes for Carrying Out the Invention,” including the drawings and the appended claims. [Brief explanation of the drawing]

[0009] [Figure 1] This invention illustrates a system for training a classification system for detecting cognitive decline based on speech samples of a subject, according to an exemplary embodiment of this application. [Figure 2] This invention provides a method for training a classification system for detecting cognitive decline based on speech samples of a subject, according to an exemplary embodiment of this application. [Figure 3] This application illustrates an ensemble classifier comprising multiple component classifiers and an ensemble module according to an exemplary embodiment of this application. [Figure 4] Figure 3 illustrates a method for independently selecting subsets of speech features for each component classifier of an exemplary ensemble classifier. [Figure 5] This application illustrates an apparatus for detecting cognitive decline based on a subject's speech sample according to an exemplary embodiment. [Figure 6] This application describes a method for detecting cognitive decline based on a subject's speech sample according to an exemplary embodiment. [Figure 7] This invention illustrates an exemplary system including an ensemble classifier for detecting cognitive decline based on speech samples from a subject according to Embodiment I of this application. [Figure 8] In an exemplary embodiment of the Rey Auditory Verbal Learning Test (RAVLT) according to Embodiment I of this application, data corresponding to the average number of words recalled across different steps are shown. [Modes for carrying out the invention]

[0010] This application relates to an apparatus and method for detecting and / or predicting cognitive decline, particularly mild cognitive impairment (MCI), by analyzing data corresponding to speech samples from a subject or patient obtained from a neuropsychological test, such as a word list recall (WLR) test, using a computer-implemented method. Neuropsychological tests can screen a subject's cognitive abilities, such as memory. Furthermore, this application includes a system and method for training a classification system configured to detect cognitive decline (e.g., MCI) based on a subject's WLR speech samples, and further / or predict the onset of cognitive decline or dementia (e.g., AD).

[0011] Figure 1 shows an exemplary embodiment of a system 100 for training a classification system to detect and / or predict cognitive decline, particularly MCI, based on a subject's speech samples. System 100 includes a database 110 for storing various types of data, including data corresponding to audio recordings of previously conducted WLR tests. Specifically, the database 110 includes training baseline speech data 112 and one or more sets of training test speech data 114, each acquired under different experimental conditions, as will be further described below. The database 110 may be stored on one or more non-temporary computer-readable storage media.

[0012] The database 110 may be operably connected to the arithmetic unit 120 to provide some or all of the data stored in the database 110 to the arithmetic unit 120, or to enable the arithmetic unit 120 to retrieve some or all of the data stored in the database 110. As shown in Figure 1, the database 110 is connected to the arithmetic unit 120 via a communication network 140 (e.g., the Internet, a wide area network, a local area network, a cellular network, etc.). However, it is also possible to connect the database 110 directly to the arithmetic unit 120 via a wired connection. In this embodiment, the arithmetic unit 120 comprises a processor 122, a computer-accessible medium 124, and an input / output device 126 for receiving data and / or instructions to the arithmetic unit 120 and / or transmitting data and / or instructions from the arithmetic unit 120. The processor 122 may include, for example, one or more microprocessors and may use instructions stored in the computer-accessible medium 124 (e.g., a memory storage device). The computer-accessible medium 124 may be, for example, a non-temporary computer-accessible medium containing executable instructions internally. The system 100 may further include a memory storage device 130 provided separately from the computer-accessible medium 124 for storing the ensemble classifier 300 generated and trained by the system 100. The memory storage device 130 may be part of the arithmetic unit 120, or it may be outside the arithmetic unit 120 and operably connected to the arithmetic unit 120. The memory storage device 130 may be connected to a separate arithmetic unit (not shown) for detecting and / or predicting cognitive decline in a subject. In another embodiment, the ensemble classifier 300 may be stored in another memory storage device (not shown) connected to a separate arithmetic unit (not shown) for detecting and / or predicting cognitive decline in a subject.

[0013] Figure 2 shows an exemplary embodiment of Method 200 for training a classification system to detect and / or predict cognitive decline in subjects based on the subjects' speech samples. In particular, Method 200 generates and trains an ensemble classifier 300 to analyze speech samples (e.g., WLR speech samples) to determine whether a sample correlates more with a normal patient or a patient with cognitive impairment (e.g., a patient with mild cognitive impairment). WLR speech samples can be obtained from a database 110 that can store data corresponding to previously recorded audio files from various WLR tests, such as experiments using RAVLT to assess verbal episodic memory. RAVLT is a word list-based tool administered by an examiner that can be used to measure verbal episodic memory. It can be used to detect and / or generate scores related to verbal memory, such as learning speed, short-term and delayed verbal memory, recall ability after distracting stimuli, cognitive memory, and learning patterns (serial position effect), which correlate with cognitive ability.

[0014] In step 202, the computing unit 110 receives from the database 110 multiple sets of training baseline speech data 112 corresponding to multiple sets of preceding audio recordings of speeches from groups of normal and cognitively impaired patients. Each set of training baseline speech data 112 corresponds to speeches from groups of normal and cognitively impaired patients in response to the same set of commands relating to listening to a word list and then playing and uttering the same word list. For example, training baseline speech data 112 corresponds to speeches from groups of normal and cognitively impaired patients in response to a first set of commands relating to listening to a first word list and then playing and uttering the first word list immediately following. In a particular embodiment, training baseline speech data 112 may include data corresponding to audio recordings from groups of normal and cognitively impaired patients from learning tests of WLR tests, such as the RAVLT test. Step 202 may utilize any number of preferred sets of training baseline speech data 112 to establish average baseline characteristics that can be compared with the training test speech data 114 to identify and / or enhance feature differentiation signals contained within the training test speech data 114. Specifically, the training baseline speech data 112 may include learning phase data from any suitable WLR test. The training baseline speech data 112, in particular the learning phase data, serves as baseline characteristics that can be compared to the training test speech data 114 to generate a quantitative representation of the cognitive load of normal and cognitively impaired patient groups. In some embodiments, at least three sets, at least five sets, or at least ten sets of the training baseline speech data 112 may be used. In one exemplary embodiment, five sets of the training baseline speech data 112 are used.

[0015] Furthermore, the computing unit 110 receives a set of training test speech data 114 from the database 110. The set of training test speech data 114 is used by method 200 to generate a training dataset 330 for training the ensemble classifier 300, and / or to generate the ensemble classifier 300, as will be further described below. The training test speech data 114 corresponds to preceding audio recordings of speeches from groups of normal and cognitively impaired patients in response to at least one different set of commands regarding reproduction and utterance, compared to the training baseline speech data 112. For example, the different set of commands could instruct a patient to hear a different list of words and immediately reproduce and utter that different list of words. As another example, a different set of instructions could instruct a patient to hear the same word list used in the training baseline speech data 112, but to reproduce and pronounce the word list after a distraction task and / or after a delay of a certain time period (e.g., at least about 10 minutes or about 10 minutes, at least about 20 minutes or about 20 minutes, or at least about 30 minutes or about 30 minutes). In a particular embodiment, the training test speech data 112 could include data corresponding to audio recordings of speeches from groups of normal and cognitively impaired patients from distraction tests, post-distraction tests, and / or time-delay tests of a WLR test, such as the RAVLT test.

[0016] In a particular embodiment, the method 200 for training a classification system to detect and / or predict cognitive decline in a subject based on the subject's speech samples can utilize clinical test data obtained as part of the RAVLT test from normal and MCI patients. Data corresponding to the average number of words reproduced across different steps of an exemplary embodiment of the RAVLT test is shown in Figure 8. As shown in Figure 8, the RAVLT test may include several different tests (e.g., tests I-V 811-815) that are part of the learning phase 802 of the RAVLT test. In step 202, the computing unit 110 may acquire speech data that is part of this learning phase 802 of the RAVLT test as training baseline speech data 112. Furthermore, Figure 8 shows that the RAVLT test includes different types of recall tests, namely, interrupted test B 822m, post-interruption test 824, and 20-minute delayed recall 826. The computing unit 110 may acquire speech data collected in any one of these recall tests as training test speech data 114 for step 202.

[0017] In step 204, the computing unit 110 analyzes and extracts multiple speech features from each of the sets of speech data received from the database 110, namely (i) multiple sets of training baseline speech data 112 and (ii) sets of training test speech data 114. In particular, the computing unit 110 extracts a set of the same type of speech features from each of the above datasets. The computing unit 110 can extract any suitable type of acoustic characteristics for analyzing the speech recording of spoken speech as speech features from the sets of speech data, such as mean and standard deviation data values ​​corresponding to acoustic characteristics. For example, a speech feature may include one or more exemplary acoustic characteristics of a speech recording enumerated and defined in Table 1 below. In some embodiments, a speech feature may include all or a subset of the mean and / or standard deviation data values ​​corresponding to these exemplary acoustic characteristics across frames of the speech recording.

[0018]

Table 1

[0019] The arithmetic unit 120 can extract any suitable number of speech features from each speech data set. When the number of speech features increases, the prediction and / or analysis performance of the system and method of the present application can be improved, but the computational load may increase. Therefore, an appropriate number of speech features can be selected to balance the prediction and / or analysis performance with the computational efficiency. In some embodiments, the arithmetic unit 120 can extract at least 24, at least 30, at least 50, or at least 100 different speech features from each speech data set. In one embodiment, the arithmetic unit 120 can extract 5 to 150 speech features, 10 to 100 speech features, or 12 to 50 speech features from each speech data set.

[0020] In step 206, the computing unit 120 generates an ensemble classifier 300 based on a set of training test speech data 114 received from the database 110 and relevant speech features relating to the training test speech data 114 acquired in the preceding step (step 204). As shown in Figure 3, the ensemble classifier 300 comprises a plurality of component classifiers 310 and an ensemble module 320. Each of the component classifiers 310 is a machine learning classifier configured to produce a component output that identifies sample data as corresponding to sample data from normal patients or patients with cognitive impairment. When the component classifiers 310 are trained using data obtained from normal patients and MCI patients, each of the component classifiers 310 is a machine learning classifier configured to produce a component output that identifies sample data as corresponding to sample data from normal patients or MCI patients. More specifically, each of the component classifiers 310 is a support vector machine (SVM). Alternatively, each of the component classifiers 310 may utilize other suitable supervised learning modules, such as a logistic regression module, a fast gradient boosting module, a random forest module, or a simple Bayes module. Each of the component classifiers 310 is generated by the computing unit 120 and analyzes a downsampled subset of speech features from the training test speech data 114. For each component classifier 310, the subset of speech features is selected separately and independently using the method 400 shown in Figure 4 and further described below. The ensemble classifier 300 can include any suitable number N of component classifiers 310. For example, the ensemble classifier 300 can include at least 10, at least 20, at least 30, or at least 50 component classifiers 310. In a particular embodiment, the ensemble classifier 300 includes 30 component classifiers 310.

[0021] The ensemble module 320 is configured to receive component outputs from all of the component classifiers 310 and generate an ensemble output 340 that identifies the sample data as corresponding to a normal patient or an MCI patient based on the component outputs. The ensemble module 320 can utilize any suitable method for determining the ensemble output 340 based on the component outputs provided by each of the component classifiers 310. For example, the ensemble module 320 can utilize a bagging or aggregating method that receives and considers the component outputs from each of the component classifiers 310 with equal weights. Alternatively, the ensemble module 320 can utilize other methods in which different weights are given to the component outputs, such as an adaptive boosting method or a gradient boosting method, for example.

[0022] Figure 4 shows an exemplary embodiment of Method 400 for independently selecting a subset of speech features for each component classifier 310. The computing unit 120 repeats Method 400 each time step 206 selects a desired downsampled subset of speech features to generate each component classifier 310. In other words, the computing unit 120 analyzes the training test speech data 114 using Method 400 to select a desired downsampled subset of speech features to generate a first component classifier 311, repeats the analysis to select another downsampled subset of speech features to generate the next component classifier 312, and so on, to generate subsequent component classifiers 313-317 until the computing unit 120 selects a downsampled subset of speech features to generate the Nth component classifier 317. Method 400 enables the computing unit 120 to generate an ensemble classifier 300 in which each component classifier 310 selects a different downsampled subset of speech features, so that each component classifier 310 models only a subset of all available speech features, but the ensemble classifier 300 as a whole provides sampling across a larger number of speech features. This structure of the ensemble classifier 300 improves computational efficiency by limiting each component classifier 310 to analyzing only the subset of speech features selected for each component classifier, while providing a module with improved predictive and / or analytical performance that incorporates a larger number of speech features.

[0023] In step 402, the computing unit 120 analyzes the training test speech data 114 to obtain a subsample of the training test speech data 114. As discussed below, this subsample is used by the computing unit 120 to identify desired parameters and features for the component classifier 310 generated for fewer speech features. The subsample includes a first number of samples of the training test speech data 114 from normal patients and a second number of samples of the training test speech data 114 from cognitively impaired patients. The first number of samples may be randomly selected from entries in the training test speech data 114 from normal patients. Similarly, the second number of samples may be randomly selected from entries in the training test speech data 114 from cognitively impaired patients. In step 402, in order to provide a substantial balance between the two different classifications of patients (i.e., normal patients and cognitively impaired patients), the second number of samples is at least 80% of the first number of samples, at least 90% of the first number of samples, or at least 95% of the first number of samples. Preferably, the subsamples are balanced such that the ratio of a first number of samples to a second number of samples is 1:1. This preferred subsampling allows the remainder of Method 400 to continue with data that is evenly balanced between the two classifications, namely normal patients and cognitively impaired patients. Typically, the training test speech data 114 may include more normal patients than cognitively impaired patients, and thus may result in unbalanced data between the two classifications. Balanced subsampling helps address any classification imbalance that may exist in the training test speech data 114, which could result in a classifier biased towards the larger classification, i.e., normal patients (which could lead to false negatives in the classifier and cause patients who should be identified as corresponding to cognitively impaired patients to be missed).Therefore, the ensemble classifier 300 combines several individual component classifiers 310 for subsets of speech features selected using substantially equal or equal proportions of normal and cognitively impaired patients, thus resulting in a balanced classifier that can cross-sample different speech features and thus learn from the entire training dataset.

[0024] In step 404, the computing unit 120 analyzes all speech features of a subsample of the training test speech data 114 and ranks the speech features based on predetermined criteria. Speech features can be ranked based on any suitable statistical criteria for identifying the features that most significantly contribute to the feature discrimination signal observed in the subsample of the training test speech data 114. Specifically, speech features can be ranked based on the importance of each speech feature to the subsample of the training test speech data 114. Specifically, if the component classifier 310 is an SVM, the computing unit 120 can rank speech features based on their importance in designating the difference (e.g., decision boundary) between normal patients and cognitively impaired patients as observed in the subsample obtained from the preceding step (step 402). More specifically, each speech feature can be ranked by a coefficient corresponding to the SVM hyperplane (e.g., decision boundary) based on the subsample of the training test speech data 114. Speech features with smaller coefficients to the SVM hyperplane are considered relatively less significant in specifying the SVM decision boundary, and therefore have lower feature importance and a lower rank.

[0025] In step 406, the computing unit 120 selects a subset of speech features from the ranking of speech features generated in the preceding step (step 404) based on a predetermined ranking threshold (for example, selecting only the top x speech features). The selected subset of x speech features is used by the computing unit 120 to generate a component classifier 310 for analyzing the downsampled x features. By selecting the top-ranked x speech features, the computing unit 120 generates a component classifier 310 that models the features that most significantly contribute to the feature discrimination signals contained in the training test speech data 114, while keeping the computational cost of the component classifier 310 to the number of downsampled x speech features. It is thought that any suitable number x of top-ranked features can be selected. For example, step 406 can select at least 10, at least 20, or at least 30 top-ranked features. In one particular embodiment, the top 20 ranked features are selected for each component classifier 310.

[0026] Returning to Method 200, in step 208, the computing unit 120 generates a training dataset 330 for training the ensemble classifier 300 generated from the preceding step (step 206). Specifically, the computing unit 120 generates a training dataset 330 that includes training baseline speech data 112, and in particular training test speech data 114 normalized by multiple sets of the training baseline speech data 112. In one embodiment, the training baseline data 112 includes learning phase data from a WLR test (e.g., a RAVLT test), and the training dataset 330 is generated by normalizing the training test speech data 114 with the learning phase data from the WLR test. The training dataset obtained from this embodiment provides quantitative values ​​corresponding to the cognitive load of groups of normal and cognitively impaired patients. Specifically, the computing unit 120 generates a training dataset 330 that includes each feature of the training test speech data 114 normalized by the average of the corresponding features across multiple sets of the training baseline speech data 112. More specifically, each feature of the training dataset 330 can be obtained by subtracting the mean of the features across multiple sets of the training baseline speech data 112 from the features of the training test speech data 114. By normalizing the speech features of the training test speech data 114 across multiple sets of the training baseline speech data 112, the feature discrimination signals contained within the training test speech data 114 can be enhanced, improving the predictive and / or analytical performance of the ensemble classifier 300 trained on such normalized data.

[0027] In step 210, the computing unit 120 trains the ensemble classifier 300 using the training dataset 330 generated in step 208 to produce a trained ensemble classifier 516. Each of the component classifiers 310 is trained with a downsampled portion of the training dataset 330 corresponding to a subset of features selected to be modeled by the component classifier. Thus, a unique downsampled portion of the training dataset 330 is provided for training each component classifier 310. In other words, the downsampled portion of the training dataset 330 for training component classifier 311 is different from the downsampled portion of the training dataset 330 for training the other component classifiers 312-317. The trained ensemble classifier 516 comprises component classifiers 310, each of which has a selected subset of speech features, along with corresponding weighting coefficient values ​​for each feature generated from training the component classifiers 310 with the training dataset 330. The trained ensemble classifier 516 can be stored in any suitable memory and loaded into the user device to analyze new patient speech data in order to detect and / or predict cognitive decline in subjects, particularly mild cognitive impairment (MCI).

[0028] Figure 5 shows an exemplary apparatus 500 for detecting and / or predicting cognitive decline in a subject based on a sample of the subject's speech. The apparatus 500 utilizes a trained ensemble classifier 516, which is generated and trained by the method 200 described above. The apparatus 500 includes a voice input device 506 for receiving speech signals and generating data corresponding to a recording of the speech signals. For example, the voice input device 506 may include a microphone for capturing the subject's spoken speech. Furthermore, the apparatus 500 may include a voice output device 508 for generating speech output. For example, the voice output device 508 may include a speaker for providing commands for a WLR test in an audible manner to the subject. The apparatus 500 further includes a display 512 for generating visual output to the user. The apparatus 500 may further include an input / output device 510 for receiving data or commands to the apparatus 500 and / or transmitting data or commands from the apparatus 500. The device 500 further comprises a processor 502 and a computer-accessible medium 504. The processor 502 is operably connected to an audio output device 508 and a display 512, and controls the audio and visual outputs of the device 500. Furthermore, the processor 502 is operably connected to an audio input device 506, and receives and analyzes data corresponding to audio recordings captured by the audio input device 506. The processor 502 may include, for example, one or more microprocessors and may use instructions stored on the computer-accessible medium 504 (e.g., a memory storage device). The computer-accessible medium 504 may be, for example, a non-temporary computer-accessible medium containing executable instructions internally. The device 500 may further include a memory storage device 514 provided separately from the computer-accessible medium 504 for storing a trained ensemble classifier 516. The memory storage device 514 may be part of the device 500, or it may be outside the device 500 and operably connected to the device 500.

[0029] Figure 6 shows an exemplary method 600 for detecting and / or predicting cognitive decline in a subject based on a subject speech sample according to an exemplary embodiment of the present application. In particular, the method 600 obtains a WLR speech sample from the subject and determines whether the sample correlates more strongly with that of a normal patient or a cognitively impaired patient. In step 602, the subject is provided with a first set of instructions multiple times by a clinician or a processor 502 directing the audio output equipment 506 to operate the device 500 to audibly deliver the first set of instructions to the patient, and the processor 502 receives subject baseline speech data from the audio input equipment 506 corresponding to multiple audio recordings of the subject's speech in response to the first set of instructions. The subject is also provided with a second set of instructions by the clinician or the processor 502, and the processor 502 directs the audio output equipment 506 to audibly deliver the second set of instructions, and the processor 502 receives audio data from the audio input equipment 506 corresponding to the subject test speech data in accordance with the second set of instructions. The first instruction set can correspond to the instructions used in step 202 above to generate the training baseline speech data 112, and the second instruction set can correspond to the instructions used in step 202 to generate the training test speech data 114.

[0030] In step 604, the processor 502 analyzes and extracts multiple features from the subject baseline speech data and the subject test speech data in a manner similar to that described above with respect to step 204. In step 606, the processor 502 generates subject test data in a manner similar to that described above with respect to the training dataset in step 208, in which each feature of the subject test speech data is normalized with the corresponding feature of the subject baseline speech data. Similar to step 208, in one embodiment, the subject baseline speech data includes learning phase data from a WLR test (e.g., a RAVLT test) performed on the subject, and the subject test data is generated by normalizing the subject test speech data with the learning phase data from the WLR test. The subject test data obtained from this embodiment provides quantitative values ​​corresponding to the cognitive load of the subject. In step 608, the processor 502 analyzes the subject test data using a trained ensemble classifier 516. Each of the multiple component classifiers 310 of the trained ensemble classifier 516 analyzes the subject test data and generates a component output that identifies the subject test data as corresponding to a normal patient or a cognitively impaired patient. The ensemble module 320 of the trained ensemble classifier 516 receives component outputs from the component classifier 310, analyzes the component outputs, and generates an ensemble output that identifies subject test data as corresponding to normal patients or patients with cognitive impairment. Furthermore, the processor 502 can also generate an output indicating whether the subject is at high risk of neurodegeneration and / or likely to have cognitive impairment, based on the ensemble output generated from the analysis of the subject test data by the trained ensemble classifier 516. The output can indicate that the subject is at high risk of neurodegeneration and / or likely to have cognitive impairment if the ensemble output identifies the subject test data as corresponding to patients with cognitive impairment.In contrast, the output may indicate that the subject is not at high risk of neurodegeneration and / or is not likely to have cognitive impairment, if the ensemble output identifies the subject test data as corresponding to a normal patient. In step 610, the processor 502 directs the display 512 to provide a visual representation of the output.

[0031] Ensemble outputs generated from the analysis of subject test data by a trained ensemble classifier 516, and / or outputs generated by the apparatus 500, can be provided to clinicians to screen and identify patients at high risk of neurodegeneration, and / or patients with cognitive impairment, including but not limited to, impaired verbal episodic memory, slower learning speed, impaired short-term verbal memory, impaired delayed verbal memory, impaired recall after distracting stimuli, and impaired cognitive memory. Subjects identified by the subject test data as corresponding to patients with cognitive impairment through the ensemble outputs may be at high risk of developing neurodegenerative diseases such as AD or other types of dementia. Therefore, ensemble outputs generated from the analysis of subject test data, and / or outputs generated by the apparatus 500, can assist clinicians in directing patients in a critical condition to further cognitive tests to further confirm the outputs from the ensemble outputs and / or method 600. Thus, the apparatus and method of this application enable earlier identification of patients at risk of neurodegeneration, better care for these patients, and / or initiation of treatment at an earlier stage of cognitive decline, compared to conventional diagnoses of neurodegenerative diseases, dementia, or AD. Furthermore, if the ensemble output indicates that the subject test data corresponds to patients with cognitive impairment, and / or if the output generated by the device 500 indicates that the subject is at high risk of neurodegeneration and / or likely to have cognitive impairment, the processor 502 generates an instruction for administering treatment to the subject. In some embodiments, the instruction can automatically direct the administration of treatment without requiring intervention from an intervening user.

[0032] Suitable treatments include those for improving cognitive abilities and / or those for correcting cognitive decline. In one example, the treatment may include non-pharmacological treatments to improve the subject's cognitive abilities, such as digital treatments (e.g., brain training modules). In some embodiments, digital treatment can be automatically initiated when the ensemble output indicates that the subject's test data corresponds to a patient with cognitive decline, and / or the output generated by the device 500 indicates that the subject is at high risk of neurodegeneration and / or is likely to have cognitive decline. The brain training module includes a set of instructions that can be executed by a processor to administer brain training exercises to improve the subject's cognitive abilities. The brain training module may be connected to a user interface to display the brain training exercise instructions to the subject and to receive input from the user in response to such instructions.

[0033] In another example, treatment may include the administration of one or more drugs effective for preventing and / or improving the progression of neurodegeneration. In particular, treatment may include the administration of one or more drugs effective for preventing and / or improving the progression of dementia, or more specifically, for preventing and / or improving the progression of Alzheimer's disease (AD). Suitable drugs effective may include one or more drugs effective for preventing or reducing the aggregation of beta-amyloid or tau proteins in the brain, improving the resilience of brain synapses or cells, regulating the expression of the ApoE4 gene, regulating neuroinflammation-related pathways, etc. In one example, drugs effective may include antipsychotics, acetylcholinesterase inhibitors, etc.

[0034] Those skilled in the art will understand that the exemplary embodiments described herein can be implemented in any number of ways, such as as separate software modules or as a combination of hardware and software. For example, the exemplary methods may be embodied in one or more programs which include lines of code that are stored in a non-temporary storage medium and, when compiled, can be executed by one or more processor cores or separate processors. A system according to one embodiment comprises a plurality of processor cores and an instruction set that is executed on these plurality of processor cores to perform the exemplary methods described above. The processor cores or separate processors may be incorporated into or communicate with any suitable electronic device, such as an onboard processing unit inside the device, or an external processing unit that can communicate with at least a portion of the device, such as a mobile computing device, smartphone, computing tablet, computing device, etc. [Examples]

[0035] (Example I) In Example I, an exemplary word list recall test was administered manually by examiners (similar to those used in the RAVLT test) and also by computer to a total of 106 patients, 84 of whom were normal and 22 of whom had MCI. Each patient was provided with a first set of commands to listen to a first word list and immediately recall and pronounce it. The first word list included the following words: drum, helmet, music, coffee, school, parent, machine, garden, radio, farmer, nose, sailor, color, house, and river. Speech was recorded for each patient, and a set of baseline speech data corresponding to the speech recording was generated. The presentation of the first word list and its immediate recall were repeated a total of five times for the entire group of patients, generating five different baseline speech datasets.

[0036] Next, the patients were given a distraction task. Specifically, each patient was given a second set of commands to listen to a second word list and then reproduce and pronounce the second word list. The second word list included the following words: desk, ranger, bird, shoe, stove, mountain, glasses, towel, cloud, boat, lamb, bell, pencil, church, and fish. A speech was recorded for each patient, and a set of test speech data corresponding to the speech record for the distraction test was generated. After the distraction task, the patients were asked to reproduce and pronounce the first word list. A speech was recorded for each patient, and a set of test speech data corresponding to the speech record for the post-distraction test was generated. After a 20-minute delay, the patients were again asked to reproduce and pronounce the first word list. A speech was recorded for each patient, and a set of test speech data corresponding to the speech record for the delayed reproduction test was generated.

[0037] For example, exemplary sets of speech features, such as the mean and standard deviation of exemplary acoustic properties (as described above in Table 1) across speech frames of speech audio recordings, were extracted from each of the following sets of baseline speech data: a set of test speech data for the distraction test, a set of test speech data for the post-distraction test, and a set of test speech data for the delayed recall test. An exemplary ensemble classifier 700 of Example I is shown in Figure 7. As shown in Figure 7, the ensemble classifier 700 of Example I comprises 30 component classifiers 701-730, each of which is a support vector machine (SVM) generated based on a downsampled subset of the top 20 features determined based on a subsample of training data balanced between training data correlated with control patients and training data correlated with MCI patients. Specifically, the subsample includes 20 entries corresponding to control patients and 20 entries corresponding to MCI patients.

[0038] As shown in Figure 7, the ensemble classifier 700 is trained and validated by 10-fold cross-validation using a set of data 760, each containing features of the test speech data normalized by the mean of the corresponding features across five sets of baseline speech data. As shown in Figure 7, the data 760 is randomly distributed and divided into 10 equal-sized partitions. Of these 10 partitions, data from 9 partitions are used as training data 770 for the ensemble classifier 700, and the remaining partitions are used as validation data 780. Training and validation are repeated using another partition as validation data 780 and the remaining partition as training data 770 until each partition has been used once as validation data 780. Example I utilizes the 10-fold cross-validation method. It is conceivable that the ensemble classifier 700 could be validated using k-fold cross-validation, where k is any suitable positive integer.

[0039] Table 2 below shows the mean values ​​of specific performance measures generated based on validation data across each of the 10-segment sets for each of the interfered tests, post-interference tests, and delayed recall tests. As shown in the performance measures in Table 2, the post-interference test has the highest differential signal between MCI patients and normal patients, followed by the interfered test, and then the delayed recall test.

[0040] [Table 2]

[0041] The specific embodiments disclosed herein are intended to illustrate some aspects of the invention, and the invention described and claimed herein is not limited in scope by these embodiments. Any equivalent embodiments are intended to be within the scope of the invention. In fact, various modifications of the invention, in addition to those shown and described herein, will be apparent to those skilled in the art from the foregoing description. Such modifications are also intended to be within the scope of the accompanying claims. All publications cited herein are incorporated herein by reference in their entirety.

Claims

1. A control method using a system for detecting cognitive decline in a subject, wherein the system comprises a voice input device and a processor. The aforementioned method, The method involves acquiring subject baseline speech data corresponding to multiple audio recordings of the subject's speech in response to a first set of commands provided to the subject, using the voice input equipment, wherein the first set of commands corresponds to a first word list playback test. The method involves acquiring subject test speech data corresponding to further audio recordings of the subject's speech in response to a second set of commands provided to the subject using the voice input equipment, wherein the second set of commands corresponds to a first word list recall test and corresponds to a post-interruption test. The process involves using the aforementioned processor to extract a plurality of features from the subject baseline speech data and the subject test speech data, wherein the features correspond to the acoustic characteristics of the speech data. The method for generating subject test data is to use the aforementioned processor and normalize the subject test speech data using the subject baseline speech data, wherein the normalization of the subject test speech data is characterized by subtracting the average value of each of the features across the plurality of audio recordings of the speech that constitute the subject baseline speech data from the corresponding features of the subject test speech data. The method involves analyzing the subject test data using the aforementioned processor, characterized in that the analysis uses a trained ensemble classifier. Includes, The aforementioned trained ensemble classifier, Multiple component classifiers and ensemble modules Includes, Each of the plurality of component classifiers is configured to generate a component output that identifies the subject test data as corresponding to a normal patient or a patient with cognitive impairment, and each component classifier is configured to analyze a subset of features selected from the plurality of features. The ensemble module is configured to receive the component output from the component classifier and generate an ensemble output that identifies the subject test data as corresponding to normal patients or patients with cognitive impairment based on the component output. The subset of features of the trained ensemble classifier is selected separately and independently for each of the multiple component classifiers during training by the steps of ranking the multiple features for a subsample of training test speech data based on predetermined criteria and selecting the subset of features from the multiple features based on predetermined ranking thresholds. A method wherein the plurality of component classifiers are trained using training baseline speech data corresponding to prior audio recordings of speeches from groups of normal and cognitively impaired patients in response to a first set of instructions, and training test speech data corresponding to prior audio recordings of speeches from the same groups of normal and cognitively impaired patients in response to a second set of instructions.

2. The method according to claim 1, wherein the subsample comprises a first number of samples of the training test speech data corresponding to normal patients and a second number of samples of the training test speech data corresponding to patients with cognitive impairment.

3. The method according to claim 2, wherein the second number of samples is at least 80% of the first number of samples.

4. The method according to claim 3, wherein the ratio of the sample of the first number to the sample of the second number is 1:

1.

5. The method according to claim 2, wherein the aforementioned predetermined criterion is the importance of the feature.

6. The method according to claim 1, wherein the plurality of component classifiers are trained using a training dataset generated by normalizing training test speech data using training baseline speech data, the normalization of the training test speech data is characterized by subtracting the average value of each of the features across the plurality of preceding audio recordings of speech forming the training baseline speech data from the corresponding features of the training test speech data.

7. The method according to claim 1, wherein mild cognitive impairment (MCI) is detected in the subject, and the group of normal patients and cognitively impaired patients consists of normal patients and MCI patients.

8. The method according to claim 1, wherein each of the plurality of component classifiers includes a plurality of weighted feature coefficients trained using the training baseline speech data and the training test speech data.

9. The method according to claim 1, wherein each of the plurality of component classifiers is a machine learning classifier.

10. The method according to claim 9, wherein each of the plurality of component classifiers is a support vector machine (SVM).

11. The system further includes a display, and the method uses the display, When the ensemble output identifies the subject test data as corresponding to patients with cognitive impairment, an output indicating that the subject is likely to have cognitive impairment is displayed; and when the ensemble output identifies the subject test data as corresponding to patients with normal cognitive function, an output indicating that the subject is unlikely to have cognitive impairment is displayed. The method according to claim 1, further comprising:

12. When the ensemble output identifies the subject test data as corresponding to a patient with cognitive impairment, the ensemble output is used to generate an output indicating that the subject has impaired verbal episodic memory. The method according to claim 1, further comprising:

13. A device for detecting cognitive decline in subjects, Audio output equipment configured to generate audio output, A voice input device configured to receive an audio signal and generate data corresponding to the recording of the audio signal, The display and Processor, and The audio output device is directed to provide the subject with a first set of commands corresponding to the first word list playback test multiple times in an audible manner. Subject baseline speech data corresponding to multiple audio recordings of the subject's speech in response to the first set of commands is received from the voice input device. The audio output equipment is directed to provide the subject with an audible second set of commands corresponding to the first word list recall test and the post-interruption test. The system receives subject test speech data from the voice input device, corresponding to further audio recordings of the subject's speech in response to the second set of commands. From the subject baseline speech data and the subject test speech data, a plurality of features are extracted, wherein the features correspond to the acoustic characteristics of the speech data. Subject test data is generated by normalizing the subject test speech data using the subject baseline speech data, wherein the normalization of the subject test speech data is characterized by subtracting the average value of each of the features across the plurality of audio recordings of the speech that form the subject baseline speech data from the corresponding feature of the subject test speech data. The subject test data is analyzed using a trained ensemble classifier to generate an output indicating whether the subject is likely to have cognitive impairment. The display is directed to provide the user with a visual representation of the output. A non-temporary computer-readable storage medium containing an instruction set executable by the processor capable of operating in such a manner, A memory configured to store the aforementioned trained ensemble classifier and Equipped with, The aforementioned trained ensemble classifier, Multiple component classifiers and ensemble modules Equipped with, Each of the plurality of component classifiers is configured to generate a component output that identifies the subject test data as corresponding to a normal patient or a patient with cognitive impairment, and each component classifier is configured to analyze a subset of features selected from the plurality of features. The subset of features of the trained ensemble classifier is selected separately and independently for each of the multiple component classifiers during training by the steps of ranking the multiple features for a subsample of training test speech data based on predetermined criteria and selecting the subset of features from the multiple features based on predetermined ranking thresholds. The ensemble module is configured to receive the component output from the component classifier and generate an ensemble output that identifies the subject test data as corresponding to normal patients or patients with cognitive impairment based on the component output. The apparatus wherein the plurality of component classifiers are trained using training baseline speech data corresponding to prior audio recordings of speeches from groups of normal and cognitively impaired patients in response to a first set of commands, and training test speech data corresponding to prior audio recordings of speeches from the same groups of normal and cognitively impaired patients in response to a second set of commands.

14. The apparatus according to claim 13, wherein the subsample comprises a first number of samples of the training test speech data corresponding to normal patients and a second number of samples of the training test speech data corresponding to patients with cognitive impairment.

15. A computer-implemented method for training a classification system configured to detect cognitive decline in a subject based on a speech sample of the subject, The objective is to acquire training baseline speech data corresponding to audio recordings of speeches from groups of normal and cognitively impaired patients in response to a first set of commands corresponding to a first word list recall test, and training test speech data corresponding to audio recordings of speeches from groups of normal and cognitively impaired patients in response to a second set of commands corresponding to a post-interruption test. (i) Extract from the training baseline speech data and (ii) the training test speech data a plurality of features, wherein the features correspond to the acoustic characteristics of the speech data, To generate an ensemble classifier comprising multiple component classifiers and ensemble modules, Each of the aforementioned component classifiers is configured to generate a component output that identifies the sample data as corresponding to a normal patient or a patient with cognitive impairment, and each component classifier is configured to analyze a subset of features selected from the aforementioned features. The subset of features of the trained ensemble classifier is selected separately and independently during training for each of the multiple component classifiers by the steps of ranking the multiple features for a subsample of the training test speech data based on predetermined criteria and selecting the subset of features from the multiple features based on predetermined ranking thresholds. The ensemble module is configured to receive the component output from the component classifier and generate an ensemble output that identifies the sample data as corresponding to normal patients or patients with cognitive impairment based on the component output. To generate an ensemble classifier, A training dataset is generated by normalizing the training test speech data using the training baseline speech data, wherein the normalization of the training test speech data is characterized by subtracting the average value of each of the features across the plurality of preceding audio recordings of the speech forming the training baseline speech data from the corresponding features of the training test speech data. Training the ensemble classifier using the aforementioned training dataset A method implemented in a computer, including [this].

16. The computer-implemented method according to claim 15, wherein the subsample comprises a first number of samples of the training test speech data corresponding to normal patients and a second number of samples of the training test speech data corresponding to patients with cognitive impairment.

17. A system for training a classification system configured to detect cognitive decline in subjects based on the subjects' speech samples, A database configured to store training baseline speech data and training test speech data from groups of normal and cognitively impaired patients, wherein the training baseline speech data corresponds to audio recordings of speeches from the groups of normal and cognitively impaired patients in response to a first set of commands corresponding to a first word list recall test, and the training test speech data corresponds to audio recordings of speeches from the groups of normal and cognitively impaired patients in response to a second set of commands corresponding to the first word list recall test and a post-interruption test. It comprises a processor and a non-temporary computer-readable storage medium containing an instruction set that can be executed by the processor, which is operablely connected to communicate with the database, The baseline training speech data and the test training speech data are received from the database. (i) From the training baseline speech data and (ii) from the training test speech data, a plurality of features are extracted, wherein the features correspond to the acoustic characteristics of the speech data. To generate an ensemble classifier comprising multiple component classifiers and ensemble modules, Each of the aforementioned component classifiers is configured to generate a component output that identifies the sample data as corresponding to a normal patient or a patient with cognitive impairment, and each component classifier is configured to analyze a subset of features selected from the aforementioned features. The ensemble module is configured to receive the component output from the component classifier and generate an ensemble output that identifies the sample data as corresponding to normal patients or patients with cognitive impairment based on the component output. Generate an ensemble classifier, A subset of features for each of the plurality of component classifiers is selected separately and independently from the plurality of features by the step of ranking the plurality of features for a subsample of the training test speech data based on predetermined criteria, and selecting the subset of features from the plurality of features based on predetermined ranking thresholds. A training dataset is generated by normalizing the training test speech data using the training baseline speech data, wherein the normalization of the training test speech data is characterized by subtracting the average value of each of the features across the plurality of preceding audio recordings of the speech forming the training baseline speech data from the corresponding features of the training test speech data. The ensemble classifier is trained using the aforementioned training dataset. A computing device capable of operating in such a manner, A memory configured to store the aforementioned trained ensemble classifier, A system for training a classification system, equipped with the necessary features.

18. A system for training the classification system according to claim 17, wherein the subsamples consist of a first number of samples of the training test speech data corresponding to normal patients and a second number of samples of the training test speech data corresponding to patients with cognitive impairment.

Citation Information

Patent Citations

  • Cognitive dysfunction danger computing device, cognitive dysfunction danger computing system, and program

    JP2011255106A

  • Voice-Based Medical Assessment

    JP2020522028A

  • Selecting speech features for building models for detecting medical conditions

    US20180322894A1

  • Allosteric corticotropin-releasing factor receptor 1 (CRFR1) antagonists that decrease p-TAU and improve cognition

    WO2018048953A1