Alzheimer's disease diagnostic device and method

The Alzheimer's disease diagnosis device and method enhance diagnostic accuracy by analyzing speech data to generate classification results and cognitive function assessment scores, addressing the limitations of existing methods in early detection and management.

JP2026502148APending Publication Date: 2026-01-21EMOCOG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025536421
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-23
Filing Date
2023-12-05
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Existing methods for diagnosing Alzheimer's disease lack accuracy in early detection and management, particularly through analysis of speech characteristics.

Method used

An Alzheimer's disease diagnosis device and method that analyzes speech data to extract features using machine learning algorithms, generating classification results and cognitive function assessment scores based on these features.

Benefits of technology

Improves the accuracy of diagnosing Alzheimer's disease through early detection and management by analyzing speech characteristics, enabling precise prediction of cognitive function assessment scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502148000001_ABST
    Figure 2026502148000001_ABST
Patent Text Reader

Abstract

The present invention relates to an apparatus and method for diagnosing Alzheimer's disease based on analysis of a speaker's speech data. The Alzheimer's disease diagnosis method according to the present invention is a method for diagnosing Alzheimer's disease performed by a processor of an Alzheimer's disease diagnosis device, and includes the steps of collecting speech data of a speaker, extracting features of the speech data, and generating one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result based on the features of the speech data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus and method for diagnosing Alzheimer's disease based on analysis of a speaker's speech data. [Background technology]

[0002] Dementia can refer to a neurological condition in which cognitive abilities, memory, thinking ability, judgment, etc. decline due to a gradual decline in brain function. Dementia mainly occurs during the aging process, and as it progresses, it can seriously affect social, occupational, and daily life functioning.

[0003] Dementia has various forms and causes, with Alzheimer's disease being the most common. Alzheimer's disease is a chronic brain disease that occurs when damage to brain tissue reduces the function and connectivity of neurons (nerve cells). Symptoms of dementia initially begin with minor memory loss and confusion, and can gradually progress to a more severe decline in language ability, judgment, reasoning ability, and the ability to carry out daily activities.

[0004] The above-mentioned background art is technical information that the inventor possessed for the purpose of deriving the present invention or that he acquired in the process of deriving the present invention, and is not necessarily publicly known art that was made public to the general public prior to the filing of the present application. Summary of the Invention [Problem to be solved by the invention]

[0005] An object of the present invention is to analyze the speech characteristics of a speaker collected through a speech task and, based on the results, improve the accuracy of diagnosing Alzheimer's disease.

[0006] One object of the present invention is to analyze a speaker's voice characteristics collected through a speech task and accurately predict a cognitive function assessment score based on the analysis.

[0007] The problems to be solved by the present invention are not limited to those described above, and other unmentioned problems and advantages of the present invention can be understood from the following description and will be more clearly understood in the embodiments of the present invention. Furthermore, it will be understood that the problems to be solved and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims. [Means for solving the problem]

[0008] An Alzheimer's disease diagnosis device according to the present invention includes a processor and a memory operatively connected to the processor and storing at least one code executed by the processor, wherein the memory can store code that, when executed via the processor, causes the processor to collect speech data of a speaker, extract features of the speech data, and generate one or more of an Alzheimer's disease classification result and a predicted result of a cognitive function assessment score based on the features of the speech data.

[0009] The Alzheimer's disease diagnosis method according to the present invention is a method for diagnosing Alzheimer's disease performed by a processor of an Alzheimer's disease diagnosis device, and includes the steps of collecting speech data of a speaker, extracting features of the speech data, and generating one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result based on the features of the speech data.

[0010] In addition, other methods and systems for implementing the present invention, and computer-readable recording media storing computer programs for executing the methods can also be provided.

[0011] Other aspects, features, and advantages beyond those described above will become apparent from the following drawings, claims, and detailed description of the invention. [Effects of the Invention]

[0012] According to the present invention, the speech characteristics of a speaker collected through a speech task are analyzed, and based on this, the accuracy of diagnosing Alzheimer's disease is improved, which can be useful for the early diagnosis and management of Alzheimer's disease.

[0013] Furthermore, analyzing the speech characteristics of a speaker collected through a speech task and accurately predicting cognitive function assessment scores based on the analysis may be useful for the early diagnosis and management of Alzheimer's disease.

[0014] The effects of the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. [Brief explanation of the drawings]

[0015] [Figure 1] FIG. 1 is an exemplary diagram of an Alzheimer's disease diagnosis environment according to this embodiment. [Figure 2] FIG. 2 is a block diagram showing an outline of the configuration of the Alzheimer's disease diagnosis apparatus according to this embodiment. [Figure 3] FIG. 3 is a block diagram showing an outline of the configuration of the diagnosis management unit of the Alzheimer's disease diagnostic apparatus of FIG. [Figure 4] FIG. 4 is a configuration diagram for explaining the outline of the configuration of the diagnosis management unit shown in FIG. [Figure 5] FIG. 5 is a diagram illustrating an example of a speech task for collecting speech data of a speaker according to this embodiment. [Figure 6] FIG. 6 is a table comparing the characteristics of speech data between a group of Alzheimer's disease patients and a normal group according to this embodiment. [Figure 7] FIG. 7 is a table comparing features of voice data between Alzheimer's disease patients and healthy elderly people for each voice task according to this embodiment. [Figure 8] FIG. 8 is a table comparing the performance of the Alzheimer's disease classification results and the cognitive function assessment score prediction results according to this embodiment. [Figure 9]FIG. 9 is a block diagram showing an outline of the configuration of an Alzheimer's disease diagnosis apparatus according to another embodiment. [Figure 10-12] 10-12 are flowcharts for explaining the Alzheimer's disease diagnosis method according to this embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0016] The advantages and features of the present invention, as well as methods for achieving them, will become more apparent from the following detailed description of the embodiments accompanied by the accompanying drawings. However, the present invention is not limited to the embodiments presented below, and may be realized in different forms, and should be understood to include all modifications, equivalents, and alternatives within the spirit and technical scope of the present invention. The embodiments presented below are provided to fully disclose the present invention and to fully convey the scope of the invention to those skilled in the art to which the present invention pertains. In describing the present invention, if it is determined that a detailed description of related publicly known technology may hinder the gist of the present invention, such a detailed description will be omitted.

[0017] The terms used in this application are merely used to describe particular embodiments and are not intended to limit the present invention. The singular includes the plural unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "have" specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described herein, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. Terms such as "first," "second," and "second" can be used to describe various components, but the components should not be limited by these terms. These terms are used solely to distinguish one component from another.

[0018] Furthermore, in this application, a "module" may be a hardware component such as a processor or circuitry, and / or a software component executed by a hardware component such as a processor.

[0019] Hereinafter, an embodiment of the present invention will be described in detail with reference to the accompanying drawings. In the description with reference to the accompanying drawings, the same or corresponding components will be given the same drawing numbers, and duplicate descriptions will be omitted.

[0020] In the following embodiments, terms such as first and second are not used in a limiting sense but to distinguish one component from another.

[0021] In the following embodiments, singular expressions include plural expressions unless the context clearly indicates otherwise.

[0022] In the following embodiments, terms such as "comprise" or "have" mean the presence of features or components described in this specification, and do not preclude the possibility that one or more other features or components may be added.

[0023] Certain process sequences may be performed out of the order described when an embodiment is otherwise feasible, for example, two steps described as successive may be performed substantially simultaneously or in the reverse order from that described.

[0024] 1 is a diagram illustrating an example of an Alzheimer's disease diagnosis environment according to this embodiment. Referring to FIG. 1, the Alzheimer's disease diagnosis environment 1 may include an Alzheimer's disease diagnosis apparatus 100, a user terminal 200, and a network 300.

[0025] The Alzheimer's disease diagnosis apparatus 100 can collect speech data of a speaker from a speaker terminal (hereinafter referred to as a user terminal 200) provided by the speaker. The Alzheimer's disease diagnosis apparatus 100 can collect first to third speech data uttered by the speaker in response to first to third speech tasks. The first speech task in this embodiment can include a command requesting the speaker to respond to a predetermined question. Here, the first speech data can include speech data collected through an interview task. The second speech task in this embodiment can include a command requesting the speaker to output a reading result for a predetermined story and speak in accordance with the reading result. Here, the second speech data can include speech data collected through a repetition task. The third speech task can include a command requesting the speaker to recall a predetermined story. Here, the third speech data can include speech data collected through a recall task.

[0026] The Alzheimer's disease diagnosis apparatus 100 can extract features of collected voice data, i.e., first voice data to third voice data. The Alzheimer's disease diagnosis apparatus 100 can separate the collected voice data into speech segments and pause segments, and extract features of the voice data included in the speech segments. The Alzheimer's disease diagnosis apparatus 100 in this embodiment can extract one or more of frequency-related features, loudness-related features, temporal features, and spectrum features from the voice data included in the speech segments.

[0027] The Alzheimer's disease diagnosis apparatus 100 can generate one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result based on the extracted features of the voice data. The Alzheimer's disease diagnosis apparatus 100 in this embodiment can select features of the voice data before generating one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result. The Alzheimer's disease diagnosis apparatus 100 can use an analysis of variance (ANOVA) algorithm to select features of the voice data.

[0028] The Alzheimer's disease diagnosis device 100 can generate one or more of an Alzheimer's disease classification result and a predicted cognitive function assessment score using the feature selection results of the voice data. Here, the cognitive function assessment score can include, in one embodiment, an MMSE (mini-mental state examination) score. The MMSE is a standardized testing tool for assessing cognitive function and can be used to diagnose neurological and psychiatric disorders such as Alzheimer's disease. The MMSE is typically scored on a scale of 0 to 30, with higher scores indicating better cognitive function. A score of 27 to 30 points is within the normal range and indicates the absence of any cognitive impairment or problems. A score of 21 to 26 points indicates severe cognitive decline, which may indicate minor cognitive problems. A score of 11 to 20 points indicates moderate cognitive decline, which may indicate significant cognitive impairment. A score of 0 to 10 points indicates severe cognitive decline, which may indicate serious cognitive impairment or suspected dementia.

[0029] The Alzheimer's disease diagnosis device 100 can use an artificial intelligence algorithm to generate one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result. Here, artificial intelligence (AI) is a field of computer engineering and information technology that studies methods to enable computers to think, learn, self-develop, and the like that are possible with human intelligence, and can mean enabling computers to imitate human intelligent behavior.

[0030] In addition, AI does not exist on its own, but is directly or indirectly connected to many other fields of computer science. In particular, in modern times, there have been many active attempts to introduce AI elements into various fields of information technology and utilize them to solve problems in those fields.

[0031] Machine learning is a branch of artificial intelligence that includes research into giving computers the ability to learn without being explicitly programmed. Specifically, machine learning is the study and construction of systems and algorithms that learn, make predictions, and improve their performance based on empirical data. Rather than following strictly defined, static program instructions, machine learning algorithms can build specific models to derive predictions or decisions based on input data.

[0032] Both unsupervised and supervised learning methods can be used for machine learning in artificial neural networks. Deep learning, a type of machine learning, can learn at multiple levels based on data. Deep learning can represent a collection of machine learning algorithms that extract core data from multiple data sets.

[0033] In this embodiment, the Alzheimer's disease diagnosis apparatus 100 may exist independently in the form of a server, or the Alzheimer's disease diagnosis function provided by the Alzheimer's disease diagnosis apparatus 100 may be implemented in the form of an application and installed in the user terminal 200. Furthermore, the Alzheimer's disease diagnosis apparatus 100 may be a database server that provides data necessary for applying various artificial intelligence algorithms.

[0034] The user terminal 200 can access an Alzheimer's disease diagnostic application and / or an Alzheimer's disease diagnostic site provided by the Alzheimer's disease diagnostic apparatus 100 to receive an Alzheimer's disease diagnostic service.

[0035] Such a user terminal 200 may include a communication terminal capable of performing the functions of a computing device (not shown), and may be, but is not limited to, a desktop computer 201, a smartphone 202, a laptop 203 operated by a user, as well as a tablet PC, a smart TV, a mobile phone, a personal digital assistant (PDA), a media player, a microserver, a global positioning system (GPS) device, an e-book reader, a digital broadcasting terminal, a navigation system, a kiosk, an MP3 player, a digital camera, a home appliance, and other mobile or non-mobile computing devices. The user terminal 200 may also be a wearable terminal such as a watch, glasses, a hair band, or a ring equipped with communication and data processing functions. The user terminal 200 is not limited to the above, and any terminal capable of web browsing may be worn without limitation.

[0036] The network 300 can serve to connect the Alzheimer's disease diagnosis apparatus 100 and the user terminal 200. Such a network 300 can include, for example, wired networks such as a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or an integrated service digital network (ISDN), as well as wireless networks such as a wireless LAN (WLAN), code-division multiple access (CDMA), or satellite communication, but the scope of the present invention is not limited thereto. Furthermore, the network 300 can transmit and receive information using short-range communication and / or long-range communication. Here, short-range communication may include Bluetooth, RFID (radio frequency identification), IrDA (infrared data association), UWB (ultra-wideband), ZigBee, and Wi-Fi technologies, and long-range communication may include CDMA (code-division multiple access), FDMA (frequency-division multiple access), TDMA (time-division multiple access), OFDMA (orthogonal frequency-division multiple access), and SC-FDMA (single carrier frequency-division multiple access) technologies.

[0037] Network 300 may include connections of network elements such as hubs, bridges, routers, switches, etc. Network 300 may include one or more connected networks, e.g., a multi-network environment, including public networks such as the Internet and private networks such as secure corporate private networks. Access to network 300 may be provided via one or more wired or wireless access networks.

[0038] Additionally, network 300 may support CAN (controller area network) communications, V2I (vehicle to infrastructure) communications, V2X (vehicle to everything) communications, WAVE (wireless access in vehicular environment) communications technologies, and IoT (Internet of Things) networks for exchanging information between distributed components such as things and / or 5G communications.

[0039] Fig. 2 is a block diagram showing an outline of the configuration of the Alzheimer's disease diagnosis apparatus according to this embodiment. In the following description, parts that overlap with the description of Fig. 1 will be omitted. Referring to Fig. 2, the Alzheimer's disease diagnosis apparatus 100 can include a communication unit 110, a storage medium 120, a program storage unit 130, a database 140, a diagnosis management unit 150, and a control unit 160.

[0040] The communication unit 110 may provide a communication interface necessary for providing transmission and reception signals between the Alzheimer's disease diagnosis apparatus 100 and the user terminal 200 in the form of packet data in cooperation with the network 300. Furthermore, the communication unit 110 may serve to receive a predetermined information request signal from the user terminal 200 and to transmit information processed by the diagnosis management unit 150 to the user terminal 200. Here, the communication interface serves as a medium for connecting the Alzheimer's disease diagnosis apparatus 100 and the user terminal 200, and may include a path that provides a connection path so that the user terminal 200 can transmit and receive information after connecting to the Alzheimer's disease diagnosis apparatus 100. Furthermore, the communication unit 110 may be a device including hardware and software necessary for transmitting and receiving signals such as control signals and data signals to and from other network devices via wired or wireless connections.

[0041] The storage medium 120 temporarily or permanently stores data processed by the control unit 160. The storage medium 120 may include a magnetic storage medium or a flash storage medium, but the scope of the present invention is not limited thereto. The storage medium 120 may include an internal memory and / or an external memory, and may include a volatile memory such as a DRAM, SRAM, or SDRAM; a non-volatile memory such as a one-time programmable ROM (OTPROM), PROM, EPROM, EEPROM, mask ROM, flash ROM, NAND flash memory, or NOR flash memory; a flash drive such as an SSD, a compact flash (CF) card, an SD card, a Micro-SD card, a Mini-SD card, an XD card, or a memory stick; or a storage device such as an HDD.

[0042] The program memory unit 130 is equipped with control software that performs operations such as collecting the speaker's voice data, separating the collected voice data into speech sections and pause sections, extracting features of the voice data included in the speech sections, selecting features of specific voice data from the extracted features of the voice data, and using the selected features of the voice data to generate one or more of Alzheimer's disease classification results and cognitive function assessment score prediction results.

[0043] The database 140 may include an administrative database that stores various information for diagnosing Alzheimer's disease. The administrative database may store speech tasks for collecting voice data. The administrative database may store an algorithm (e.g., Voice Studio 2.0) for separating speech segments from collected voice data. The administrative database may store an algorithm (e.g., eGeMAPS) for extracting features of voice data. The administrative database may store an algorithm (e.g., ANOVA) for selecting features of voice data. The administrative database may store an artificial intelligence algorithm for generating one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result.

[0044] The database 140 may also include a user database that stores information about users (same as the above-mentioned speakers) who receive the Alzheimer's disease diagnostic service. Here, the user information may include basic information about the user, such as the user's name, affiliation, personal details, gender, age, contact information, email address, address, and image, as well as information about user authentication (login) such as ID (or email address) and password, and information about the connection, such as the country of connection, location of connection, information about the device used for connection, and the connected network environment.

[0045] The user database may store the first to third voice data collected from the user for the purpose of diagnosing Alzheimer's disease. The user database may also store user-specific information, information and / or category history provided by the user who connected to the Alzheimer's disease diagnostic application or Alzheimer's disease diagnostic site, environment setting information set by the user, resource usage information used by the user, and billing and settlement information corresponding to the user's resource usage.

[0046] The diagnosis management unit 150 can collect speech data of a speaker. The diagnosis management unit 150 can separate the collected speech data into speech segments and pause segments. The diagnosis management unit 150 can extract features of speech data included in the speech segments. The diagnosis management unit 150 can select features of specific speech data from the extracted features of speech data. The diagnosis management unit 150 can generate one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result using the selected features of speech data.

[0047] The control unit 160, as a type of central processing unit, can drive control software stored in the program storage unit 130 to control the overall operation of the Alzheimer's disease diagnosis apparatus 100. The control unit 160 can include any type of device capable of processing data, such as a processor. Here, the term "processor" may refer to a data processing device implemented in hardware, having a physically structured circuit for executing functions expressed by code or instructions included in a program. Examples of such data processing devices implemented in hardware include a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc., but the scope of the present invention is not limited thereto.

[0048] Fig. 3 is a block diagram for schematically explaining the configuration of the diagnostic management unit of the Alzheimer's disease diagnostic device of Fig. 2, Fig. 4 is a configuration diagram for roughly explaining the structure of the diagnostic management unit shown in Fig. 3, and Fig. 5 is an example diagram of a speech task for collecting speech data of a speaker according to this embodiment. In the following explanation, parts that overlap with the explanation of Figs. 1 and 2 will not be explained again. With reference to Figs. 3 to 7, the diagnostic management unit 150 can include a data collection unit 151, a first data processing unit 152, a second data processing unit 153, and a generation unit 154.

[0049] The data collection unit 151 can collect speech data of the speaker. The data collection unit 151 can collect first to third speech data uttered by the speaker in response to the first to third speech tasks. The first to third speech data collected by the data collection unit 151 may be stored in the database 140 as first to third speech files.

[0050] The data collection unit 151 can collect first speech data uttered by a speaker in response to a first speech task that requests the speaker's response to one or more preset questions. 510 in Fig. 5 shows examples of questions included in the first speech task, such as questions about age and date of birth, questions about educational background, questions about diet, questions about post-meal behavior, and questions about mood. In this embodiment, the content of the questions is not limited to the above examples and can be changed.

[0051] The data collection unit 151 can output a reading result of a preset story and collect second voice data spoken by the speaker in response to a second speech task that requests speech in accordance with the reading result. 520 in FIG. 5 shows examples of Shim Cheongjeong and Kongji Patch as preset stories included in the second and third speech tasks. In this embodiment, the content of the preset stories is not limited to the above examples and can be changed. In this embodiment, the reading result of the preset story can be output according to a delimiter (" / "), and second voice data spoken by the speaker in response to the second speech task that requests speech in accordance with the reading result can be collected.

[0052] The data collection unit 151 can collect third speech data uttered by the speaker in response to the third speech task that requires recollection of the above-mentioned story. In this embodiment, the data collection unit 151 can collect the third speech data of the speaker in response to the third speech task immediately after collecting the second speech data, i.e., without any delay time.

[0053] The first data processing unit 152 can extract features of the voice data from the first to third voice data of the speaker. In this embodiment, the first data processing unit 152 can include a 1-1 data processing unit 152-1 and a 1-2 data processing unit 152-2.

[0054] The 1-1 data processing unit 152-1 can separate the first to third voice data of the speaker into a speech section and a pause section.

[0055] The first-first data processing unit 152-1 may use an algorithm such as Voice Studio 2.0 to separate audio data into speech segments and pause segments. Voice Studio 2.0 may be a program that allows users to separate speech segments by clicking and dragging. After uploading an audio file to the program, users can drag to specify the speech segments they want to separate. Once separation of speech segments and pause segments for an audio file is complete, timestamps corresponding to the speech segments may be stored in a JSON file format. Based on the timestamps of the files, a Python library may be used to separate and store audio files corresponding to each speech segment.

[0056] The first-second data processing unit 152-2 can extract features of the first to third voice data included in the speech section, where the features of the voice data can include one or more of frequency-related features, intensity-related features, time features, and spectral features.

[0057] Frequency-related features relate to the pitch of the audio and may include fundamental frequency (F0), jitter, etc. Intensity-related features relate to the amplitude of the audio and may include volume and shimmer, etc. Temporal features relate to the speed of the audio and may include duration, etc. Spectral features indicate how much energy is present at various frequencies in audio with various amplitudes and frequencies and may include mel-frequency cepstral coefficients (MFCC), Alpha Ratio, and Hammarberg Index, etc.

[0058] MFCCs may be used for speech recognition and speech processing. MFCCs primarily use mel-scale frequency bands and can divide the frequency bands into mel filter banks. This allows for modeling the human auditory characteristics of speech. Alpha Ratios may be used to evaluate the frequency band energy distribution of speech. Alpha Ratios can calculate energy ratios by dividing a frequency band into multiple intervals. Hammarberg Indexes may be used to evaluate speech quality and pronunciation characteristics. Compared to MFCCs and Alpha Ratios, Hammarberg Indexes can use more general frequency bands and calculate the energy distribution of specific frequency bands.

[0059] In this embodiment, an algorithm such as the extended geneva minimalistic acoustic parameter set (eGeMAPS) can be used to extract features from speech data. Functions for calculating and extracting frequency, intensity, speech duration, and spectral features are built into the Python library opensmile, and features can be extracted by inputting the first to third speech data included in the speech section into this library. In this embodiment, by applying the eGeMAPS algorithm, 88 speech features can be extracted per speech section.

[0060] Here, the 88 speech features may include fundamental frequency-related derived variables, loudness-related derived variables, spatial characteristics-related variables, voice quality-related variables, formant-related variables, and speaking rate-related variables, as follows:

[0061] The derived variables related to fundamental frequency (F0) were F0semitoneFrom27.5Hz_sma3nz_amean (mean value of F0 (fundamental frequency) extracted from the semitone frequency scale based on 27.5Hz, indicating the pitch of the voice), F0semitoneFrom27.5Hz_sma3nz_stddevNorm (normalized standard deviation of F0, indicating the degree of variability of the pitch of the voice), F0semitoneFrom27.5Hz_sma3nz_percentile20.0 (20th percentile of F0), F0semitoneFrom27.5Hz_sma3nz_percentile50.0 (50th percentile of F0), F0se mitoneFrom27.5Hz_sma3nz_percentile80.0 (80th percentile of F0), F0semitoneFrom27.5Hz_sma3nz_pctlrange0-2 (range of 0-2% of total F0), F0semitoneFrom27.5Hz_sma3nz_meanRisingSlope (mean value of F0 rising slope), F0semitoneFrom27.5Hz_sma3nz_stddevRisingSlope (standard deviation of F0 rising slope), F0semitoneFrom27.5Hz_sma3nz_meanFallingSlope (mean value of F0 falling slope), F0semitoneFrom27.5Hz_sma3nz_meanFallingSlope5Hz_sma3nz_stddevFallingSlope (standard deviation of F0 falling slope), logRelF0-H1-H2_sma3nz_amean (average energy ratio of the first F0 harmonic (H1) to the highest harmonic energy at the second harmonic frequency (H2)), logRelF0-H1-H2_sma3nz_stddevNorm (ratio of the first F0 harmonic (H1) to the highest harmonic energy at the second harmonic frequency (H2)) 1), logRelF0-H1-A3_sma3nz_amean (mean energy ratio of the first F0 harmonic (H1) to the highest harmonic energy in the third formant range (A3)), and logRelF0-H1-A3_sma3nz_stddevNorm (standard deviation of the energy ratio of the first F0 harmonic (H1) to the highest harmonic energy in the third formant range (A3)).

[0062] Derived loudness-related variables include loudness_sma3_amean (mean volume of the audio signal, indicating the loudness of the voice), loudness_sma3_stddevNorm (standard deviation of the audio signal volume, indicating the variability of the voice loudness), loudness_sma3_percentile20.0 (20th percentile of the audio signal volume), loudness_sma3_percentile50.0 (50th percentile of the audio signal volume), loudness_sma3_percentile80.0 (80th percentile of the audio signal volume), and loudness_sma3_pc These can include tlrange0-2 (range of 0 to 2% of the total audio signal volume), loudness_sma3_meanRisingSlope (mean value of the audio signal volume rising slope), loudness_sma3_stddevRisingSlope (standard deviation of the audio signal volume rising slope), loudness_sma3_meanFallingSlope (mean value of the audio signal volume falling slope), loudness_sma3_stddevFallingSlope (standard deviation of the audio signal volume falling slope), and loudnessPeaksPerSec (number of peak values ​​of the audio signal volume measured per second).

[0063] The spatial characteristic related variables are spectralFlux_sma3_amean (mean value of spectral flux (an index showing the difference in frequency spectrum between consecutive audio frames, used to measure audio changes and energy fluctuations)), spectralFlux_sma3_stddevNorm (standard deviation of spectral flux), mfcc1_sma3_amean (MFCCs (Mel-Frequency Computed Coefficients), which is one of the main audio features showing the frequency history of the audio signal). mfcc1_sma3_stddevNorm (standard deviation of the first coefficient of the MFCC), mfcc2_sma3_amean (mean of the second coefficient of the MFCC), mfcc2_sma3_stddevNorm (standard deviation of the second coefficient of the MFCC), mfcc3_sma3_amean (mean of the third coefficient of the MFCC), mfcc3_sma3_stddevNorm (standard deviation of the third coefficient of the MFCC), mfcc4_sma3_amean (mean of the fourth coefficient of the MFCC), mfcc4_sma3_stddevNorm (standard deviation of the fourth coefficient of the MFCC), alphaRatioV_sma3nz_amean (alpha ratio (alpha) which means the energy distribution within the harmonic frequency band of the audio signal) ration), alphaRatioV_sma3nz_stddevNorm (standard deviation of alpha ratio), hammerbergIndexV_sma3nz_amean (mean of Hammarberg index, which is an index that measures the spectral characteristics of an audio signal and evaluates the flatness and energy distribution of the audio), hammerbergIndexV_sma3nz_stddevNorm (standard deviation of Hammarberg index), slopeV0-500_sma3nz_amean (mean of slope values ​​extracted from the frequency range from 0Hz to 500Hz), slopeV0-500_sma3nz_stddevNorm (standard deviation of slope values ​​extracted from the frequency range from 0Hz to 500Hz), slopeV500-1500_sma3nz_amean (mean of slope values ​​extracted from the frequency range from 500Hz to 1500Hz),slopeV500-1500_sma3nz_stddevNorm (standard deviation of slope values ​​extracted from the frequency range from 500Hz to 1500Hz), spectralFluxV_sma3nz_amean (spectrum plus mean value during speech periods in the given audio data), spectralFluxV_sma3nz_stddevNorm (standard deviation during speech periods in the given audio data), mfcc1V_sma3nz_amean (mean of the first coefficients of MFCCs during speech periods in the given audio data), mfcc1V_sma3nz_stddevNorm (standard deviation of the first coefficients of MFCCs during speech periods in the given audio data), mfcc2V_sma3nz_amean (mean of the second coefficients of MFCCs during speech periods in the given audio data), mfcc2V_sma3nz_stddevNorm (standard deviation of the second coefficients of MFCCs during speech periods in the given audio data), mfcc3V_sma 3nz_amean (mean of the third coefficient of the MFCCs in the speech sections of the given audio data), mfcc3V_sma3nz_stddevNorm (standard deviation of the third coefficient of the MFCCs in the speech sections of the given audio data), mfcc4V_sma3nz_amean (mean of the fourth coefficient of the MFCCs in the speech sections of the given audio data), mfcc4V_sma3nz_stddevNorm (standard deviation of the fourth coefficient of the MFCCs in the speech sections of the given audio data), alphaRatioUV_sma3nz_amean (mean of the alpha ratio in the speech sections of the given audio data), hammerbergIndexUV_sma3nz_amean (mean of the Hammarberg index in the speech sections of the given audio data), slopeUV0-500_sma3nz_amean (mean of the slope values ​​extracted from the frequency range from 0 Hz to 500 Hz in the speech sections of the given audio data),slopeUV500-1500_sma3nz_amean (the average of the slope values ​​extracted from the frequency range from 500 Hz to 1500 Hz in the voice section of the given audio data).

[0064] Voice quality related variables may include jitterLocal_sma3nz_amean (mean value of jitter, which indicates the degree of irregularity in frequency per speech cycle), jitterLocal_sma3nz_stddevNorm (standard deviation of jitter), shimmerLocaldB_sma3nz_amean (mean value of shimmer, which indicates the degree of irregularity in voice intensity per speech cycle), shimmerLocaldB_sma3nz_stddevNorm (standard deviation of shimmer), HNRdBACF_sma3nz_amean (mean value of Harmonic-to-Noise Ratio (HNR), which provides information on the overall noise characteristics of voice), and HNRdBACF_sma3nz_stddevNorm (standard deviation of HNR).

[0065] The formant-related variables were F1frequency_sma3nz_amean (mean of the first formant frequency), F1frequency_sma3nz_stddevNorm (standard deviation of the first formant frequency), F1bandwidth_sma3nz_amean (mean of the first formant bandwidth), F1bandwidth_sma3nz_stddevNorm (standard deviation of the first formant bandwidth), and F1amplitudeLogRelF0_sma3nz_amean (mean of the first formant bandwidth). F1amplitudeLogRelF0_sma3nz_stddevNorm (the mean of the amplitude of the first formant converted to a logarithm and expressed relative to F0), F1frequency_sma3nz_amean (the mean of the amplitude of the second formant frequency), F2frequency_sma3nz_stddevNorm (the standard deviation of the amplitude of the second formant frequency), F2bandwidth_sma3nz_amean (the mean of the amplitude of the second formant frequency), F2amplitudeLogRelF0_sma3nz_stddevNorm (mean of second formant bandwidth), F2bandwidth_sma3nz_stddevNorm (standard deviation of second formant bandwidth), F2amplitudeLogRelF0_sma3nz_amean (mean of second formant amplitudes converted to logarithms and expressed relative to F0), F2amplitudeLogRelF0_sma3nz_stddevNorm (standard deviation of second formant amplitudes converted to logarithms and expressed relative to F0), F3frequency_sma3nz_amea n (mean value of the third formant frequency), F3frequency_sma3nz_stddevNorm (standard deviation of the third formant frequency), F3bandwidth_sma3nz_amean (mean value of the third formant bandwidth), F3bandwidth_sma3nz_stddevNorm (standard deviation of the third formant bandwidth), F3amplitudeLogRelF0_sma3nz_amean (mean of the logarithmic value of the third formant amplitude expressed relative to F0),F3amplitudeLogRelF0_sma3nz_stddevNorm (standard deviation of logarithmic third formant amplitudes relative to F0) can be included.

[0066] Speech rate-related variables may include VoicedSegmentsPerSec (the number of voice segments in the audio data measured in 1-second increments), MeanVoicedSegmentLengthSec (the average value of the voice segments in the audio data measured in 1-second increments), StddevVoicedSegmentLengthSec (the standard deviation of the voice segments in the audio data measured in 1-second increments), MeanUnvoicedSegmentLength (the average value of the silent segments in the audio data measured in 1-second increments), and StddevUnvoicedSegmentLength (the standard deviation of the silent segments in the audio data measured in 1-second increments).

[0067] The second data processing unit 153 can select features of the voice data that satisfy preset conditions from the extraction results of the features of the voice data. In other words, the second data processing unit 153 can select features of the voice data to be used for generating the classification result of Alzheimer's disease and the prediction result of the cognitive function assessment score.

[0068] The second data processing unit 153 can use an analysis of variance (ANOVA) algorithm to select features of the voice data. In statistics, when two or more groups are to be compared with each other, the analysis of variance can include a method of performing hypothesis testing using an F distribution created by comparing the variance within a group, the grand mean, and the variance between groups caused by the difference in the mean of each group.

[0069] The second data processing unit 153 can analyze 88 correlations between the features of the voice data and the presence or absence of Alzheimer's disease through repeated measurements of analysis of variance, with the features of the voice data extracted by the first-second data processing unit 152-2 as independent variables and the presence or absence of Alzheimer's disease as the dependent variable. From the results of analyzing the 88 correlations, the second data processing unit 153 selects features of the voice data that are below a predetermined significance level (ρ<0.001), and can use these to generate Alzheimer's disease classification results and cognitive function assessment score prediction results.

[0070] The generation unit 154 may generate one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result based on the features of the voice data selected by the second data processing unit 153. In this embodiment, the generation unit 154 may simultaneously generate an Alzheimer's disease classification result and a cognitive function assessment score prediction result. The generation unit 154 may include a first generation unit 154-1 and a second generation unit 154-2.

[0071] The first generation unit 154-1 may generate a classification result of Alzheimer's disease corresponding to the features of the voice data that are below the predetermined significance level, using a first deep neural network model that has been pre-trained to classify the presence or absence of Alzheimer's disease according to the features of the voice data. Here, the first deep neural network model may be a model that has been trained by a supervised learning method using training data that uses the features of the voice data as input and labels the presence or absence of Alzheimer's disease.

[0072] The first generation unit 154-1 can train an initially set first deep neural network model in a supervised learning manner using the labeled training data. Here, the initially set first deep neural network model is an initial model designed to be configured as a model capable of classifying the presence or absence of Alzheimer's disease, and parameter values ​​are set to arbitrary initial values. The initial model is trained using the above training data, and the parameter values ​​are optimized, so that it can be completed as a classification model that can accurately classify the presence or absence of Alzheimer's disease.

[0073] Specifically, the first generator 154-1 is a weighted voting classifier that combines random forest and logistic regression to classify the presence or absence of Alzheimer's disease based on the features of speech data. Weights may be added depending on the performance of the classifier. Hyperparameters may be adjusted using training data through 10-fold cross-validation, and the resulting model may be evaluated using test data. Accuracy, precision, sensitivity, specificity, and F1-score may be measured to evaluate and compare performance with other speech tasks.

[0074] The second generation unit 154-2 may generate a predicted result of the cognitive function assessment score corresponding to the features of the voice data that are below a predetermined significance level, using a second deep neural network model that has been pre-trained to predict the cognitive function assessment score corresponding to the features of the voice data. Here, the second deep neural network model may be a model that has been trained by a supervised learning method using training data that uses the features of the voice data as input and the cognitive function assessment score as labels.

[0075] The second generation unit 154-2 may train an initially set second deep neural network model in a supervised learning manner using the labeled training data. Here, the initially set second deep neural network model is an initial model designed to be configured as a model capable of predicting cognitive function assessment scores, and parameter values ​​are set to arbitrary initial values. The initial model is trained using the above training data to optimize the parameter values, and may be completed as a classification model capable of accurately predicting cognitive function assessment scores.

[0076] Specifically, the second generation unit 154-2 may be constructed using the same methodology as the first generation unit 154-1. The second generation unit 154-2 is a weighted regression model, and may be constructed by combining a random forest and a support vector machine.

[0077] In this embodiment, the generation unit 154 can generate three Alzheimer's disease classification results and three predicted results of cognitive function assessment scores. The generation unit 154 can generate a first Alzheimer's disease classification result and a predicted result of a first cognitive function assessment score based on the feature selection result of the speech section of the first voice data. The generation unit 154 can generate a second Alzheimer's disease classification result and a predicted result of a second cognitive function assessment score based on the feature selection result of the speech section of the second voice data. The generation unit 154 can generate a third Alzheimer's disease classification result and a predicted result of a third cognitive function assessment score based on the feature selection result of the speech section of the third voice data.

[0078] In an optional embodiment, the generation unit 154 may generate a classification result of Alzheimer's disease based on the feature selection results of the speech sections of the first speech data and the second speech data. The generation unit 154 may determine whether to generate a prediction result of a cognitive function assessment score based on the classification result of Alzheimer's disease. In this embodiment, the generation unit 154 may determine to generate a prediction result of a cognitive function assessment score when the classification result of Alzheimer's disease for one or more of the features of the first speech data and the second speech data is generated as a first value ("1", Alzheimer's disease). However, the generation unit 154 may also determine not to generate a prediction result of a cognitive function assessment score when the classification result of Alzheimer's disease for the features of the first speech data and the second speech data is generated as a second value ("0", normal). When it is decided to generate a predicted result of the cognitive function assessment score, the generation unit 154 can generate a predicted result of the cognitive function assessment score based on the feature selection result for the speech section of the third audio data excluding the features of the first audio data and the features of the second audio data.

[0079] 6 is a table comparing features of speech data between a group of Alzheimer's disease patients and a normal group according to this embodiment. FIG. 6 may include experimental results comparing features using speech data collected from multiple Alzheimer's disease patients and multiple healthy older adults as a normal group. Referring to FIG. 6, AD (Alzheimer's disease) indicates Alzheimer's disease patients, HC (healthy older adults) indicates healthy older adults, and ρ indicates a statistical correlation coefficient.

[0080] First, for frequency-related features, the mean F0 (AD = 30.14 ± 3.13, HC = 31.6 ± 3.94, ρ < 0.001), the 20th percentile F0 (AD = 26.31 ± 3.15, HC = 28.55 ± 4.98, ρ < 0.001), and the 50th percentile F0 (AD = 30.16 ± 3.35, HC = 31.32 ± 4.02, ρ = 0.011) indicate speech pitch. In all cases, the values ​​of AD patients were lower than those of healthy older adults. The F0 percentile range indicated a larger F0 range in AD patients than in healthy older adults (AD = 7.4 ± 1.77, HC = 5.73 ± 3.36, ρ < 0.001), and the higher F0 standard deviation indicated a higher F0 variability in AD patients than in healthy older adults (AD = 0.19 ± 0.04, HC = 0.16 ± 0.06, ρ < 0.001). AD patients also showed higher normal periodicity (jitter) than healthy older adults (AD = 0.04 ± 0.01, HC = 0.03 ± 0.02, ρ < 0.001), indicating greater pitch variability. In summary, frequency-related features indicated that AD patients had lower pitch and greater pitch variability in their speech compared to healthy older adults.

[0081] The Alzheimer's disease patients had the following characteristics regarding speech intensity: lower 80th percentile intensity (AD = 0.55 ± 0.27, HC = 0.68 ± 0.28, ρ < 0.001), lower 50th percentile intensity (AD = 0.22 ± 0.14, HC = 0.27 ± 0.12, ρ = 0.008), and lower mean intensity (AD = 0.33 ± 0.14, HC = 0.37 ± 0.14, ρ = 0.012). The intensity peak per second (AD = 2.37 ± 0.57, HC = 2.82 ± 0.59, ρ < 0.001), intensity percentile range (AD = 0.46 ± 0.25, HC = 0.62 ± 0.28, ρ < 0.001), intensity decline slope (AD = 4.6 ± 2.25, HC = 5.73 ± 2.62, ρ < 0.001), and intensity rise slope (AD = 5.33 ± 2.61, HC = 6.55 ± 2.82, ρ < 0.001) all indicate intensity variability and were reduced in AD patients. Intensity-related features indicate that AD patients produce speech with reduced intensity and more monotonous sounds compared with normal elderly subjects.

[0082] In terms of temporal characteristics, AD patients had shorter durations of speech segments compared with healthy older adults (AD = 0.3 ± 0.11, HC = 0.37 ± 0.14, ρ < 0.001) and a greater number of speech segments compared with healthy older adults (AD = 2.06 ± 0.4, HC = 1.79 ± 0.82, ρ < 0.001). This suggests that AD patients struggle with word continuity, speak at a slower pace, and show more pauses and hesitations than healthy older adults.

[0083] FIG. 7 is a table comparing features of voice data between Alzheimer's disease patients and healthy elderly people for each voice task according to this embodiment.

[0084] Referring to Figure 7, certain trends in frequency-related features were observed in the speech of Alzheimer's disease patients, regardless of the speech task. In all speech tasks, Alzheimer's disease patients exhibited lower speech pitch (e.g., F0 mean, 20th, and 50th percentiles) and greater pitch variability (F0 percentile range, normalized F0 standard deviation). However, in the interview (first speech data) and repetition task (second speech data), these trends were not sufficiently discernible to be statistically significant. However, in the recollection task (third speech data), most features showed statistical significance. Furthermore, in the analysis results regardless of the previous speech task, there was no significant difference between the F0 ascending slope and the F0 descending slope. However, after separating the data by speech task, these features were found to be significantly different in the recollection task.

[0085] Consistent trends were also observed for loudness-related features across the audio tasks. In all audio tasks, Alzheimer's disease patients produced lower intensities (mean intensity, 20th, 50th, and 80th percentiles) and more monotonous sounds (intensity percentile range, intensity peaks per second, intensity rise slope, and intensity fall slope). However, in the interview (first audio data) and repetition task (second audio data), most intensity-related features showed no statistically significant differences. In the recollection task (third audio data), on the other hand, sufficient distinction was observed to demonstrate statistical significance.

[0086] Regarding temporal features, the memory task did not show good discrimination: Alzheimer's disease patients had shorter durations of vocalizations and more transient interruptions in the interview (first audio data) and the recollection task (third audio data).

[0087] FIG. 8 is a table comparing the performance of the Alzheimer's disease classification results and the cognitive function assessment score prediction results according to this embodiment.

[0088] Referring to FIG. 8, 810 compares the performance of the first deep neural network model (Alzheimer's disease classification model) on the first audio data (interview task), the second audio data (repetition task), and the third audio data (recollection task). The recollection task (third audio data) achieved an accuracy of 81.4%, outperforming all other audio tasks, including the interview (first audio data) and the repetition task (second audio data). The repetition task (second audio data) achieved the second highest accuracy of 78.5%, and the interview task (first audio data) achieved an accuracy of 76.1%. The first deep neural network model uses a weighted evaluation classifier that combines logistic regression and random forests, with the best results shown in bold.

[0089] 820 also compares the performance of the second deep neural network model (a cognitive function assessment (e.g., MMSE) score prediction model) for the first audio data (interview task), the second audio data (repetition task), and the third audio data (recall task). MAE (mean absolute error) and RMSE (root mean square error) are prediction error measurement indices; the lower the indices, the higher the predictability; the best results are shown in bold.

[0090] We investigated the acoustic characteristics of the speech of AD patients and healthy elderly people in Figures 6–8 and examined how these characteristics changed under different speech tasks and memory loads. We found that AD patients possess speech characteristics that distinguish them from healthy elderly people, such as lower pitch, increased pitch variability, lower volume, monotonous volume, and slower speech rate. Furthermore, we demonstrated that the addition of high memory load may accentuate the speech characteristics of Alzheimer's disease. In particular, we demonstrated that the acoustic characteristics of AD speech may vary significantly depending on the speech task, with the most significant differences observed in the recollection task involving high memory load. Furthermore, the recollection task with high memory load performed better than other tasks with low memory load in the Alzheimer's disease classification model and cognitive function assessment score prediction model.

[0091] Therefore, the Alzheimer's disease diagnosis device 100 according to this embodiment can be useful for the early diagnosis and management of Alzheimer's disease by analyzing the voice characteristics of a speaker and accurately predicting the diagnostic accuracy of Alzheimer's disease and the cognitive function assessment score based on the results.

[0092] 9 is a block diagram showing an outline of the configuration of an Alzheimer's disease diagnosis apparatus according to another embodiment. In the following description, parts that overlap with the description of FIGS. 1 to 8 will be omitted. Referring to FIG. 9, the Alzheimer's disease diagnosis apparatus 100 according to another embodiment can include a processor 170 and a memory 180.

[0093] In this embodiment, the processor 170 can process the functions performed by the communication unit 110, storage medium 120, program storage unit 130, database 140, diagnosis management unit 150, and control unit 160 disclosed in Figures 2 and 3.

[0094] Such a processor 170 can control the overall operation of the Alzheimer's disease diagnosis apparatus 100. Here, the term "processor" can refer to a data processing device built into hardware, having a circuit physically structured to execute functions expressed by codes or instructions included in a program. Examples of such data processing devices built into hardware include a microprocessor, a central processing unit, a processor core, a multiprocessor, an ASIC, an FPGA, and the like, but the scope of the present invention is not limited thereto.

[0095] The memory 180 is operatively connected to the processor 170 and is capable of storing at least one code related to operations to be performed by the processor 170 .

[0096] Furthermore, memory 180 may temporarily or permanently store data processed by processor 170 and may include data constructed in database 140. Here, memory 180 may include a magnetic storage medium or a flash storage medium, although the scope of the present invention is not limited thereto. Such memory 180 may include internal memory and / or external memory, and may include storage devices such as volatile memory such as DRAM, SRAM, or SDRAM; non-volatile memory such as OTPROM, PROM, EPROM, EEPROM, mask ROM, flash ROM, NAND flash memory, or NOR flash memory; flash drives such as SSDs, CF cards, SD cards, Micro-SD cards, Mini-SD cards, xD cards, or Memory Sticks; or HDDs.

[0097] 10 is a flowchart illustrating a method for diagnosing Alzheimer's disease according to one embodiment. In the following description, parts that overlap with the description of FIGS. 1 to 9 will be omitted. The method for diagnosing Alzheimer's disease according to this embodiment will be described on the assumption that the Alzheimer's disease diagnosis apparatus 100 executes the method using the processor 170 with the help of peripheral components.

[0098] Referring to FIG. 10 , in step S1010, processor 170 may collect speech data of a speaker. Processor 170 may collect first speech data spoken by the speaker in response to a first speech task requiring the speaker to respond to one or more predetermined questions. Here, the first speech data may include speech data collected through an interview task. Processor 170 may output a reading result of a predetermined story and collect second speech data spoken by the speaker in response to a second speech task requiring the speaker to speak in accordance with the reading result. Here, the second speech data may include speech data collected through a repetitive task. Processor 170 may collect third speech data spoken by the speaker in response to a third speech task requiring the speaker to recall the above-mentioned predetermined story. Here, the third speech data may include speech data collected through a recall task.

[0099] In step S1020, processor 170 may extract features of the collected audio data. Here, before extracting the features of the audio data, processor 170 may separate the collected audio data into speech segments and pause segments. Processor 170 may extract one or more of frequency-related features, intensity-related features, temporal features, and spectral features from the audio data included in the separated speech segments. In any embodiment, after extracting the features of the audio data, processor 170 may select the features of the audio data to be used for generating Alzheimer's disease classification results and cognitive function assessment score prediction results. To select the features of the audio data, processor 170 may analyze the correlation between the features of the audio data and the presence or absence of Alzheimer's disease through repeated measures analysis of variance, with the features of the audio data as independent variables and the presence or absence of Alzheimer's disease as the dependent variable. Processor 170 may select the features of the audio data for which the correlation analysis result is less than a predetermined significance level (e.g., ρ<0.001).

[0100] In step S1030, the processor 170 may generate one or more of an Alzheimer's disease classification result and a cognitive function assessment score prediction result based on the features of the selected audio data. The processor 170 may simultaneously generate an Alzheimer's disease classification result and a cognitive function assessment score prediction result corresponding to the features of the audio data that are below a predetermined significance level. The processor 170 may generate an Alzheimer's disease classification result corresponding to the features of the audio data that are below a predetermined significance level using a first deep neural network model pre-trained to classify the presence or absence of Alzheimer's disease corresponding to the features of the audio data. Here, the first deep neural network model may be a model trained in a supervised learning manner using training data that uses the features of the audio data as input and labels the presence or absence of Alzheimer's disease. The processor 170 may generate a cognitive function assessment score prediction result corresponding to the features of the audio data that are below a predetermined significance level using a second deep neural network model pre-trained to predict a cognitive function assessment score corresponding to the features of the audio data that are below a predetermined significance level. Here, the second deep neural network model may be a model trained by a supervised learning method using training data that uses features of speech data as input and cognitive function assessment scores as labels.

[0101] Furthermore, the processor 170 can provide one or more of the Alzheimer's disease classification result and the predicted result of the cognitive function assessment score to the user terminal 200. When providing one or more of the Alzheimer's disease classification result and the predicted result of the cognitive function assessment score to the user terminal 200, instructions (such as exercise and dietary habits) that the user can implement to prevent Alzheimer's disease can be provided.

[0102] 11 is a flowchart illustrating a method for diagnosing Alzheimer's disease according to another embodiment. In the following description, parts that overlap with the description of FIGS. 1 to 10 will be omitted. The method for diagnosing Alzheimer's disease according to this embodiment will be described on the assumption that the Alzheimer's disease diagnosis apparatus 100 executes the method using the processor 170 with the help of peripheral components.

[0103] Referring to FIG. 11, in step S1110, processor 170 can collect first speech data spoken by the speaker in response to a first speech task, second speech data spoken by the speaker in response to a second speech task, and third speech data spoken by the speaker in response to a third speech task.

[0104] In step S1120, processor 170 can separate each of the first, second, and third audio data into a speech section and a pause section.

[0105] In step S1130, the processor 170 can extract features of the audio data, including one or more of a frequency-related feature, an intensity-related feature, a temporal feature, and a spectral feature, for each of the first, second, and third audio data included in the speech section. As an example, the processor 170 can extract 88 features of the audio data for each of the speech sections of the first, second, and third audio data.

[0106] In step S1140, processor 170 may select features of the first, second, and third voice data included in the speech section. To select the features of the voice data, processor 170 may analyze the correlation between the features of the first, second, and third voice data and the presence or absence of Alzheimer's disease through repeated measures analysis of variance, with the features of the first, second, and third voice data as independent variables and the presence or absence of Alzheimer's disease as the dependent variable. From the results of the correlation analysis, processor 170 may select features of the first, second, and third voice data that are less than a predetermined significance level (e.g., ρ<0.001).

[0107] In step S1150, the processor 170 can use the feature selection results of the first, second, and third audio data to generate classification results for the first, second, and third Alzheimer's disease, and can use the feature selection results of the first, second, and third audio data to generate prediction results for the first, second, and third cognitive function assessment scores.

[0108] 12 is a flowchart illustrating a method for diagnosing Alzheimer's disease according to yet another embodiment. In the following description, parts that overlap with the description of FIGS. 1 to 11 will be omitted. The method for diagnosing Alzheimer's disease according to this embodiment will be described on the assumption that the Alzheimer's disease diagnosis apparatus 100 executes the method using the processor 170 with the help of peripheral components.

[0109] Referring to FIG. 12, in step S1201, the processor 170 can collect first speech data spoken by a speaker in response to a first speech task, and second speech data spoken by a speaker in response to a second speech task.

[0110] In step S1203, the processor 170 can separate the first and second audio data into speech segments and pause segments.

[0111] In step S1205, the processor 170 can extract features of the first and second audio data included in the speech section. The processor 170 can extract audio data features including one or more of frequency-related features, intensity-related features, temporal features, and spectral features for each of the first and second audio data included in the speech section. As an example, the processor 170 can extract 88 audio data features for each of the speech sections of the first and second audio data.

[0112] In step S1207, processor 170 may select features of the first and second voice data. To select the features of the voice data, processor 170 may analyze the correlation between the features of the first and second voice data and the presence or absence of Alzheimer's disease through repeated measures analysis of variance, with the features of the first and second voice data as independent variables and the presence or absence of Alzheimer's disease as the dependent variable. From the results of the correlation analysis, processor 170 may select features of the first and second voice data that are less than a predetermined significance level (e.g., ρ<0.001).

[0113] In step S1209, processor 170 can generate an Alzheimer's disease classification result using the feature selection results of the first and second speech data. Processor 170 can simultaneously generate an Alzheimer's disease classification result and a predicted cognitive function assessment score corresponding to speech data features below a predetermined significance level. Processor 170 can generate an Alzheimer's disease classification result corresponding to speech data features below a predetermined significance level using a first deep neural network model pre-trained to classify the presence or absence of Alzheimer's disease in response to speech data features. Here, the first deep neural network model can be a model trained using a supervised learning method using training data that uses speech data features as input and labels indicating the presence or absence of Alzheimer's disease.

[0114] In step S1211, processor 170 can determine whether the Alzheimer's disease classification result generated using the feature selection results of the first and second voice data is a first value. Here, the first value can indicate Alzheimer's disease. Processor 170 can terminate if all of the Alzheimer's disease classification results generated using the feature selection results of the first and second voice data are not the first value, i.e., if they are the second value (normal).

[0115] In step S1213, if the classification result of Alzheimer's disease generated using the feature selection results of the first and second voice data is a first value, processor 170 can collect third voice data uttered by the speaker in response to the third speech task. In this embodiment, if the classification result of Alzheimer's disease generated using the feature selection results of one or more of the features of the first voice data and the second voice data is a first value, processor 170 can collect third voice data uttered by the speaker in response to the third speech task.

[0116] In step S1215, the processor 170 may separate the third audio data into speech segments and pause segments.

[0117] In step S1217, the processor 170 can extract features of the third audio data included in the speech section. The processor 170 can extract features of the audio data including one or more of a frequency-related feature, an intensity-related feature, a temporal feature, and a spectral feature for each piece of the third audio data included in the speech section. As an example, the processor 170 can extract 88 features of the audio data for each piece of the speech section of the third audio data.

[0118] In step S1219, processor 170 may select features of the third voice data. To select the features of the voice data, processor 170 may analyze the correlation between the features of the third voice data and the presence or absence of Alzheimer's disease through repeated measures analysis of variance, with the features of the third voice data as independent variables and the presence or absence of Alzheimer's disease as the dependent variable. From the results of the correlation analysis, processor 170 may select features of the third voice data that are less than a predetermined significance level (e.g., ρ<0.001).

[0119] In step S1221, processor 170 can generate a predicted result of the cognitive function assessment score based on the feature selection result of the third voice data. Processor 170 may generate a predicted result of the cognitive function assessment score corresponding to the features of the third voice data that are below a predetermined significance level using a second deep neural network model pre-trained to predict cognitive function assessment scores corresponding to the features of the voice data. Here, the second deep neural network model may be a model trained using a supervised learning method with training data that uses the features of the voice data as input and the cognitive function assessment scores as labels.

[0120] Compared with Figures 10 and 11, Figure 12 enables earlier diagnosis of Alzheimer's disease. In Figures 10 and 11, features of the first to third speech data are extracted and selected to generate an Alzheimer's disease classification result and a cognitive function assessment score prediction result. On the other hand, in Figure 12, features of the first and second speech data are extracted and selected, and only after classification as Alzheimer's disease is it predicted that the speaker is an Alzheimer's disease patient, and features of the third speech data are extracted and selected to generate a cognitive function assessment score prediction result, thereby enabling earlier diagnosis of Alzheimer's disease.

[0121] The above-described embodiments of the present invention may be realized in the form of a computer program executable by various components on a computer, and such a computer program may be recorded on a computer-readable medium, which may include a hardware device specially configured to store and execute program instructions, such as a magnetic medium such as a hard disk, floppy disk, or magnetic tape, an optical medium such as a CD-ROM or DVD, a magneto-optical medium such as a floptical disk, a ROM, a RAM, or a flash memory.

[0122] On the other hand, the computer programs may be those specially designed and constructed for the purposes of the present invention, or they may be those well known and available to those skilled in the computer software art. Examples of computer programs include not only machine code, such as produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter, etc.

[0123] The use of the term "said" and similar indicators in the present specification (particularly in the claims) can correspond to both the singular and the plural. Furthermore, when a range is stated in the present invention, it is deemed to include inventions to which individual values ​​belonging to said range are applied (unless otherwise specified), and is the same as if each individual value constituting said range were stated in the detailed description of the invention.

[0124] Unless explicitly stated or stated to the contrary, steps constituting a method according to the present invention may be performed in any suitable order. The order in which the steps are described is not necessarily intended to limit the scope of the present invention. The use of all examples or exemplary terms (e.g., etc.) in the present invention is merely for the purpose of explaining the present invention in detail, and the scope of the present invention is not limited by such examples or exemplary terms unless otherwise limited by the claims. Furthermore, those skilled in the art will recognize that various modifications, combinations, and variations can be made within the scope of the claims or their equivalents, depending on design conditions and factors.

[0125] Therefore, the concept of the present invention should not be limited to the above-described embodiments, and all scopes equivalent to or modified equivalently from the scope of the claims, as well as the scope of the claims described below, can be said to fall within the scope of the concept of the present invention.

Claims

1. 1. A method for diagnosing Alzheimer's disease executed by a processor of an Alzheimer's disease diagnosis apparatus, comprising: collecting speech data of a speaker; extracting features from the speech data; generating one or more of a classification result of Alzheimer's disease and a prediction result of a cognitive function assessment score based on the features of the speech data; How to diagnose Alzheimer's disease.

2. The collecting step is collecting first speech data uttered by a speaker in response to a first speech task that requests a response from the speaker to one or more predetermined questions; a step of outputting a reading result of a preset story, and collecting second speech data uttered by the speaker in response to a second speech task that requests the speaker to speak in accordance with the reading result; and collecting third speech data uttered by the speaker in response to a third speech task requiring the speaker to recall a story. The method for diagnosing Alzheimer's disease according to claim 1.

3. Prior to the step of extracting features from the audio data, further comprising a step of separating the voice data of the speaker into speech segments and pause segments; The step of extracting features of the speech data includes: extracting one or more of frequency-related features, intensity-related features, temporal features, and spectral features from speech data included in a speech section; The method for diagnosing Alzheimer's disease according to claim 1.

4. After the step of extracting features from the audio data, The method further includes selecting features of the speech data to be used in generating a classification result of Alzheimer's disease and a prediction result of the cognitive function assessment score. The method for diagnosing Alzheimer's disease according to claim 1.

5. The step of selecting features of the audio data includes: loading features of speech data included in a speech section separated from speech data of a speaker; analyzing the correlation between the features of the speech data and the presence or absence of Alzheimer's disease by repeated measures analysis of variance, with the features of the speech data as independent variables and the presence or absence of Alzheimer's disease as a dependent variable; and selecting, from the results of the correlation analysis, features of the speech data that are less than a predetermined significance level. The method for diagnosing Alzheimer's disease according to claim 4.

6. The generating step includes: generating a classification result of Alzheimer's disease corresponding to the features of the speech data that are below a predetermined significance level, using a first deep neural network model that has been pre-trained to classify the presence or absence of Alzheimer's disease corresponding to the features of the speech data; generating a predicted result of the cognitive function assessment score corresponding to the features of the voice data that are less than the predetermined significance level using a second deep neural network model that has been pre-trained to predict the cognitive function assessment score corresponding to the features of the voice data; The first deep neural network model is The model is trained using a supervised learning method with training data that uses features of speech data as input and labels indicating the presence or absence of Alzheimer's disease. The second deep neural network model is This model is trained using a supervised learning method with training data that uses features from speech data as input and cognitive function assessment scores as labels. The method for diagnosing Alzheimer's disease according to claim 5.

7. The generating step includes: and simultaneously generating a classification result of Alzheimer's disease and a predicted result of a cognitive function assessment score in response to features of the speech data that are less than a predetermined significance level. The method for diagnosing Alzheimer's disease according to claim 6.

8. The generating step includes: generating a first Alzheimer's disease classification result and a first cognitive function assessment score prediction result based on a feature selection result of a speech section of first speech data spoken by a speaker in response to a first speech task that requires the speaker to respond to one or more predetermined questions; a step of outputting a reading result for a predetermined story, and generating a second Alzheimer's disease classification result and a second cognitive function assessment score prediction result based on a feature selection result for a speech section of second audio data spoken by a speaker in response to a second speech task that requests the speaker to speak in accordance with the reading result; generating a third Alzheimer's disease classification result and a third cognitive function assessment score prediction result based on a feature selection result for a speech section of third speech data spoken by a speaker in response to a third speech task requiring recollection of a story; The method for diagnosing Alzheimer's disease according to claim 1.

9. The generating step includes: generating a classification result of Alzheimer's disease based on the feature selection results of the speech segments of the first speech data and the second speech data; determining whether to generate a prediction result of a cognitive function assessment score based on the Alzheimer's disease classification result; generating a predicted result of the cognitive function assessment score based on a feature selection result for the speech section of the third speech data excluding the features of the first speech data and the features of the second speech data, when it is determined to execute generation of a predicted result of the cognitive function assessment score; The method for diagnosing Alzheimer's disease according to claim 2.

10. The determining step comprises: determining to generate a predicted result of the cognitive function assessment score when a classification result of Alzheimer's disease of one or more of the features of the first speech data and the second speech data is generated as a first value; The method for diagnosing Alzheimer's disease according to claim 8.

11. A computer-readable recording medium storing a computer program for causing a computer to execute the method according to any one of claims 1 to 10.

12. As an Alzheimer's disease diagnostic device, a processor; a memory operatively connected to the processor and storing at least one code for execution by the processor; The memory, when executed by the processor, causes the processor to collect speech data of a speaker; Extract features from the voice data, storing code for generating one or more of a classification result of Alzheimer's disease and a prediction result of a cognitive function assessment score based on the features of the speech data; Alzheimer's disease diagnostic device.

Citation Information

Patent Citations

  • Portable water purifier bottle with filtration and sterilization function

    KR1020220111588A

  • Film heater

    KR1020230111720A