Voice-based method for diagnosing parkinson's disease

The method employs neural networks to analyze voice recordings, addressing limitations in existing voice analysis for Parkinson's disease diagnosis by increasing accuracy and simplifying the diagnostic process, enabling remote and accessible diagnosis.

WO2025105987A1PCT designated stage expired Publication Date: 2025-05-22OBSHCHESTVO S OGRANICHENNOI OTVETSTVENNOSTIU TSIFROVYE RESHENIIA DLIA DIAGNOSTIKI ZABOLEVANII MOZGA BREINFON
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/RU2024/050277
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-10-31
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing methods for diagnosing Parkinson's disease and other Parkinsonism syndromes using voice analysis are limited by their reliance on resource-intensive machine learning algorithms and the need for manual extraction of acoustic characteristics, which can lead to inaccurate results due to unstructured data and small sample sizes.

Method used

A method utilizing artificial intelligence and neural network architectures to analyze voice recordings, where audio recordings are processed by a neural network that extracts implicit features from mel-spectrograms, allowing for accurate diagnosis and monitoring of Parkinson's disease and other Parkinsonism syndromes without the need for manual feature extraction.

Benefits of technology

This approach increases the accuracy of diagnostics and monitoring while simplifying the procedure, achieving more accurate diagnostic results compared to solutions limited by pre-known features, and enabling remote and accessible diagnosis using a smartphone and internet connection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000011_0000
    Figure 00000011_0000
  • Figure 00000011_0001
    Figure 00000011_0001
  • Figure 00000012_0000
    Figure 00000012_0000
Patent Text Reader

Abstract

The invention relates to the field of information and communications technology (ICT) specially designed for medical diagnostic and monitoring purposes. A voice-based method for diagnosing and monitoring Parkinson's disease and other parkinsonian disorders using artificial intelligence envisages the steps of: recording a subject's voice or uploading a pre-recorded sound clip of the subject's voice via an interface of a client electronic device; transmitting the recording to a server via a web service; processing the recording using a neural network which subsequently provides a response containing the probability of the presence of signs of Parkinson's disease and / or parkinsonism in the subject; receiving, on the client device, the neural network response in the form of a probability of the presence of parkinsonism and a classification of the recording according to two or more classes, including but not limited to the labels "healthy person" or "signs of parkinsonism". The technical result of the invention is that of enhancing diagnostic accuracy and monitoring, as well as simplifying the procedure itself without detriment to quality.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD OF DIAGNOSIS OF PARKINSON'S DISEASE BY VOICE

[0002] The invention relates to the field of information and communication technologies (ICT), specifically designed for medical diagnostics and monitoring.

[0003] State of the art.

[0004] The prior art discloses a speech recognition system for people with Parkinson's disease based on integrated dimensionality reduction, a patent for an invention (CN111210846B, published on July 5, 2022). The invention is a speech recognition system for Parkinson's disease based on integrated dimensionality reduction and comprises: 1) a data collection unit: training, verification and test data; 2) a classifier module that classifies and identifies test data; 3) an output module that is used to output the final recognition result. This invention is a method for extracting speech characteristics using the LPP (Local Preserved Projection) algorithm, which largely preserves the essence of the data while reducing its dimensionality.After data extraction, one of the two proposed machine learning algorithms is applied: SVM (Support Vector Machine), RF (Random Forest); or a single-layer neural network ELM (Extreme Leaning Machine).

[0005] Differences from the declared technical solution: emphasis on preparation and extraction of characteristics from data, the classification model plays a lesser role, therefore, resource-intensive machine learning algorithms (they do not extract characteristics from data independently) and a single-layer neural network were chosen. The quality of the approach proposed by the authors depends heavily on the quality of work with data and to a lesser extent on the selected classification algorithms.

[0006] The state of the art also includes systems and methods for screening neurological and other diseases using the speech behavior of a subject, a patent for an invention (US20120265024A1, published on July 22, 2014). In this invention, the authors of the presented patent extract characteristics from speech recordings and / or sound recordings (phonemes) using an algorithm, by which the patient's data is compared with reference data. As such an algorithm, the authors propose: a statistical approach, pattern recognition and / or machine learning algorithms (Hidden Markov models, SVM, neural networks). It is worth noting the year of the proposed solution (not patenting), to which no modern neural network architectures have yet been presented that could qualitatively extract the necessary features from patient records.

[0007] The biggest technological difference is that the claimed technical solution has neural network architectures for solving the problem of diagnosing patients, the fine-tuning of which does not require manually extracting acoustic characteristics from audio recordings. It should be noted that the claimed technical solution is used for screening diagnostics using audio recordings in patients not only with Parkinson's disease, but also in patients with parkinsonism-plus syndromes, as well as secondary parkinsonism.

[0008] It should be noted that many solutions use models trained on a very small number of audio recordings of subjects' voices (less than 50 subjects), which is not enough to train more complex algorithms, and is also not a representative sample of patients. The data is often unstructured, patients can perform different tasks during audio recordings, and such data do not always allow for an accurate determination of Parkinson's disease, unlike the case of all patients performing the same type of tasks (for example, pronouncing one phoneme for some time).

[0009] The above technical solutions, as well as the claimed invention, allow one way or another to analyze audio recordings, however, the principle and procedure of the analysis differs significantly from that implemented in the present invention.

[0010] Disclosure of invention

[0011] For a better understanding of the present invention, the main terms used in the present description of the invention are provided and explained below. Unless otherwise defined, technical and scientific terms in this application have the standard meanings generally accepted in the scientific and technical literature.

[0012] In order to ensure sufficient disclosure of the invention in relation to the claimed technical solution, a list of terms used in the description of the claimed invention is provided below.

[0013] API (API) is an application programming interface (a set of classes, procedures, functions, structures, or constants) by which one computer system can interact with another system, as well as a way of using these elements using some protocol. Dysarthria is a speech disorder caused by discoordination of the respiratory muscles, vocal cords, larynx, palate, tongue, lips in case of disruption of the innervation of the speech apparatus at any level: from the cerebral cortex to the peripheral nerves, as well as at the level of the cerebellum or subcortical nuclei. In Parkinson's disease, specific hypokinetic dysarthria develops, which is characterized by decreased voice volume, monotony, reduced fundamental frequency range, inaccuracy of articulation of consonants and vowels, shallow breathing, short bursts of speech, and irregular pauses.

[0014] Semantic analysis is an important subtask of Natural Language Processing (NLP), a stage in the sequence of actions of the algorithm for automatic understanding of texts, which consists of identifying semantic relationships and forming a semantic representation of texts

[0015] Phonation - pronunciation, the process of articulation of human speech

[0016] Articulation is the joint work of the speech organs, necessary for the pronunciation of speech sounds.

[0017] Prosody - The system of pronunciation of stressed and unstressed, long and short syllables in speech

[0018] A biomarker is any parameter that can be reliably measured and that can be used to learn something about a person's health or mortality status: for example, the presence of a disease, a physiological change, a response to treatment, or a psychological disorder.

[0019] Parkinsonism-plus syndromes - a syndrome characterized by the presence of parkinsonism in the patient and other signs, such as early imbalance / falls, poor response to levodopa drugs, early development of cognitive impairment and autonomic failure, etc.

[0020] Machine learning is the science of developing algorithms and statistical models that computer systems use to perform tasks without explicit instructions, relying instead on patterns and inference.

[0021] A neural network / neural network model is a simplified model of the nervous system of living organisms, the basic units of which are called neurons and are usually grouped into layers. Neurodegenerative diseases are a heterogeneous group of disorders of the nervous system that arise due to the progressive degeneration and death of certain groups of neurons, which leads to disruption of the synapses, glial cells, and the networks that they together form.

[0022] A mel spectrogram is a regular spectrogram where the frequency is expressed not in Hz, but in mels, the transition to which is achieved by applying mel filters (triangular functions uniformly distributed on the mel scale) to the original spectrogram.

[0023] A client device is an electronic device (smartphone, laptop, computer or other device) that provides voice audio recording, communication with the server and output of results. The device must provide recording of audio data that meets the following minimum specifications:

[0024] - PCM or WAV or OGG format;

[0025] - Sampling frequency 48 kHz;

[0026] - Number of channels 1 (mono);

[0027] - Quantization depth 16 bits (2 bytes) per sample;

[0028] - Byte order (reverse);

[0029] - Signed Integer numbers;

[0030] - Bitrate Constant (constant), 768 kb / s.

[0031] The audio sampling frequency after conversion can be reduced to 16 kHz, which is sufficient for solving many problems related to speech analysis.

[0032] The task that the claimed technical solution is aimed at solving is the rapid, remote and accessible diagnosis and monitoring of Parkinson's disease and other diseases with Parkinsonism syndrome based on the analysis of voice audio recordings.

[0033] The technical result of the invention is to increase the accuracy of diagnostics and monitoring, as well as to simplify the procedure itself without loss of quality.

[0034] This is achieved by the fact that the claimed method for diagnosing and monitoring Parkinson's disease and other diseases with parkinsonism syndrome by a person's voice based on artificial intelligence provides for the following stages: recording a voice or downloading a pre-recorded audio recording of a subject through the interface of a client electronic device, transmitting the recording to the server through a web service and processing it by a neural network, which then issues a response with the probability of the presence of signs of Parkinson's disease and / or parkinsonism in the subject, receiving a response from the neural network on the client device in the form of the probability of the presence of parkinsonism and classifying the recording into two or more classes, including, but not limited to, the labels "healthy person" or "signs of parkinsonism".

[0035] The invention is illustrated by drawings:

[0036] Fig. 1 - An example of a mel-spectrogram of a phoneme (a) in a healthy person (a) and a patient with parkinsonism (Parkinson's disease) (b);

[0037] Fig. 2 - Block diagram of an example of the method operation.

[0038] Implementation of the invention

[0039] The invention is a method, including a web service that implements internal processing of requests and access to a neural network model, with which you can configure integration via REST API, or use a ready-made web interface that can be interacted with in any modern browser on a desktop or smartphone. Thus, the invention can be presented as a mobile application, a widget on a website, integrated with a call center, or connected to any other software, including that installed in medical centers. A server that receives requests from a client device and processes the audio recording and sends the results to the client device.Server requirements: OS must support Docker and docker-compose (Ubuntu, Debian, CentOS, RHEL), minimum hardware requirements: 1 vCPU (Intel or AMD processor with x86_64 architecture), 3 GB RAM, 2 GB vRAM (NVIDIA graphics card - any model from the Ampere, Ada Lovelace, Turing, Volta series).

[0040] - Using the web interface, you can make an audio recording of your voice or upload previously recorded audio files of any format, and get instant diagnostic results. Audio recordings of patients' voices, obtained in various ways when performing tasks on arbitrary speech or pulling out phonemes, developed with the participation of Parkinson doctors, are fed to the input in requests to the web service, automatically converted into audio samples with the required sampling frequency and bit depth. Then the raw signal is decomposed into a spectrum using the Fourier transform, and the resulting audio spectrogram, reduced to a small scale that focuses attention on the part of the spectrum with the voice, containing maximum information on prosodies, articulation and other voice features, is fed to the input of the neural network model for classification.The neural network extracts features from the audio spectrogram in an implicit form, which allows, with a sufficiently large training sample (from 1000 subjects and more), to achieve more accurate diagnostic results compared to solutions limited to the use of only pre-known and algorithmically extracted explicit features.

[0041] - Also, to eliminate anomalies in predictions obtained from audio recordings that contain fragments with no speech (silence), the web service includes a Voice Activity Detector (VAD) model for detecting voice activity in audio, which marks and filters fragments that do not contain the patient’s voice.

[0042] - The formation of the dataset, determination of the minimum length of audio recordings, as well as the composition of tasks that subjects must complete were developed with the participation of neurologists and pairs of kinsonologists.

[0043] - The server, which includes the web service, returns a response with the forecast of the neural network model of the classifier and the diagnosis formed on its basis. In addition to determining the class label to which the subject was assigned, the web service returns the degree of confidence of the model in its forecast, obtained as a real numerical value from 0 to 1 after applying the Softmax activation function.

[0044] In the proposed method, classifiers can divide patients into at least the following classes - "healthy person" or "signs of Parkinsonism" with possible detailing of the type of Parkinsonism, staging, etc. In the method, conclusions about the patient's illness can be presented in the form of one or more conclusions indicating the probability of their reliability.

[0045] Features and advantages of the proposed solution:

[0046] 1. accessibility (can be integrated into an accessible shell - chat bot, website, mobile application, call center protocols);

[0047] 2. ubiquitous use (can be used anywhere, if you have a smartphone and the Internet);

[0048] 3. mass character (can be used as a screening tool);

[0049] 4. training the neural network on audio recordings of native Russian speakers; training on a large sample of patients (more than 1000 subjects);

[0050] 5. extraction of implicit features from mel-spectrograms of audio fragments with voice by a neural network, which allows achieving higher quality indicators compared to solutions using compressed ready-made features calculated from audio recordings, in which some information may be lost;

[0051] 6. cross-platform and universal service: the ability to install on servers and client devices with any systems that support work with Docker; the ability to process audio files of any format containing the patient’s voice;

[0052] 7. focus on diagnosing Parkinson's disease, which does not exclude the expansion of the invention's diagnostic capabilities to other types of neurodegenerative diseases using audio with additional collection of relevant data and training of additional models.

[0053] Example 1.

[0054] The presented solution can be used at least in medical institutions (general and special purpose), research scientific centers, in telemedicine, as well as directly by patients with neurological diseases and people wishing to assess their health.

[0055] The method claimed for registration as an invention can be implemented in the form of integration into a call center.

[0056] In this case, the system will work as follows.

[0057] When calling the call center, the patient will be asked if he wants his voice to be assessed for Parkinsonism (Parkinson's disease or other diseases associated with Parkinsonism syndrome). After receiving consent, instructions will be given on the necessary recording, which will be transferred to the neural network circuit upon completion.

[0058] Below is an example of converting an audiogram obtained from a patient into a mel-spectrogram and a visual comparison with the mel-spectrogram of a healthy person.

[0059] After analysis, the neural network returns a response with a certain probability of the presence of Parkinsonism in a particular subject, which can be sent in various ways: SMS, email or any other available way.

[0060] The server returns a response with the forecast of the neural network classifier model and the conclusion formed on its basis. In addition to determining the class label to which the subject was assigned, the server returns the degree of confidence of the model in its forecast, obtained as a real numerical value from 0 to 1 after applying the Softmax activation function. The threshold value at which the conclusion “healthy person” is formed is less than 0.5; at a probability of 0.5 and higher, the conclusion “signs of Parkinsonism” is formed.

Claims

CLAUSE OF THE INVENTION 1. A method for diagnosing and monitoring Parkinson's disease and other diseases with parkinsonism syndrome by voice based on artificial intelligence includes the following steps: recording a voice or uploading a pre-recorded audio recording of a subject through the interface of a client electronic device, transmitting the recording to a server through a web service and processing it by a neural network, which then issues a response with the probability of the presence of signs of Parkinson's disease and / or parkinsonism in the subject, receiving a response from the neural network on the client device in the form of the probability of the presence of parkinsonism and classifying the recording into two or more classes, including, but not limited to, the labels "healthy person" or "signs of parkinsonism".

Citation Information

Patent Citations

  • ASSESSMENT OF PARKINSON'S DISEASE CONDITION(S) BASED ON VOICE RECOGNITION

    RU2022132120A

  • Computer-aided diagnostic system for early diagnosis of prostate cancer

    US11495327B2

  • Method and apparatus for providing a predictive healthcare service

    US20140343955A1

  • Wearable personal digital device for facilitating mobile device payments and personal use

    US20170293740A1

  • Clinical Pathway Integration and Clinical Decision Support

    US20220359091A1