Neurocognitive Impairment Detection via Speech Pause Probability Histograms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting neurocognitive impairment, such as mild cognitive impairment (MCI), are time-consuming, labor-intensive, and require trained professionals, with existing automated solutions failing to accurately account for hesitation pauses and providing high error rates, especially in early stages of Alzheimer's disease.

Innovation Solution

An automated speech analysis method that generates decision information based on probability values for silent and filled pauses, using a histogram representation of speech samples, which is processed by a machine learning algorithm to accurately detect neurocognitive impairment without the need for trained personnel.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated speech analysis is used to detect neurocognitive impairment, then productivity and accessibility are improved, but measurement precision deteriorates due to high error rates in detecting hesitation pauses

Engineering Contradiction:
Improvescreening capacityVSAvoiddetection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts the probability threshold for classifying filled pauses based on the specific speech sample and context, rather than using a fixed threshold. This allows the measurement criteria to adapt to varying speech patterns and conditions, improving detection accuracy while maintaining automated processing

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The invention changes the parameter used for pause classification from simple acoustic features to probability values derived from multiple speech recognition engines. By transforming the input parameters and using histogram analysis of probability distributions, the system achieves more precise differentiation between silent pauses and filled pauses, resolving the measurement precision issue

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If simple acoustic parameter analysis is used, then device complexity is reduced, but measurement precision deteriorates due to inability to accurately distinguish pause types

Engineering Contradiction:
Improvesystem simplicityVSAvoidpause classification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The invention introduces probability values as an intermediary between the raw speech signal and the final pause classification. Instead of directly analyzing acoustic parameters, the system uses probability distributions from speech recognition engines as a mediating layer to indirectly and more accurately determine pause types, achieving high precision without complex hardware

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual evaluation by trained professionals is used, then measurement precision is improved, but productivity deteriorates due to time-consuming processes

Engineering Contradiction:
Improvediagnostic accuracyVSAvoidscreening speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-evaluation by automatically analyzing speech samples without requiring trained professionals. It uses multiple speech recognition engines to generate probability values and automatically classifies pauses based on histogram analysis of these probabilities, enabling both high precision and rapid processing of large numbers of patients

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention replaces the mechanical system of manual professional evaluation with an automated computational system. Instead of human listeners analyzing speech samples, the system uses speech recognition engines and probability-based algorithms to automatically detect and classify pauses, achieving both speed and accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Ease of operation

If probability-based classification with fixed thresholds is used, then ease of operation is improved, but measurement precision deteriorates due to high false positive rates

Engineering Contradiction:
Improveautomation levelVSAvoidfalse positive rate
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system dynamically determines classification thresholds based on the specific speech sample being analyzed, rather than using fixed predetermined thresholds. This dynamic adaptation reduces false positives by adjusting the criteria to match the actual speech patterns and context of each patient

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The invention adds a new dimension to the classification process by using histogram analysis of probability distributions across multiple speech recognition engines. Instead of relying on a single threshold value, the system analyzes the distribution of probability values across different engines and time points, providing a more robust and accurate classification that reduces false positives

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12161481B2Automatic detection of neurocognitive impairment based on a speech sample
Publication Date: 2024.12.10 SZTE TTC ZRT
  • US12161481B2 patent drawing
  • US12161481B2 patent drawing
  • US12161481B2 patent drawing

AI summary

The invention is a method for automatic detection of neurocognitive impairment, comprising,generating, in a segmentation and labelling step (11), a labelled segment series (26) from a speech sample (22) using a speech recognition unit (24); andgenerating from the labelled segment series (26), in an acoustic parameter calculation step (12), acoustic parameters (30) characterizing the speech sample (22).The method is characterised bydetermining, in a probability analysis step (14), in a particular temporal division of the speech sample (22), respective probability values (38) corresponding to silent pauses, filled pauses and any types of pauses for respective temporal intervals thereof;calculating, in an additional parameter calculating step (15), a histogram by generating an additional histogram data set (42) from the determined probability values (38) by dividing a probability domain into subdomains and aggregating durations of the temporal intervals corresponding to the probability values falling into the respective subdomains; andgenerating, in an evaluation step (13), decision information (34) by feeding the acoustic parameters (30) and the additional histogram data set (42) into an evaluation unit (32), the evaluation unit (32) using a machine learning algorithm.The invention is furthermore data processing system, a computer program product and a computer-readable storage medium for carrying out the method.