Method and apparatus for evaluating neurolinguistic disorder using explainable artificial intelligence

Explainable AI is used to automatically diagnose dysarthria severity by extracting specific features from voice data, providing visual explanations, and classifying severity, addressing the inefficiencies of current subjective assessments, thereby improving accuracy and patient motivation.

WO2026029275A1PCT designated stage Publication Date: 2026-02-05HAII CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017381
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-30
Filing Date
2024-11-06
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current face-to-face assessments of post-stroke dysarthria are time-consuming, resource-intensive, and reliant on subjective evaluations, making it difficult to discern subtle changes and varying symptoms, and the reliability of results depends on the evaluator's experience.

Method used

Utilizing explainable artificial intelligence (XAI) to automatically diagnose dysarthria severity by extracting dysarthria-specific features from voice data, classifying severity using a tree-based machine learning algorithm, and providing visualized explanations through techniques like SHAP and LLM-generated sentences.

Benefits of technology

Enables rapid, accurate, and personalized assessments, reducing time and resource requirements while enhancing the reliability of diagnostic results and motivating patients towards rehabilitation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017381_05022026_PF_FP_ABST
    Figure KR2024017381_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for evaluating the severity of a neurolinguistic disorder using explainable artificial intelligence, implemented by a server. The present method may comprise the steps of: receiving, by a server, a patient's voice data collected through a user terminal; performing audio preprocessing and voice segmentation on the received voice data; extracting dysarthria-specific features from the processed voice data; classifying the severity of dysarthria on the basis of the extracted features using a tree-based machine learning algorithm; generating visualization information by analyzing and visualizing the contribution of the features that influenced the prediction using explainable artificial intelligence techniques; and generating natural language-type large language model (LLM)-generated statements explaining the rationale for a prediction result using an LLM.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for evaluating neurolinguistic disorders using explainable artificial intelligence

[0001] The present disclosure relates to a method and device for determining the severity and condition of a patient with dysarthria after stroke by utilizing eXplainable AI technology.

[0002] Accurate speech assessments of post-stroke dysarthria can help determine the severity of dysarthria and identify the underlying neurological problems based on the patient's speech characteristics. This can aid in identifying appropriate interventions and monitoring prognosis. Because post-stroke dysarthria presents with varying symptoms depending on the location and size of the lesion, accurately assessing each patient's condition and providing tailored speech therapy are crucial. However, current face-to-face assessments of dysarthria rely on auditory-perceptual assessments by speech-language pathologists. This process is time-consuming and resource-intensive, making it difficult to discern subtle changes in patients, and the reliability of the results can vary depending on the evaluator's experience level.

[0003] AI technology can complement the subjective limitations of existing diagnoses and capture more subtle changes, enabling more precise and accurate assessments and personalized treatments. XAI aims to help humans understand the decision-making process of AI systems. It can develop a diagnostic system that can explain the severity and error patterns of individual patients' speech production subsystems to them. Using AI technology, the system can automatically diagnose items assessed by speech therapists, and explainable AI and visualization techniques can identify the patient's specific problems and provide diagnostic evidence and explanations. This can present personalized assessment results to each patient, enhancing their understanding and focus, motivating them to pursue their own rehabilitation, and ultimately leading to effective behavioral change. Furthermore, it is expected to assist medical professionals in making decisions related to diagnosis and treatment based on objective results.

[0004] [Prior Art Literature]

[0005] [Patent Document]

[0006] (Patent Document 1) Registered Patent No. 10-2539049: Method and Device for Evaluating Neurospeech Disorders

[0007] The objective of this disclosure is to utilize explainable artificial intelligence to accurately assess the severity and condition of patients with post-stroke dysarthria and provide customized speech therapy based on this assessment. Furthermore, this disclosure utilizes AI technology to automatically diagnose assessment items required by speech therapists. Furthermore, explainable AI and visualization techniques are used to assess patients' problems and provide diagnostic evidence and explanations. This approach aims to reduce the time and resources required for the assessment process and enhance the reliability of the assessment. This approach can motivate patients to pursue rehabilitation and provide objective and reliable assessment results to medical professionals, supporting their diagnostic and treatment decision-making.

[0008] In one embodiment of the present disclosure, a method for evaluating the severity of a neurolanguage disorder using explainable artificial intelligence by a server is disclosed, the method comprising: a step of receiving, by the server, a patient's voice data collected through a user terminal; a step of performing audio preprocessing and voice segmentation processing on the voice data by the server; a step of extracting, by the server, a dysarthria-specific feature from the processed voice data; a step of classifying, by the server, the severity of the dysarthria based on the extracted specific feature using a tree-based machine learning algorithm; a step of generating, by the server, visualized information that analyzes and visualizes the contribution of the specific feature that has influenced the prediction using an explainable artificial intelligence technique; and a step of generating, by the server, an LLM (Large Language Model)-generated sentence in a natural language format as a basis for the prediction result.

[0009] In one embodiment of the present disclosure, the specific feature may include at least one of pitch, intensity, jitter (%), shimmer (%), harmonic-to-noise ratio, speaking rate, articulation rate, pause frequency ratio, pause length, and character accuracy.

[0010] In one embodiment of the present disclosure, the server may further include a step of providing the visualization information and the LLM generation statement to the user terminal.

[0011] In one embodiment of the present disclosure, the tree-based machine learning algorithm may be LightGBM (Light Gradient Boosting Machine).

[0012] In one embodiment of the present disclosure, the explainable artificial intelligence technique may be the SHapley Additive exPlanations (SHAP) technique.

[0013] In one embodiment of the present disclosure, the visualization information may include at least one of a graph, a chart, and text.

[0014] In one embodiment of the present disclosure, the step of performing the audio preprocessing and voice segmentation processing may include the step of segmenting the speech segment through noise removal and voice activity detection (VAD).

[0015] In another aspect of the present disclosure, a neurolanguage disorder evaluation device for performing the above neurolanguage disorder evaluation method is disclosed, the neurolanguage disorder evaluation device including a processor and a memory.

[0016] A method for assessing neurolinguistic disorders according to one embodiment of the present disclosure includes a step of automatically assessing the degree of a user's neurolinguistic disorder based on voice data received from a user terminal using explainable artificial intelligence. This enables rapid, convenient, and highly accurate assessment. This enables the development of customized treatment plans, reduction of time and resources in the assessment process, enhanced understanding and reliability of diagnostic results, increased motivation for rehabilitation, and flexible application in various medical settings.

[0017] FIG. 1 is a schematic diagram showing a server and a terminal using a neurolinguistic disorder evaluation method according to one embodiment of the present disclosure.

[0018] FIG. 2 is a flowchart illustrating steps of a neurolinguistic disorder assessment method according to one embodiment of the present disclosure.

[0019] FIG. 3 is a detailed diagram of a data processing pipeline of a neurolanguage disorder assessment method according to one embodiment of the present disclosure.

[0020] FIG. 4 is a detailed diagram of a data processing pipeline for paragraph reading in a neurolanguage disorder assessment method according to one embodiment of the present disclosure.

[0021] FIG. 5A is a speech assessment result user terminal display screen providing the overall results classifying the severity of speech impairment after stroke according to one embodiment of the present disclosure.

[0022] FIG. 5b is a user terminal display screen that evaluates the severity of speech impairment after stroke and shows detailed results for each area according to one embodiment of the present disclosure.

[0023] FIG. 5c is a user terminal display screen that visually provides results by region in a neurolanguage disorder evaluation device according to one embodiment of the present disclosure.

[0024] FIG. 6 is a user terminal display screen visually providing detailed results of excellent areas and improvement areas in a neurolanguage disorder evaluation device according to one embodiment of the present disclosure.

[0025] FIG. 7 is a user terminal display screen detailing one area requiring improvement in a neurospeech disorder assessment device according to one embodiment of the present disclosure.

[0026] FIG. 8 is a user terminal display screen of a comparison result provided by a neurolanguage disorder evaluation device according to one embodiment of the present disclosure.

[0027] FIG. 9 is a user terminal display screen visually representing a patient's speech impairment score prediction result according to one embodiment of the present disclosure.

[0028] FIG. 10 is a user terminal display screen visually showing the cumulative results of a patient's speech impairment evaluation score according to one embodiment of the present disclosure.

[0029] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail together with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below, but may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure is complete and to fully inform those skilled in the art of the scope of the invention, and the present disclosure is defined solely by the scope of the claims.

[0030] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present disclosure. For example, a component expressed in the singular should be understood to include plural components unless the context clearly indicates only the singular. Furthermore, in the specification of the present disclosure, terms such as "comprise" or "have" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and the use of such terms does not exclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0031] Additionally, unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0032] Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless expressly defined in the specification of the present disclosure.

[0033] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the attached drawings. However, in the following description, specific descriptions of widely known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.

[0034] The prior art documents described in this disclosure are incorporated herein by reference in their entirety, and it will be understood that the contents of the prior art documents can be applied by a person having ordinary skill in the art to the parts briefly described in this disclosure.

[0035] Hereinafter, a neuro-language disorder evaluation device (10) according to an embodiment of the present disclosure will be described with reference to the drawings.

[0036] FIG. 1 is a schematic diagram showing a server (10) and a user terminal (20) for evaluating the severity of a neurolanguage disorder according to one embodiment of the present disclosure, and FIG. 2 is a flowchart showing steps (S210 to S260) of a neurolanguage disorder evaluation method (200) according to one embodiment of the present disclosure.

[0037] The server (10) may include a processor (not shown), memory (not shown), a communication module (not shown), etc., and may be configured to communicate with a user terminal (20).

[0038] A user terminal (20) refers to at least one of a patient terminal, a medical staff terminal, and a caregiver terminal. Examples of user terminals (20) include mobile phones, smartphones, tablet PCs, personal digital assistants (PDAs), laptop computers, smartwatches, and other electronic devices capable of connecting to the Internet. These user terminals (20) can be utilized by various users, including patients, medical staff, and caregivers, and can perform functions such as collecting and transmitting voice data.

[0039] The user terminal (20) may be configured to communicate with the server (10). The terminal (20) includes a processor (not shown), memory (not shown), a display (21), a speaker (not shown), and a microphone (not shown).

[0040] The above processor may be configured to perform a task according to instructions stored in the memory.

[0041] The above microphone may be configured to detect the user's speech so as to record the user's voice.

[0042] The server (10) is a device capable of performing tasks such as SHAP (SHapley Additive exPlanations), LightGBM (Light Gradient Boosting Machine), and Global / Local interpretability, and may be configured as at least one of, for example, an on-premise server installed within a company, a cloud server accessible via the Internet such as AWS (Amazon Web Services), Google Cloud Platform, and Microsoft Azure, a hybrid server that is a combination of an on-premise server and a cloud server, a GPU server equipped with a GPU (Graphics Processing Unit), an HPC cluster (High Performance Computing Cluster), which is a cluster that connects multiple high-performance computers and uses them as a single system, and an edge server that processes data near a data generation location.

[0043] SHAP is a method for explaining the predictions of a machine learning model. It can quantitatively evaluate the degree to which each feature contributes to the predicted result, and based on the Shapley value of game theory, it can fairly calculate the contribution of each feature.

[0044] LightGBM is a gradient boosting framework developed by Microsoft. It boasts high efficiency and fast learning speed, is suitable for processing large data sets, and can provide various hyperparameter tuning options.

[0045] Global interpretability refers to understanding the model's predictive behavior across the entire dataset. It provides insight into how the model performs overall and is primarily used to assess the model's complexity and predictive consistency.

[0046] Local interpretability refers to understanding a model's predictions for specific data points. It explains how individual predictions were made and is primarily used to assess the reliability of each prediction and the detailed behavior of the model.

[0047] The communication module of the server (10) can receive patient voice data collected through the user terminal (20). In addition, the communication module can be configured to provide visualization information and LLM generation statements to the user terminal (20).

[0048] For example, the communication method of the communication module may utilize a network built according to GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), etc.), WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), 5G, etc., but is not limited thereto and may include all transmission method standards to be developed in the future.

[0049] The memory of the server (10) is configured to store instructions executed by the processor. The memory may be configured to store voice data and evaluation results.

[0050] In one embodiment, memory may include computer-readable storage media, such as data storage devices, that are accessible by the computing device and provide persistent storage of data and executable instructions (e.g., software applications, programs, functions, etc.). Examples of memory include volatile memory and non-volatile memory, fixed and removable media devices, and any suitable memory device or electronic data storage that maintains data for access by the computing device. Memory may include various embodiments of random access memory (RAM), read-only memory (ROM), flash memory, and other types of storage media in various memory device configurations. Memory may be configured to store executable software instructions (e.g., computer-executable instructions) that are executable with the processor, or the same software application that may be implemented as a module.

[0051]

[0052] The neurolanguage disorder evaluation method (200) utilizing explainable artificial intelligence of the present disclosure may include a step (S210) of receiving, by a server (10), a patient's voice data collected through a user terminal, a step (S220) of performing audio preprocessing and voice segmentation processing on the voice data, a step (S230) of extracting dysarthria-specific features from the processed voice data, a step (S240) of classifying the severity of dysarthria based on the extracted specific features using a tree-based machine learning algorithm, a step (S250) of generating visualization information that analyzes and visualizes the contribution of the specific features that have influenced the prediction using an explainable artificial intelligence technique, and a step (S260) of generating an LLM-generated sentence in a natural language format as a basis for the prediction result using an LLM (Large Language Model).

[0053] Distinctive features are characteristics or attributes of data that provide important or meaningful information about a specific situation, condition, or task. Distinctive features may include at least one of speech patterns, pronunciation accuracy, speech rate, loudness, formation band, jitter, shimmer, pitch, quality, consistency, articulatory clarity, rhythm, and fluency, and may refer to specific elements of speech data directly related to a patient's speech status. These features play a key role in classifying and assessing the severity of speech disorders and can be used to develop effective diagnosis and treatment plans based on this data.

[0054]

[0055] Audio preprocessing and speech segmentation are a series of processing steps performed to transform collected speech data into a form suitable for analysis. These steps may include the following:

[0056] - Noise Reduction: Voice data may contain background noise or static. Noise reduction is the process of removing this unwanted noise to improve the quality of the voice signal. Methods such as noise filtering, wavelet transform, and frequency-domain noise removal can be used.

[0057] - Volume Normalization: This process minimizes the impact of volume differences during analysis by adjusting the volume of voice data to a consistent level. Methods such as root mean square (RMS)-based normalization and peak normalization can be used.

[0058] Voice Activity Detection (VAD): This is the process of distinguishing between actual speech and silent sections in voice data. Energy-based, frequency-based, and machine learning-based VAD techniques can be used.

[0059] - Resampling: This is the process of converting the sampling rate of voice data to a standard rate suitable for analysis. Downsampling and upsampling can be used.

[0060] Framing: This process divides audio data into time intervals to create frames. This allows each frame to be analyzed individually. Typically, frames are divided into 20 to 40 millisecond intervals.

[0061] Feature Extraction: This is the process of extracting useful features for analysis, such as frequency, energy, and voice characteristics, from voice data. Mel-Frequency Cepstral Coefficients (MFCC), spectrograms, and formant frequencies can be used.

[0062] Segmentation: This is the process of dividing speech data into individual utterance units by identifying utterance segments. This allows for independent analysis of each utterance unit. Segmentation can be performed using time-based, pronunciation-based, or phoneme-based methods.

[0063] Audio preprocessing is an essential step in preparing speech data for analysis. This preprocessing process improves the quality of speech data and enables more accurate and reliable results during subsequent analysis. In medical fields such as neurospeech disorder assessment, preprocessing is particularly crucial for improving data accuracy and consistency. The speech data collected through this audio preprocessing process is used to extract specific features of dysarthria and then converted into a format suitable for severity assessment using machine learning algorithms.

[0064]

[0065] In one embodiment, the neurolanguage disorder evaluation method (200) may further include a step of providing the visualization information and the LLM generation sentence to a user terminal (20) by the server (10).

[0066] LLMs are artificial intelligence models capable of performing natural language processing tasks by learning from large-scale text data. LLMs support a variety of language understanding and generation tasks, offering capabilities such as sentence generation, question answering, translation, and summarization. These models utilize deep learning technology to learn complex patterns in text data and possess human-level language understanding and generation capabilities. For example, models like GPT-3 utilize billions of parameters to perform a wide range of language tasks.

[0067] The above tree-based machine learning algorithm may be LightGBM (Light Gradient Boosting Machine).

[0068] LightGBM is a gradient boosting framework developed by Microsoft. It is specifically optimized for efficiently processing large data sets and learning quickly. This algorithm trains and combines multiple decision trees in stages to create a final predictive model. Key features of LightGBM include:

[0069] 1. Tree-based learning: LightGBM uses decision trees to learn data. This is used to build a predictive model, and multiple trees are combined to create the final model.

[0070] 2. Efficiency and speed: LightGBM is designed to perform data segmentation and training very quickly, and performs particularly well on large data sets.

[0071] 3. Accuracy: LightGBM provides high accuracy and often performs better in prediction than other tree-based algorithms.

[0072] 4. Hyperparameter Tuning: LightGBM provides various hyperparameters to optimize model performance.

[0073] 5. Memory efficiency: LightGBM minimizes memory usage and works efficiently even on large data sets.

[0074] The server (10) of the present disclosure utilizes LightGBM to extract features specific to dysarthria from speech data and classify the severity of dysarthria using a tree-based machine learning algorithm. LightGBM enables rapid and accurate analysis of large-scale speech data, significantly improving the efficiency and accuracy of neurospeech disorder assessment.

[0075] The above explainable AI technique may be the SHAP technique. SHAP is an explainable AI technique based on the Shapley value of game theory, which quantitatively assesses the contribution of each unique feature to a predicted outcome. SHAP is used to interpret and explain the model's prediction results, and it fairly calculates the contribution of each feature, thereby enhancing model transparency. SHAP is particularly effective in explaining the predictions of complex machine learning models and can be utilized in a variety of domains.

[0076]

[0077] An example Python code for one embodiment of the present disclosure is as follows. It should be noted that the Python code below is intended for understanding and reproducing the operating method, not for actual use. The code is divided into a training section and an inference section. The training section tests multiple models to determine the LightGBM model as the final model, while the inference section includes functions that extract SHAP values ​​using the LightGBM model and explain the prediction results. The training section consists of the files exemplified below.

[0078] <train.py>

[0079]

[0080]

[0081]

[0082] <load_config 함수>

[0083] - Reads the JSON configuration file at the given path and converts the JSON data into a Python OrderedDict object.

[0084] <balance_dataset 함수>

[0085] - Deal with data imbalance by intentionally reducing data of a specific class in the dataset.

[0086] <Main Executive>

[0087] - Specify the configuration file path and load the configuration file to obtain configuration options.

[0088] - Load learning, verification, and test data and adjust the dataset through imbalance processing.

[0089] - Set feature variables and target variables in learning, verification, and test data.

[0090] - LightGBM model training is performed using learning and verification data.

[0091] - Make predictions on test data using the learned model.

[0092]

[0093] The LightGBM model training process in the above code requires multiple experiments to identify hyperparameter combinations that yield high performance evaluation metrics. To achieve this, libraries like Optuna and MLJAR can be used to iterate through multiple iterations to identify the hyperparameter combination that yields the highest model performance. Furthermore, please note that the code may be subject to modifications depending on changes in the target for classification or prediction.

[0094] The inference unit consists of the files exemplified below. These files contain a simplified version of the core code required for deploying the model in an actual service. Please note that the feature types and preprocessing steps specified in these files may be subject to change to improve model performance in the future.

[0095] <main_ml.py>

[0096]

[0097]

[0098] <AutumnPredictor 클래스>

[0099] * __init__ method

[0100] - Initialize the model by receiving 'result_path'.

[0101] - Create extract_ac, extract_sr, and extract_word instances using the classes written in the extract_acoustic.py, extract_speechrate.py, and extract_wer.py files.

[0102] * predict method

[0103] - Input a voice file and extract sound, speech rate, and word error rate through the 'extract_ac', 'extract_sr', and 'extract_word' instances.

[0104] - Combine the extracted features into one data frame and select the columns required for inference by referring to the data_info.json file.

[0105] - Perform predictions using the trained model, return prediction results and SHAP values, and save visualization results.

[0106] <generate_column_dict 함수>

[0107] - Create a dictionary by receiving the same features used in the training section as column names and a list of values ​​for the corresponding columns.

[0108] - If the value is a decimal, it is rounded to the third decimal place and stored.

[0109] <Main Executive>

[0110] - Set the path for saving voice files and results.

[0111] - Create an AutumnPredictor instance and perform a prediction.

[0112] - Outputs the feature list and prediction results.

[0113]

[0114] <extract_acoustic.py>

[0115]

[0116]

[0117]

[0118] <extract_acoustic_feature 클래스>

[0119] * __init__ method

[0120] - Initialize the dictionary for feature extraction.

[0121] - Set the sampling rate and the sound format to use.

[0122] * extract method

[0123] - Extract the following features from the audio file and return them in the form of a data frame.

[0124] - length: length of the audio file

[0125] - Formant: Average value of the first formant, average value of the second formant

[0126] - pitch: pitch value, mean, standard deviation, maximum, minimum

[0127] - intensity: negative intensity value, mean, standard deviation, maximum

[0128] - jitter: frequency fluctuation rate value, DDP (second derivative of adjacent jitter values), PPQ5 (variation of jitter values ​​in four adjacent analysis intervals)

[0129] - shimmer: Amplitude change rate value, APQ (amplitude change rate within 11 periods)

[0130] - hnr: negative sharpness value, logarithmic value

[0131] * _load_formant_pitch_intensity method

[0132] - Extract formant, pitch, and intensity from audio files.

[0133] * averaging_pitch method

[0134] - Convert pitch values ​​to mean, standard deviation, maximum, and minimum values.

[0135] * _averaging_formant method

[0136] - Convert the pitch value to the average according to the pitch number.

[0137]

[0138] <extract_speechrate.py>

[0139]

[0140]

[0141]

[0142] <extract_sr_feature 클래스>

[0143] * __init__ method:

[0144] - Initialize the dictionary for feature extraction.

[0145] * extract method

[0146] - Detect speech sections from a voice file and calculate the speech rate according to the following steps.

[0147] - Converts audio files to TextGrid objects and detects silence and speech segments.

[0148] - Calculate the length of the detected utterance section to obtain the total utterance time.

[0149] - Find the section where the intensity is above a certain level in the firing section and calculate the valid peak.

[0150] - The firing rate is measured by calculating the section considered as firing during the valid peak.

[0151] - Using the calculated firing rate, the following firing rate-related features are extracted and returned in the form of a data frame.

[0152] - speaking_rate: ratio of effective speaking peaks to total speaking length

[0153] - articulation_rate: effective articulation peak ratio to articulation interval length

[0154] - npause: Number of silence intervals detected

[0155] - asd: ratio of effective utterance peaks to total utterance interval length

[0156] - phon_ratio: ratio of utterance segment to total utterance length

[0157] - pause_dur: The length of the silence period, calculated by subtracting the length of the speech period from the total speech period.

[0158] - pause_freq_ratio: The ratio of the number of silence periods divided by the total speech length.

[0159] - pause_dur_ratio: Ratio of the number of silence periods to the length of the silence period.

[0160] - pauserate: the ratio of the length of silence to the total length of speech

[0161]

[0162] <extract_wer.py>

[0163]

[0164]

[0165] <extract_errorrate 클래스>

[0166] * __init__ method

[0167] - Initialize the dictionary for feature extraction.

[0168] - Set the sampling rate and whether to use the GPU.

[0169] - Load the Whisper model for speech to text tasks.

[0170] * extract method

[0171] - Load the Whisper model to the GPU or CPU.

[0172] - Converts speech files to text using the Whisper model using the _stt_whisper method.

[0173] - Calculate WER and CER by comparing the converted text with the predefined reference text (GROUND_TRUTH).

[0174] - Convert the calculated WER and CER values ​​to percentages, and limit the maximum value to 100.

[0175] * _stt_whisper method

[0176] - Converts speech files to text using the Whisper model.

[0177]

[0178] <data_info.json>

[0179]

[0180] data_info.json is a file that defines the list of columns to be used for inference, data type definitions, and parameters to be used by the inferor. The above code is an example for understanding the structure, and some content has been omitted.

[0181]

[0182] The inference unit using the above files operates in the following steps:

[0183] - The main_ml.py file instantiates and uses the extract_acoustic_feature, extract_sr_feature, and extract_errorrate classes defined in extract_acoustic.py, extract_speechrate.py, and extract_wer.py, respectively.

[0184] - When creating an instance of the AutumnPredictor class in main_ml.py, specify the path where the prediction model is stored.

[0185] - Call the predict method to receive a voice file as input and extract acoustic features, speech interval features, and word error rate features.

[0186] The extracted features are combined and input into a prediction model. The severity is predicted based on the predicted results and SHAP values, and the used features, predicted results, and SHAP values ​​are returned in dictionary form.

[0187]

[0188] The specific steps for classifying data by LightGBM, which was used as a classification model in the training unit, are as follows. The prediction function f(x) of the LightGBM model is expressed as a set of multiple trees. T is the number of trees, and γ is the number of trees. t is the weight of tree t, h t (x) represents the predicted value in tree t.

[0189]

[0190] LightGBM's learning process consists of the following steps:

[0191]

[0192] 1. Set the initial value of the model prediction value. Typically, the initial value is set to the average of the target column values.

[0193]

[0194] 2. Calculate the difference between the predicted value of the current model and the actual value to obtain the residual error (r i ( t )) and find a new tree h that minimizes the residual error. t Learn (x).

[0195]

[0196] 3. In the data sampling stage, the Gradient-based One-Side Sampling (GOSS) algorithm is used to further select samples with large residual errors calculated in the previous stage, and only some samples with small residual errors are selected to reduce the computational cost.

[0197] - GOSS assumes that in tree-based learning algorithms, instances with large gradients (differential values ​​at the point where the residual error occurs) are instances that have not been sufficiently learned, and that the more such instances are included in learning, the more the loss function can be lowered.

[0198] - When I is the training data, a is the sampling rate of data with large gradients, and b is the sampling rate of data with small gradients, select Ia as many top instances and Ib as many random instances.

[0199] 4. In the feature selection and combining process, the Exclusive Feature Bundling (EFB) algorithm is used to bundle mutually exclusive features into a single feature, thereby reducing the dimensionality and reducing the computational cost.

[0200] - EFB calculates the degree of mutual exclusivity between features (the degree to which when a specific feature has a value at a data point, another specific feature has a value of 0), creates a bundle of features, and creates a new feature by combining the features within the bundle.

[0201] 5. Train a new tree to predict the residual error using data with GOSS and EFB applied.

[0202] 6. Optimal value of leaf nodes (maximum number of leaves a tree can have) (γ) j ) is calculated.

[0203]

[0204] 7. Add the learned tree to the model to update the predicted values.

[0205]

[0206] 8. Repeat steps 2 through 7 to determine the model with the highest performance. At this point, an early stopping condition can be set to prevent overfitting to the training data.

[0207]

[0208] The SHapley Additive exPlanations (SHAP) value is based on the idea of ​​how to fairly distribute the contributions of each player in game theory. It calculates the importance of individual features by combining multiple features and represents the average change (average marginal contribution) depending on the presence or absence of the feature. In other words, the SHAP value of an individual feature is proportional to the gradient that the feature changes during model training, meaning that it is a value that can explain which feature played a major role when explaining the model's prediction results.

[0209] The specific process of calculating the SHAP values ​​used to explain the classification results in the inference section is as follows.

[0210] - Generate possible feature permutations. Let the number of features be i.

[0211]

[0212] Number of possible feature permutations = i!

[0213] - For each permutation, calculate the difference between the model predictions when feature i is added and when it is not added. Let v(S) be the model prediction for feature set S.

[0214]

[0215] Difference between model predictions =

[0216]

[0217] - Calculate the SHAP value by averaging the contributions from all permutations. The SHAP value for feature i is i When you say,

[0218]

[0219] Here, N represents the entire feature set, S represents a subset of features excluding feature i, v(S) represents the model prediction value for feature set S, |S| represents the size of set S, and |N| represents the total number of features.

[0220] FIG. 3 is a detailed diagram of a data processing pipeline of a neurolanguage disorder evaluation method according to one embodiment of the present disclosure, and FIG. 4 is a detailed diagram of a data processing pipeline in the case of paragraph reading in a neurolanguage disorder evaluation method according to one embodiment of the present disclosure.

[0221] Referring to Figure 3, the data processing pipeline consists of four main stages: an audio input stage, an audio preprocessing stage, a task-specific analysis stage, and a prediction and explanation stage.

[0222] 1. Audio input stage

[0223] The audio input stage receives recorded voice data for each diagnostic task. These tasks include maximum prolonged phonation time (MPT), direct dialect movement (DDK), word reading, and paragraph reading.

[0224] The MPT task involves taking a deep breath and pronouncing the vowels / a / , / i / , and / u / for as long as possible at a comfortable pitch and intensity. The MPT task is used to assess breathing and vocalization.

[0225] The DDK task consists of two subtasks: one requiring rapid and accurate repetition of one syllable each of / 퍼 / , / 터 / , and / 커 / , and the other requiring rapid and accurate repetition of three syllables of / 퍼터커 / . The DDK task is used to comprehensively assess the movement, speed, accuracy, and consistency of the articulatory organs.

[0226] The word reading task consists of 30 words based on the Urimal-Test of Articulation and Phonology (U-TAP). The task examines 43 consonants and 10 short vowels within a word, consisting of initial, medial, and final consonants. This task is used to identify the impact of the phonological environment, including consonant type and position, on the patient's articulation and to identify articulation error patterns.

[0227] The paragraph reading task involves reading "Autumn," a standardized Korean paragraph reading test. This task is used to assess various aspects of speech, including voice quality, speech rate, pauses, and speech intelligibility.

[0228] 2. Audio preprocessing stage

[0229] Recorded voice data for each task is input, and noise removal and segmentation are performed. The length of the input voice varies by task, and even within the same task, the length can vary significantly depending on the patient's severity, from approximately 30 seconds to 3 minutes. Therefore, segmentation into a manageable size is necessary. In the audio preprocessing stage, noise removal and voice activity detection (VAD) are used to identify speech segments, and speech is segmented into sentences based on areas where breathing is unusually long.

[0230] Afterwards, feature values ​​for modeling each task are extracted from the preprocessed voice data. This step utilizes acoustic and phonological features, and speech recognition results are also used as a feature value.

[0231] 3. Task-specific analysis steps

[0232] In the task-specific analysis step, features extracted through preprocessing are used to train a model for each task, and predicted results are output. Each model predicts unique quantitative measurements (e.g., number of pauses, number of repetitions, maximum phonation duration, etc.) and the patient's severity. The results from multiple tasks are combined to determine the patient's final severity.

[0233] 4. Prediction and explanation stage

[0234] In the prediction and explanation phase, detailed factors influencing the diagnostic results are analyzed and output to ensure reliability and transparency. The model's predicted outcomes and feature importance are analyzed to reveal the values ​​that impacted the patient's condition. Furthermore, data statistics, feature values, and predicted results are utilized to visualize the issues for each task at the speech production subsystem level. A large-scale language model (LLM) is used to provide analytical explanations for the diagnostic basis.

[0235] Each task proceeds through a specific preprocessing and feature extraction process, followed by analysis and prediction. For example, paragraph reading follows the flow chart in Figure 4.

[0236] For paragraph reading tasks, long audio segments of up to three minutes can be input. The input audio is segmented using VAD, which divides the audio into multiple segments. Each audio segment segmented by VAD is processed individually.

[0237] Each segment undergoes speech recognition using a model for the Speech-to-Text task (the Whisper ASR model in the example above), calculating the Word Error Rate (WER) and Character Error Rate (CER). Additionally, features are extracted from the speech data for each segment using a feature extractor.

[0238] Afterwards, the predicted severity results for each segment are synthesized, and the final prediction results are produced through an ensemble technique.

[0239] The above visualization information may include at least one of a graph, a chart, and text.

[0240] In another aspect of the present disclosure, a neurolanguage disorder assessment device for performing the neurolanguage disorder assessment method described above is provided, the neurolanguage disorder assessment device including a processor (not shown) and a memory (not shown).

[0241]

[0242] Figure 5a shows a speech assessment results screen that provides overall results classifying the severity of speech impairment after stroke.

[0243] A dashboard (1) is an interface designed to visually grasp important information for users at a glance. Dashboard (1) summarizes and displays various data and performance indicators, helping users quickly understand their current status and take necessary actions. For example, this dashboard (1) divides health information into various categories across tabs at the top, and uses a gauge display in the center to indicate advice or status, indicating the user's current health status at a "caution" level. Below this, there is an area that provides buttons for taking action or links to additional information. Thus, dashboards play a crucial role in integrating information and streamlining the user experience.

[0244] The dashboard (1) has a top where you can select the horse evaluation results (1-1) and training results.

[0245] The degree of speech impairment (1-2) can be expressed by classifying the severity into 0 / 1 / 2.

[0246] The severity levels are as follows:

[0247] 0 = normal (green)

[0248] 1 = Caution (yellow)

[0249] 2 = Severe (red)

[0250] A speech impairment degree visualization graph (1-3) is provided with a phrase (“Caution” in Fig. 5a).

[0251] The text is color-coded according to severity:

[0252] 0 = [Normal] = Green

[0253] 1 = [Caution] = Yellow

[0254] 2 = [Serious] = Red

[0255] Descriptive comments (1-4) regarding the degree of speech impairment are provided in text.

[0256] Example in Figure 5a: "Mr. Cheolsu's speech impairment level is [Caution]."

[0257] The View Description button (1-5) is a button that provides a summarized description in a pop-up form when the user clicks it.

[0258] Scroll indicators (1-6) contain animations that indicate to the user that they can scroll down.

[0259] This screen provides users with an intuitive assessment of the severity of their post-stroke speech impairment, and includes a description of the condition and the ability to view additional information.

[0260]

[0261] Figure 5b shows a screen that evaluates the severity of speech impairment after stroke and shows the results for each area in detail.

[0262] Results by area (2) indicate whether there are many areas that need improvement compared to the current normal category or many areas that are doing well among the factors that affected the accuracy of classification of speech impairment severity along with the speech impairment score.

[0263] When the user clicks the area-specific result viewer selection toolbar (2-1 to 2-3), only the results for the corresponding area are displayed (excellent area (2-1), overall results (2-2), and improvement area (2-3).

[0264] Scores and graphs (2-4) display the ratio by excellent / improvement area in a pie chart.

[0265] - Blue = Excellent area

[0266] - Red = Area for improvement

[0267] For example, a speech impairment score of 48 points is displayed, with excellent areas colored blue and areas of improvement colored red.

[0268] The result comment area (2-6) provides comments for areas with a higher percentage of excellent / improvement areas.

[0269] Example: "Our analysis of your speech indicates that many elements appear in [area for improvement]."

[0270] When the user clicks the Score Calculation Criteria button (2-7), for example, it moves to the Score Calculation Criteria pop-up.

[0271] This screen provides users with a visual representation of the severity and current status of their speech impairment in each area, and includes the ability to view detailed descriptions and additional information about the condition.

[0272]

[0273] FIG. 5c illustrates a screen visually providing results by region in a neurolanguage disorder evaluation device according to one embodiment of the present disclosure.

[0274] Results by field (3) show areas of improvement and areas of excellence in the speech production subsystems—respiration, generation, articulation, and prosody—among factors that influenced the classification of speech impairment severity. This visualizes which areas are performing well and which areas are experiencing the most problems.

[0275] The bar graph (3-1) of areas of excellence / improvement by field is indicated by color.

[0276] - Blue = Excellent area

[0277] - Red = Area for improvement

[0278] For example, graphs are shown for the four areas of breathing, generation, articulation, and prosody, and each area is distinguished by excellent and improved areas.

[0279] The indicator for the field with the most areas (3-2) is highlighted with an asterisk in that area.

[0280] An example of a comment (3-3) for the most extensive field is provided below.

[0281] - Example: "Of the four areas, there are many observations in the [vocalization] section."

[0282] When the user clicks the field description button (3-4), for example, it moves to a field-specific description pop-up.

[0283] - Example: It is presented in the form of a question, "What are the four areas?"

[0284] This screen helps users understand the results of the speech impairment severity assessment in detail by area, and provides the function to visually check the areas in need of improvement and areas of excellence in each area.

[0285]

[0286] FIG. 6 illustrates a screen visually providing detailed results of excellent areas and improvement areas in a neurolanguage disorder evaluation device according to one embodiment of the present disclosure.

[0287] That is, Fig. 6 is a detailed result, and displays the three factors with the highest contribution among the excellent and improved areas as the excellent area ranking (4) and the improved area ranking (5).

[0288] The ranking of excellent areas (4) shows the three areas with the highest contribution in the excellent area feature and sorts them in ascending order.

[0289] The feature name (4-1) indicates the name of the superior domain. Example: "Pitch"

[0290] The feature icon (4-2) visually displays the icon in the corresponding area.

[0291] Feature Situational Comments (4-3) provide situational comments for the relevant area. Example: "The sound pitch is close to normal."

[0292] The See More button (4-4) is a button that moves to a detailed description screen for each feature.

[0293] The improvement area ranking (5) shows the three most contributing features in the improvement area, sorted in descending order.

[0294] The feature name (5-1) indicates the name of the improvement area. For example, "Sound volume"

[0295] The feature icon (5-2) visually displays the icon in the corresponding area.

[0296] Feature Situational Comments (5-3) provide situational comments for the relevant area. Example: "The volume was a bit low."

[0297] The See More button (5-4) is a button that moves to a detailed description screen for each feature.

[0298] This screen helps users identify specific areas of strength and areas requiring improvement in their neurolanguage impairment assessment results. By providing detailed analysis results for each area, it provides users with the information they need to develop specific improvement plans and strengthen their strengths.

[0299]

[0300] FIG. 7 illustrates a detailed description screen for one of the areas requiring improvement in a neurolanguage disorder assessment device according to one embodiment of the present disclosure.

[0301] Figure 7 is a detailed description screen (7), and when a feature of a detailed result is clicked one by one in the overall result (2-2) (see Figure 5b), it moves to a detailed description screen and provides a detailed description for each feature.

[0302] When the back button (7-1) is clicked, the screen returns to the full results (2-2) (see Fig. 5b).

[0303] Feature Name and Icon (7-2) displays the name and icon of the feature. For example, when you click on the "Voice Pitch" feature, an icon related to the description of the sound pitch is displayed.

[0304] The Improvement Area 1 (7-3) area provides a ranking of three areas for improvement. For example, among the three areas requiring improvement, pitch is indicated as the first area requiring improvement.

[0305] In the "Voice Pitch" section (7-4), a descriptive comment about the feature is provided, providing a detailed explanation of the feature. For example, the description for "Voice Pitch" is provided as follows: "What is voice pitch? It refers to the high and low notes of the voice. If the voice pitch is not good, speech may sound unnatural or it may be difficult to convey emotion through the voice."

[0306] This screen provides users with more specific and detailed information about the results of specific assessment items, helping to increase understanding of neurolinguistic disorder assessments and develop concrete improvement plans.

[0307] FIG. 8 is a comparison result screen provided by a neurolanguage disorder evaluation device according to one embodiment of the present disclosure, which visually compares the user's current status to the extent of the problem compared to the normal range by taking into account age, gender, etc.

[0308] Figures 8a to 8c are screens comparing the average, respectively, and indicate how problematic each improvement factor is compared to the normal range, taking into account age, gender, etc.

[0309] What is my level in Fig. 8a to Fig. 8c? (8) The screen visually displays the results of the evaluation based on the user's current status.

[0310]

[0311] FIG. 8A visually represents the user's voice pitch compared to other people, according to one embodiment of the present disclosure.

[0312] The upper area (8a-1) visually displays the results evaluated based on the user's current status.

[0313] The user area (8a-2) indicates the user's voice pitch. For example, Cheolsu's voice pitch is displayed in the 30-50 Hz range.

[0314] The subfield (8a-3) visually displays the comparison to the average, allowing the user to compare their averages with those of others.

[0315] The note area (8a-4) indicates how much the user's voice pitch differs compared to the average for a man in his 50s (45-60 Hz), for example.

[0316] The comment section (8a-5) provides a description of the user's voice pitch. Example: "Cheolsu's voice pitch is 30-50 Hz (hertz). Compared to other people, it is 10-15 Hz lower."

[0317]

[0318] FIG. 8b visually illustrates a user's pronunciation accuracy compared to other people, according to one embodiment of the present disclosure.

[0319] The central bar graph (8b-1) is a color-coded bar graph, categorized as "Insufficient," "Effort," "Average," and "Excellent," visually displaying the user's current status based on the evaluation results. Each score indicates the level of pronunciation accuracy.

[0320] - Color classification:

[0321] - Deficiency: 0-25 points (red)

[0322] - Effort: 26-50 points (orange)

[0323] - Average: 51-75 points (yellow)

[0324] - Excellent: 76-100 points (green)

[0325] The user display area (8b-2) visually displays, for example, the user's pronunciation accuracy, evaluated as 73 points. A score of 73 corresponds to "average" and is highlighted with a star icon above the bar graph.

[0326] The pronunciation accuracy comment section (8b-3) provides an explanation of pronunciation accuracy. Example: "Cheolsu's pronunciation accuracy is 73 points. Compared to other people, it is 00 points lower."

[0327] Figure 8b illustrates a visualization result provided by a neurolanguage disorder assessment device, which helps users understand their pronunciation accuracy status, recognize differences from the average, and identify directions for improvement.

[0328] FIG. 8c visually represents a user's [purse] repetition count compared to other people, according to one embodiment of the present disclosure.

[0329] Graph (8c-1) visually displays the user's evaluation of the number of [pur] repetitions.

[0330] The average area (8c-2) visually displays the user's number of repetitions compared to the average.

[0331] The average description area (8c-3) shows the average number of repetitions for that age group. For example, the average for a man in his 50s is 5.5 repetitions.

[0332] The "I" area (8c-3) visually displays the user's [Per] repetition count.

[0333] The user area (8c-5) displays the user's [Per] repetition count as a number, along with a star icon. For example, *3.2 times. This indicates that the user repeats [Per] 3.2 times per second.

[0334] The comment section (8c-6) provides an explanation of the user's [Per] repetition count: "Cheolsu's [Per] repetition count is 3.2 times per second. Compared to other people, it is 00 times lower."

[0335] [Per] The repetition count measures the number of times a user repeatedly pronounces a specific syllable within a given period of time. This metric is used to assess a user's pronunciation accuracy and fluency, and primarily focuses on examining the repeatability of pronunciation.

[0336] [Per] The repetition count measures how accurately and quickly a user can repeat the syllable "per" within a given period of time. Inaccurate or slow pronunciation may indicate problems related to neurological speech disorders or dysarthria.

[0337] Evaluation Example: Figure 8c shows that Cheolsu's [per] repetition rate is 3.2 times per second, while the average [per] repetition rate for men in their 50s is 5.5 times per second. Cheolsu's lower-than-average [per] repetition rate indicates that his pronunciation is less repetitive than average. This suggests the possibility of a neurospeech disorder or pronunciation disorder, which may require additional treatment or training.

[0338] Figure 8c illustrates the visualization results provided by the neurolanguage disorder assessment device, helping users understand their condition, recognize differences from the average, and identify directions for improvement.

[0339]

[0340] Figure 9 is a screen visually representing the predicted results of a patient's speech impairment score, according to one embodiment of the present disclosure. This screen allows the user to compare their current score with their future score if certain factors are improved. This helps the patient clearly understand the factors that need improvement and the expected outcomes.

[0341] In the Title (9) area, display a phrase such as "What if the number of repetitions improves?"

[0342] The Current Score area (9-1) displays the user's current speech impairment score. For example, 48 points.

[0343] The Future Score section (9-2) shows the projected future score based on specific improvement factors. For example, 53 points.

[0344] The Score Change Prediction Area (9-3) calculates the difference between your current and future scores, indicating how much your score will increase with improvement. Example: Expected score change: +5 points

[0345] The Predictive Explanation Comment section (9-4) specifically describes the expected changes resulting from improving the relevant element. For example, "If the number of repetitions approaches the average, the speech impairment score is expected to increase by 5 points."

[0346] The current score area (9-1) visually displays the user's current score, allowing them to grasp their current status. The future score area (9-2) visually presents the expected score after improvement, suggesting the need for improvement and its goals. The score change prediction area (9-3) emphasizes the importance and expected effects of improvement through score changes. The prediction explanation comment area (9-4) conveys the specific effects of improvement in an easily understandable manner through concise explanations.

[0347] Figure 9 plays an important role in setting the direction of treatment and providing motivation by clearly showing the user the current status and potential for improvement.

[0348] Figure 10 is a screen that visually displays the cumulative results of a patient's speech impairment assessment score, according to one embodiment of the present disclosure. By graphically representing changes in the assessment score over time, it allows for a quick overview of whether the patient is making continuous improvements. Specifically, Figure 10 visualizes changes in the patient's speech impairment assessment score over time as a graph, representing the cumulative assessment results, allowing both the patient and medical staff to easily assess the effectiveness of the treatment process.

[0349] The title area (10) is labeled “Cumulative Results Trend” and visually displays the cumulative evaluation results.

[0350] In the cumulative score graph (10-1), the x-axis represents the evaluation round, the y-axis represents the evaluation score, and the scores for each evaluation round are visualized by connecting them in a line graph.

[0351] - Example: 1st round (23 / 12 / 04), 2nd round (24 / 01 / 04), 3rd round (24 / 02 / 04)

[0352] In the score display area by round (10-2), the scores for each evaluation round are displayed as dots on the graph.

[0353] The current score indicator (10-3) highlights the current evaluation score. Example: 3rd round score: 58 points

[0354] The comment section (6-4) for cumulative results describes the current evaluation score and the change from the previous evaluation at the bottom of the graph. Example: "Your score for the third evaluation is 58 points. This is a 12-point increase from the previous evaluation."

[0355] The cumulative score graph (10-1) helps you grasp the change in evaluation scores over time at a glance. It connects scores from each evaluation round with a line, providing a visual flow.

[0356] By displaying the scores for each round (10-2), the scores for each evaluation round are clearly displayed on the graph, allowing for easy comparison of the scores for each round.

[0357] The current score indicator (10-3) highlights the scores of the current round, allowing you to see the most recent evaluation results at a glance.

[0358] By explaining the change in score compared to the previous evaluation along with the current evaluation score in the comment area (10-4) for the cumulative result, it helps to clearly understand the degree of improvement.

[0359] Figure 10 provides patients and medical staff with visual representations of ongoing assessment results, providing useful information for clearly understanding the effectiveness of treatment and developing future treatment plans.

[0360]

[0361] The neurolanguage disorder assessment method (200) of the present disclosure may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, those skilled in the art will appreciate that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0362] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner (200). The software and data may be stored on one or more computer-readable recording media.

[0363] The described embodiments of the present disclosure can also be practiced in distributed computing environments, where certain tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.

[0364] Although the embodiments described above have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the above. For example, appropriate results can be achieved even if the described techniques are performed in a different order than the described method (200), and / or components of the described system, structure, device, circuit, etc. are combined or combined in a different form than the described method (200), or are replaced or substituted with other components or equivalents.

[0365] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. In a method for evaluating the severity of neurolanguage disorder using explainable artificial intelligence by a server, A step of receiving patient voice data collected through a user terminal by the above server, A step of performing audio preprocessing and voice segmentation processing on the voice data by the above server, A step of extracting speech impairment specific features from the processed speech data by the above server, A step of classifying the severity of speech impairment based on the extracted specific features using a tree-based machine learning algorithm by the above server, A step of generating visualization information by analyzing and visualizing the contribution of specific features that have influenced the prediction using an explainable artificial intelligence technique by the above server, A step of generating an LLM generation sentence in natural language format as a basis for the prediction result by using an LLM (Large Language Model) by the above server. A method for evaluating neurolinguistic disorders, characterized by including:

2. In paragraph 1, A method for evaluating a neurolanguage disorder, characterized in that the above-mentioned specific features include at least one of pitch, intensity, jitter (%), shimmer (%), harmonic-to-noise ratio, speaking rate, articulation rate, pause frequency ratio, pause length, and character accuracy.

3. In paragraph 2, A method for evaluating neurolanguage disorders, characterized in that it further comprises a step of providing the visualization information and the LLM generation statement to the user terminal by the server.

4. In paragraph 3, A neurolanguage disorder assessment method characterized in that the tree-based machine learning algorithm is LightGBM (Light Gradient Boosting Machine).

5. In paragraph 4, A neurolanguage disorder assessment method characterized in that the above-described artificial intelligence technique is the SHAP (SHapley Additive exPlanations) technique.

6. In paragraph 5, A method for evaluating neurolinguistic disorders, characterized in that the visualization information includes at least one of a graph, a chart, and text.

7. In paragraph 6, The steps of performing the above audio preprocessing and voice segmentation processing are A method for assessing neurolanguage disorders, comprising the steps of segmenting speech segments through noise removal and voice activity detection (VAD).

8. A neuro-language disorder evaluation device that performs a neuro-language disorder evaluation method according to any one of clauses 1 to 7, including a processor and memory, Neurolanguage disorders assessment device.

Citation Information

Patent Citations

  • Bearing Fault Diagnosis Device and Diagnosis Method Using First Order Deadbeat Observer

    KR1020230149109A

  • System and application for evaluation of voice and speech disorders and speech-language therapy customized for parkinson patients

    KR102668964B1

  • Touch panel with iso-resistance compensation pattern

    KR102909821B1

  • Predicting Intensive Care Transfers And Other Unforeseen Events Using Machine Learning

    US20200160998A1

  • Systems and Methods to generate a personalized medical summary (PMS) from a practitioner-patient conversation.

    US20240029901A1