Speech disorder evaluation system based on text task

CN120077429APending Publication Date: 2025-05-30ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074134.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing verbal dysfunction detection methods rely on the experience and subjective judgment of professionals. The evaluation time is long and the results fluctuate greatly; while neuroimaging analysis is accurate, it needs to be carried out in the hospital, which is complex and expensive.

Method used

Design a speech disorder assessment system based on text tasks. By collecting speech signals, signal preprocessing and noise reduction processing, calculating response time and pause situations, obtaining acoustic features, evaluating sound structure, word formation, grammar and semantic understanding abilities, and finally calculating the speech function evaluation results.

Benefits of technology

Real-time and multi-dimensional speech function evaluation is realized, reducing dependence on professionals, simplifying the detection process, and improving the accuracy and efficiency of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077429A_ABST
    Figure CN120077429A_ABST
Patent Text Reader

Abstract

A speech disorder evaluation system based on a text task comprises an acquisition module; a processing module; the first calculation module is used for calculating response time, midway abnormal pause times and pause duration, and evaluating a word finding function abnormal probability and a text understanding level according to the response time, the midway abnormal pause times and the pause duration; the acquisition module is used for acquiring basic acoustic features and deep acoustic features of the voice signals; the first evaluation module is used for evaluating a phonetic formation capability level and a word formation capability level according to the basic acoustic features and the deep acoustic features; the second evaluation module is used for evaluating a grammar function level and a semantic understanding capability level according to the original text information and the voice information; and the second calculation module is used for calculating a final evaluation result according to the word finding function anomaly probability, the text understanding level, the phonetic capability level, the word formation capability level, the grammar function level and the semantic understanding capability level. According to the speech disorder assessment system based on the text task, the speech information of the testee is collected to obtain multi-dimensional information, then the speech function of the testee is assessed by integrating the information of all the dimensions, and the speech expression ability of the testee is fed back in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Speech impairment assessment system based on text tasks Technical Field

[0001] The present invention particularly relates to a speech disorder assessment system based on text tasks. Background Art

[0002] The main types of speech dysfunction include aphasia, dysarthria, dysphonia, and other symptoms that affect normal language expression. Currently, they are mainly detected through scale analysis, neuroimaging technology, and biomarker analysis such as serum and cerebrospinal fluid testing.

[0003] However, existing technical means have the following problems: high professional requirements for testers, scale analysis is more dependent on the experience and subjective judgment of psychological assessors, the evaluation time is long, and the evaluation results fluctuate widely; neuroimaging analysis can achieve highly accurate diagnostic results, but magnetic resonance imaging (MRI) or positron emission tomography (PET) tests must be performed in a hospital, the analysis process is complex and expensive, and professional doctors' evaluation and guidance are required.

[0004] Summary of the Invention

[0005] The present invention provides a text-based task-based speech impairment assessment system to solve the above-mentioned technical problems, specifically adopting the following technical solutions:

[0006] A text-based task-based speech impairment assessment system, comprising:

[0007] Acquisition module, used for collecting voice signals;

[0008] A processing module, configured to perform signal preprocessing and noise reduction on the collected speech signal;

[0009] A first calculation module is used to calculate the response time, the number of abnormal pauses and the duration of the pauses, and to evaluate the probability of abnormal word-finding function and text comprehension level based on the response time, the number of abnormal pauses and the duration of the pauses;

[0010] An acquisition module, configured to acquire basic acoustic features and deep acoustic features of the speech signal;

[0011] a first evaluation module, configured to evaluate the articulation ability level and the word formation ability level according to the basic acoustic features and the deep acoustic features;

[0012] A second evaluation module is used to evaluate the grammatical function level and the semantic comprehension ability level based on the original text information and the voice information;

[0013] The second calculation module is used to calculate the final evaluation result according to the abnormal probability of the word-finding function, the text comprehension level, the articulation ability level, the word-forming ability level, the grammatical function level and the semantic comprehension ability level.

[0014] Furthermore, the specific method of collecting the voice signal by the collection module is:

[0015] Designing specific text tasks;

[0016] The subjects read the text aloud and repeated the content;

[0017] The speech signals are collected when the subjects perform text reading and content retelling.

[0018] Furthermore, the specific method of the processing module performing signal preprocessing and noise reduction processing on the collected voice signal is:

[0019] The collected audio data is processed using a speech signal enhancement algorithm based on the LSTM model.

[0020] Furthermore, the basic acoustic features include: fundamental frequency features, MFCC, LFCC and resonance peaks.

[0021] Furthermore, the basic acoustic features also include a proportion.

[0022] Furthermore, -32 decibels is set as the critical value between sound and silence, and sounds less than -32 decibels are regarded as silent frames to calculate the proportion.

[0023] Furthermore, the acquisition module extracts the deep acoustic features by inputting the speech signal into a deep acoustic feature extraction model.

[0024] Furthermore, the deep acoustic feature extraction model includes a feature extraction module based on convolutional neural network, a feature quantization module and a global feature extraction module based on Transformer. The feature extraction module based on convolutional neural network abstracts the one-dimensional speech signal through a convolutional neural network, performs a frame operation with overlapping areas on the speech signal, and extracts features from each frame of the speech signal, thereby obtaining a potential speech feature sequence Z. The feature quantization module quantizes the continuous speech feature sequence Z into discrete features Q through multiplication. The global feature extraction module based on Transformer globally models the potential semantic feature sequence through the self-attention mechanism of the Transformer encoder module, thereby obtaining a more global semantic representation C.

[0025] Furthermore, the specific method for the first evaluation module to evaluate the articulation ability level and the word formation ability level according to the basic acoustic features and the deep acoustic features is:

[0026] The basic acoustic features and the deep acoustic features are concatenated and input into a classification rating model to obtain the articulation ability level and the word formation ability level.

[0027] Furthermore, the specific method of the first evaluation module evaluating the grammatical function level and the semantic comprehension ability level based on the original text information and the voice information is:

[0028] Converting the voice information into text information;

[0029] Calculating the similarity between the text information and several key grammatical points of the preset original text to obtain the grammatical function level;

[0030] The similarity between the original text and the text information is calculated to obtain the semantic understanding ability level.

[0031] The benefit of the present invention lies in the text-task-based speech disorder assessment system it provides, which collects the subject's voice information in real time, obtains multi-dimensional information through the voice information, and then evaluates the tester's speech function based on the above-mentioned dimensional information, providing real-time feedback on the subject's speech expression ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0033] FIG1 is a schematic diagram of a text-based task-based speech impairment assessment system of the present invention;

[0034] FIG2 is a schematic diagram of a portion of a speech signal enhancement processing model according to the present invention;

[0035] FIG3 is a schematic diagram of another part of the speech signal enhancement processing model of the present invention;

[0036] FIG4 is a schematic diagram of a deep acoustic feature extraction model of the present invention. DETAILED DESCRIPTION

[0037] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0038] As shown in Figure 1, a text-task-based speech disorder assessment system of the present application comprises: an acquisition module, a processing module, a first calculation module, an acquisition module, a first evaluation module, a second evaluation module, and a second calculation module. The acquisition module is used to acquire speech signals. The processing module is used to perform signal preprocessing and noise reduction on the acquired speech signals. The first calculation module is used to calculate the response time, the number of abnormal pauses, and the duration of the pauses, and to assess the probability of abnormal word-finding function and the level of text comprehension based on the response time, the number of abnormal pauses, and the duration of the pauses. The acquisition module is used to acquire the basic acoustic features and deep acoustic features of the speech signal. The first evaluation module is used to assess the articulation ability level and word-forming ability level based on the basic acoustic features and deep acoustic features. The second evaluation module is used to assess the grammatical function level and semantic comprehension ability level based on the original text information and speech information. The second calculation module is used to calculate the final assessment result based on the probability of abnormal word-finding function, text comprehension level, articulation ability level, word-forming ability level, grammatical function level, and semantic comprehension ability level. Through the text-task-based speech disorder assessment system of this application, the subject's sample speech information is collected in real time, the response time and pause interval of the subject during the text reading process are recorded, and the subject's word-finding and comprehension ability is evaluated. Through speech analysis technology, the subject's pitch and rhythm during the reading process are tested, and the subject's articulation ability and comprehension ability are evaluated. Through the natural language understanding algorithm, the subject's vocabulary expression during the retelling task is analyzed, and the subject's grammatical function and semantic comprehension ability are evaluated. Through the comprehensive evaluation of the speech function modules in the above dimensions, real-time feedback on the subject's speech expression ability is provided.

[0039] Specifically, the specific method of collecting voice signals through the collection module is:

[0040] Design specific text tasks. When designing text tasks, emphasize storytelling and transitional content, focusing on assessing the test-taker's word-finding, grammar, comprehension, and pronunciation skills. The test-taker will read and retell the text. While the test-taker is reading and retelling, use an audio device such as a recorder or microphone to capture and archive the test-taker's voice signals.

[0041] Environmental noise and device background noise are major interferences in speech analysis. As a preferred embodiment, the specific method for the processing module to perform signal preprocessing and noise reduction on the collected speech signal is: using a speech signal enhancement algorithm based on the LSTM model to process the collected audio data. The output of the encoder is connected to the time series model LSTM for time series feature extraction, and the decoder uses the features extracted by LSTM and the shallow encoder features to enhance the estimation of the speech signal. As shown in Figures 2 and 3, the encoder E consists of L encoders {E1, E2, ..., E L}, any encoder E i (i>1) consists of a convolution layer (convolution kernel is K, stride is S, output channel is 2 i-1 H), ReLU activation layer, 1×1 convolution layer (step size S, input channel 2 i-2 H, output channel is 2 i-1 H) and GLU activation layer. The time series model uses LSTM, which has two layers, and the hidden layer size of each layer is 2 L-1 H.

[0042] The decoder D and the encoder adopt a symmetrical structure. For the sake of symmetry and analysis convenience, they are numbered in reverse order, that is, D = {D L ,…,D2,D1}. Decoder D i The input is composed of the previous layer decoder (which may also be an LSTM model) and E i The output of the decoder D is added to generate i The structure and E i is roughly symmetrical, in E i The downsampling convolution layer in corresponds to the transposed convolution, which achieves the upsampling effect and is used to restore the timing information to obtain the audio data after noise reduction.

[0043] The first calculation module calculates the response time τ from the start of the text narration, the number of abnormal pauses N and the duration of these pauses I, and assesses the probability of word-finding malfunction P and the text comprehension level D. A longer response time τ indicates a higher probability of word-finding malfunction P. A greater number of abnormal pauses N and the duration of these pauses I indicate a lower text comprehension level D. Based on the scores, the probability of word-finding malfunction P is assigned a rating ranging from 0 to 100% and the text comprehension level D is assigned a rating from 1 to 5.

[0044] Response time τ: The time from the start of the task to the start of the subject's voice is the response time τ.

[0045] Number of abnormal pauses N: Calculate the silence time, that is, a pause is defined as a time when the sound level is less than -32 decibels for more than 3 consecutive seconds.

[0046] Pause duration I: The total number of abnormal pauses corresponding to the duration is accumulated as pause duration I.

[0047] For example, the probability P of abnormal word search function is:

[0048] Specifically, 10 paragraphs of text are set, and the response time τ recorded at the beginning of each text paragraph is i , where τ i is the response time of paragraph i, P = ∑τ i

[0049] Points are assigned based on the response time: 0-1s: Level 1, 1-3s: Level 2, 3-6s: Level 3, 6-9s: Level 4, and more than 10s: Level 5.

[0050] For text comprehension level D:

[0051] According to the number of abnormal pauses N,

[0052] 0 times: 1 point;

[0053] 1-2 times: 2 points;

[0054] 3-5 times: 3 points;

[0055] 6-9 times: 4 points;

[0056] More than 10 times: 5 points.

[0057] The abnormal pause ratio is,

[0058] Among them, l k is the duration of the kth abnormal pause, and L is the total duration of the audio.

[0059] Points are assigned according to percentage: 0-20%: 1 point, 20-40%: 2 points, 40-60%: 3 points, 60-80%: 4 points, 80-100%: 5 points.

[0060] The number of abnormal pauses and the pause ratio calculation score are added together, where 2-3 points: Level 1, 4-5 points: Level 2, 6-7 points: Level 3, 8-9 points: Level 4, and 10 points: Level 5.

[0061] The greater the number of abnormal pauses N and the pause duration I, the lower the text comprehension level D.

[0062] In the real-time method of this application, basic acoustic features include: fundamental frequency features, MFCCs, LFCCs, and formants. By performing signal processing on the audio to extract first-order and second-order differences, acoustic features including fundamental frequency features, MFCCs, LFCCs, formants, etc. are obtained.

[0063] Preferably, the basic acoustic features further include a ratio. Specifically, -32 decibels is set as the critical value between sound and silence, and sounds less than -32 decibels are regarded as silent frames to calculate the ratio.

[0064] In the application, the acquisition module inputs the speech signal into the deep acoustic feature extraction model to extract deep acoustic features.

[0065] As shown in Figure 4, as a preferred embodiment, the deep acoustic feature extraction model includes a convolutional neural network-based feature extraction module, a feature quantization module, and a Transformer-based global feature extraction module. The convolutional neural network-based feature extraction module abstracts the one-dimensional speech signal X through a convolutional neural network (CNN). The effect is equivalent to framing the speech signal X with overlapping regions and extracting features from each frame of the speech signal, thereby obtaining a latent speech feature sequence Z. The feature quantization module quantizes the continuous speech feature sequence Z into discrete features Q through multiplication. The Transformer-based global feature extraction module globally models the latent semantic feature sequence through the self-attention mechanism of the Transformer encoder module, thereby obtaining a more global semantic representation C.

[0066] The model is pre-trained through mask prediction tasks and contrastive learning, as shown in the figure below. The predicted mask feature c t The K+1 quantitative features q around t are t Constitute a positive sample sequence, and other quantitative features Construct a negative sample sequence, and then use the formula Evaluate the model prediction results, where the similarity function is sim(a,b)=a T b / ‖a‖‖b‖.

[0067] As a preferred embodiment, the specific method for the first evaluation module to evaluate the articulation ability level and word-forming ability level based on the basic acoustic features and the deep acoustic features is: the basic acoustic features and the deep acoustic features are concatenated and input into the classification rating model to obtain the articulation ability level and the word-forming ability level. Specifically, a cascade connection method is adopted to perform a connection operation along the feature axis, and the obtained traditional acoustic features and the deep acoustic features are concatenated. Then, support vector machines (SVM), neural networks (NN) and other modules are used to evaluate and classify the features. According to the compared and calibrated graded speech features (ratings of 1 to 5), the subjects' articulation ability A is rated 1 to 5 and their word-forming ability U is rated 1 to 5.

[0068] The specific method of the first evaluation module to evaluate the grammatical function level and semantic comprehension ability level based on the original text information and voice information is as follows:

[0069] Convert speech information into text information. Based on the Automatic Speech Recognition (ASR) model, the speech signal is converted into text information.

[0070] The grammatical function level is calculated by calculating the similarity between the text and a predetermined number of key grammatical points in the original text. N key grammatical points are pre-set in the original text, and the text of the subject's speech is compared one by one to obtain a similarity comparison. If the texts are similar, the grammatical points are confirmed to be correct; if they are incorrect, M points are accumulated. The M incorrect points are divided by the total number of grammatical points N to obtain the percentage. The percentage is assigned in 20% increments: 0-20%: Level 1, 20-40%: Level 2, 40-60%: Level 3, 60-80%: Level 4, and 80-100%: Level 5.

[0071] The semantic understanding ability level is calculated by calculating the similarity between the original text and the text information. Specifically, the similarity ratio is calculated by comparing the speech-to-text information with the original text. Points are assigned based on the percentage in 20% increments: 0-20%: Level 1, 20-40%: Level 2, 40-60%: Level 3, 60-80%: Level 4, and 80-100%: Level 5.

[0072] For step S7: the final evaluation result is calculated based on the probability of abnormal word-finding function, text comprehension level, articulation ability level, word-forming ability level, grammatical function level and semantic comprehension ability level.

[0073] Specifically, the final score can be obtained by adding up the points of each indicator. The higher the score, the worse the speech ability.

[0074] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solutions obtained by equivalent replacement or equivalent transformation fall within the scope of protection of the present invention.

Claims

1. A text-based task-based speech impairment assessment system, characterized in that: Include: A collection module, used for collecting voice signals; A processing module, used for performing signal preprocessing and noise reduction processing on the collected speech signal; A first calculation module is used to calculate the response time, the number of abnormal pauses and the duration of the pauses, and to evaluate the probability of abnormal word search function and the level of text comprehension according to the response time, the number of abnormal pauses and the duration of the pauses; An acquisition module, used to acquire basic acoustic features and deep acoustic features of the speech signal; A first evaluation module, used for evaluating the articulation ability level and the word formation ability level according to the basic acoustic features and the deep acoustic features; A second evaluation module is used to evaluate the grammatical function level and the semantic comprehension ability level according to the original text information and the voice information; The second calculation module is used to calculate the final evaluation result according to the probability of abnormal word-finding function, the text comprehension level, the articulation ability level, the word-forming ability level, the grammatical function level and the semantic comprehension ability level.

2. The text-task-based speech impairment assessment system according to claim 1, characterized in that: The specific method of collecting the voice signal by the collection module is: Designing specific text tasks; The subjects read the text aloud and repeated the content; The speech signals are collected when the subjects perform text reading and content retelling.

3. The text-task-based speech impairment assessment system according to claim 1, characterized in that: The specific method of the processing module performing signal preprocessing and noise reduction processing on the collected speech signal is: The collected audio data is processed using a speech signal enhancement algorithm based on the LSTM model.

4. The text-task-based speech impairment assessment system according to claim 1, characterized in that: The basic acoustic features include: fundamental frequency features, MFCC, LFCC and resonance peaks.

5. The text-task-based speech impairment assessment system according to claim 4, characterized in that: The basic acoustic features also include a proportion.

6. The text-task-based speech impairment assessment system according to claim 5, characterized in that: -32 decibels is set as the threshold between sound and silence, and sounds less than -32 decibels are regarded as silent frames to calculate the percentage.

7. The text-task-based speech impairment assessment system according to claim 1, characterized in that: The acquisition module extracts the deep acoustic features by inputting the speech signal into a deep acoustic feature extraction model.

8. The text-task-based speech impairment assessment system according to claim 7, characterized in that: The deep acoustic feature extraction model includes a feature extraction module based on a convolutional neural network, a feature quantization module and a basic The invention relates to a global feature extraction module based on Transformer, wherein the feature extraction module based on convolutional neural network abstracts the one-dimensional speech signal through a convolutional neural network, performs a frame operation with overlapping areas on the speech signal, and extracts features from each frame of the speech signal, thereby obtaining a potential speech feature sequence Z, the feature quantization module quantizes the continuous speech feature sequence Z into discrete features Q through multiplication, and the global feature extraction module based on Transformer globally models the potential semantic feature sequence through the self-attention mechanism of the Transformer encoder module, thereby obtaining a more global semantic representation C.

9. The text-task-based speech disorder assessment system according to claim 1, characterized in that: The specific method of the first evaluation module evaluating the articulation ability level and the word formation ability level according to the basic acoustic features and the deep acoustic features is: The basic acoustic features and the deep acoustic features are concatenated and input into a classification rating model to obtain the articulation ability level and the word formation ability level.

10. The text-task-based speech impairment assessment system according to claim 1, characterized in that: The specific method of the first evaluation module evaluating the grammatical function level and the semantic understanding ability level according to the original text information and the voice information is: Converting the voice information into text information; Calculating the similarity between the text information and several key grammatical points of the preset original text to obtain the grammatical function level; The similarity between the original text and the text information is calculated to obtain the semantic understanding ability level.