Cerebral stroke early recognition method and system based on continuous vowel and continuous voice fusion

By integrating continuous vowel and continuous speech recognition methods, the VocGAN and E-CDNN-SRU models are used to solve the problems of high requirements for Mandarin and insufficient environmental adaptability in the prior art, and high-accurate stroke screening in different environments is achieved.

CN120356489APending Publication Date: 2025-07-22SOUTHEAST UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510663974.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing pronunciation-based stroke recognition technology has problems such as high requirements for Mandarin level, insufficient environmental adaptability and small amount of pathological speech data, resulting in limited recognition accuracy and practicality.

Method used

Using the recognition method based on continuous vowel and continuous speech fusion, the training data set is expanded through the VocGAN generation model, combined with the E-CDNN-SRU speech enhancement model to suppress noise, used openSMILE features and StrokeConvNet network for identification, and stroke risk assessment is performed through mutual information method and weight fusion decision.

Benefits of technology

It improves the accuracy and universality of stroke recognition, can maintain stable recognition performance in non-ideal acoustic environment, and is suitable for stroke screening in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356489A_ABST
    Figure CN120356489A_ABST
Patent Text Reader

Abstract

The invention discloses a cerebral apoplexy early recognition method and system based on continuous vowel and continuous voice fusion. The method comprises the following steps: collecting continuous vowel data and continuous voice data; generating pathological voice data; performing voice enhancement; recognizing the continuous vowels, and obtaining a stroke prediction result by adopting a recognition model based on the continuous vowels; recognizing continuous voices, and adopting a recognition model based on the continuous voices to realize classification of health and stroke risks; and performing fusion decision according to the independent weights of the continuous vowel recognition model and the continuous voice recognition model, and judging the stroke risk probability of each sample. According to the method, bimodal analysis of continuous vowels and continuous voices is fused, advantages are complementary, and the screening accuracy and universality are remarkably improved. According to the method, the speech enhancement model based on the E-CDNN-SRU is trained to suppress background noise, the collected speech signals are preprocessed, environmental noise interference is eliminated, and it is ensured that stable recognition performance can still be kept in a non-ideal acoustic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and relates to related technologies in cross fields such as intelligent medicine, artificial intelligence, robotics, instrument science, control science, computer science, sensor technology, and human-computer interaction technology. In particular, it relates to a method and system for early identification of stroke based on the fusion of sustained vowels and continuous speech. Background Art

[0002] Stroke, also known as a stroke, is characterized by rapid onset, high mortality, and high disability rate. Strokes are generally divided into two major categories: ischemic stroke and hemorrhagic stroke, and ischemic stroke is the most common type. The treatment time window for ischemic stroke is extremely critical. Intravenous thrombolysis within 4.5 hours after onset or endovascular thrombectomy within 6 hours can significantly improve the prognosis. Therefore, quickly and accurately predicting stroke is of great clinical significance for timely initiating treatment, reducing neurological damage, and improving the prognosis of patients. Pre-hospital stroke screening tools, such as the Cincinnati Prehospital Stroke Scale and the Face-Arm-Speech Test, help the public and emergency responders quickly identify strokes by observing typical symptoms such as facial asymmetry, limb weakness, and speech disorders. However, although these scales have been widely used in the early screening of emergency strokes, their use depends on certain medical knowledge, the screening results are subjective, and subtle symptoms are easily overlooked.

[0003] Speech, as an important aspect of stroke assessment, has been applied in some identification methods. The Chinese patent with the application number 201811571779.7 realizes stroke risk prediction by extracting MFCC features of specific speech segments and combining CNN deep features with a logistic regression model; the Chinese patent with the application number 201910697111.5 constructs a hybrid network architecture integrating ResNet and LSTM, and also uses MFCC features as input to realize the risk prediction of stroke dysarthria; the Chinese patent with the application number 201910695069.3 realizes the risk assessment of stroke based on the spectrogram analysis of single-syllable speech and uses a convolutional neural network; the Chinese patent with the application number 202310146221.9 generates a speech diagnosis result by combining the log-mel spectrogram of continuous speech with a VGG model. These technologies all combine the feature extraction of speech signals with deep learning models and show the ability to predict stroke risk in specific scenarios.

[0004] However, there are still common limitations in the existing technologies: First, all solutions rely on standardized speech content (such as specific words or sentences), which have certain requirements for Mandarin proficiency. It is difficult for dialect users or people with strong accents to adapt. Second, the environmental adaptability is insufficient. Only the Chinese patent with the application number 201910695069.3 clearly requires a quiet treatment room environment, and the rest of the solutions do not clearly propose noise solutions, which are vulnerable to environmental interference and affect the diagnostic accuracy in practical applications. In the scenario of early stroke recognition based on speech, patients may suddenly get sick at home, in the office, or in public places. At this time, there is noise in the surrounding environment, and the existing patent methods are difficult to accurately identify, which limits their practicality. Finally, speech recognition based on algorithms such as deep learning requires a large amount of data, while the amount of pathological speech data is small, which limits the model effects of the CNN in Patent 201811571779.7 and other solutions that require a large amount of data. The low recognition accuracy also leads to limited practicality of their technical solutions. Summary of the Invention

[0005] In view of the problems existing in the current speech-based stroke, this patent proposes an early stroke recognition method and system based on the fusion of sustained vowels and continuous speech, aiming to achieve early stroke recognition through innovative speech technologies, so as to achieve early prevention of stroke and provide an efficient and reliable technical solution for early stroke screening.

[0006] To achieve the above object, the technical solutions adopted by the present invention are as follows:

[0007] An early stroke recognition method based on the fusion of sustained vowels and continuous speech, comprising the following steps:

[0008] Step 1: Collect sustained vowel data and continuous speech data;

[0009] Step 2: Perform speech generation on the sustained vowel data and continuous speech data based on VocGAN to expand the training data set of the recognition model;

[0010] Step 3: Train a speech enhancement model based on E-CDNN-SRU to perform speech enhancement on the sustained vowel data and continuous speech data respectively;

[0011] Step 4: Recognize sustained vowels, preprocess the collected and generated sustained vowels, extract openSMILE features, and combine two demographic features of age and gender. Use the mutual information method to reduce the feature dimension, and based on a machine learning classifier, train a sustained vowel recognition model;

[0012] Step 5: Recognize continuous speech, preprocess the collected and generated continuous speech, extract Mel spectrogram features, and train a recognition model for continuous speech based on StrokeConvNet;

[0013] Step 6: Make a fusion decision based on the respective independent weights of the sustained vowel recognition model and the continuous speech recognition model to judge the stroke risk probability.

[0014] Furthermore, in the above Step 2, the establishment and training generation process of the sustained vowel and continuous speech data generation model based on VocGAN is as follows:

[0015] Select a stroke speech generation dataset, extract the Mel spectrogram as the model input, construct a VocGAN model. The generator consists of multi-scale upsampling modules, restores the time resolution layer by layer through transposed convolution, and uses residual blocks to enhance feature modeling. The discriminator is a multi-scale structure, which discriminates the generated audio and the real audio at different resolutions respectively. During the training process, the generator inputs the Mel spectrogram and outputs waveform audio, and the discriminator inputs real and generated audio and outputs discriminant results. The loss function includes joint conditional and unconditional losses, multi-resolution STFT losses, and feature matching losses, and jointly optimizes the parameters of the generator and the discriminator. After training, the model can generate speech samples that are similar but have certain differences in the time domain based on the Mel spectrograms of existing stroke speech.

[0016] Furthermore, in the above Step 3, the establishment and training process of the speech enhancement model based on E-CDNN-SRU is as follows:

[0017] Select a speech enhancement dataset, preprocess the input data, perform STFT on the noisy speech to obtain complex spectra, separate the real and imaginary part spectra as two-channel inputs, use an encoder for feature extraction and compression. The output of the encoder is passed to the decoder after temporal modeling. The decoder restores the frequency dimension layer by layer through deconvolution, and at the same time, skip connections are used to splice the low-level features output by each layer of the encoder with the high-level features of the corresponding decoder, and important regions are weighted through attention gates. Finally, the real and imaginary parts of the clean speech are output, and the output real and imaginary part spectra are combined into a complex spectrum and converted into a time-domain waveform through inverse STFT, that is, the enhanced clean speech.

[0018] Furthermore, in the above Step 4, the establishment process of the sustained vowel recognition model is as follows:

[0019] Construct a sustained vowel dataset, preprocess the sustained vowels, extract openSMILE features, and combine two demographic features of age and gender. Use the mutual information method to reduce the feature dimension to N dimensions, train multiple machine learning classifiers respectively, select the optimal classifier for each type of sustained vowel according to the recognition performance (accuracy, sensitivity, specificity, F1 value, and AUC values of ROC and PR curves) of the test set, and take the mean of all indicators to determine the weight ω of each type of sustained vowel i , and calculate the stroke risk probability through the following formula:

[0020]

[0021] Among them, when P > 0.5, it indicates that the classifier C i predicts a risk of stroke; otherwise, it is determined to be risk-free.

[0022] Furthermore, the preprocessing of the sustained vowel includes: pre-emphasis, framing and windowing, and silent segment detection. The multiple machine learning classifiers include: KNN, SVM, DT, RF, AdaBoost, and XGBoost. The recognition performance includes: accuracy, sensitivity, specificity, F1 value, and the AUC values of the ROC and PR curves.

[0023] Furthermore, in step 5, the establishment process of the recognition model based on continuous speech is as follows:

[0024] Construct a continuous speech dataset, preprocess the continuous speech, and extract Mel-spectrum features, and then input them into the StrokeConvNet network for training to obtain the optimal recognition model.

[0025] Furthermore, two demographic features of age and gender are also incorporated into the fully connected layer during the training process of the StrokeConvNet network.

[0026] Furthermore, in step 6, the weight ω of the recognition model based on the sustained vowel vowel is the mean value of the classifier weights of each category of sustained vowel. The weight ω of the recognition model based on continuous speech speech is the mean value of the comprehensive performance indicators of the continuous speech on the test set. For fusion decision-making, for each sample x, the final stroke risk probability is calculated as follows:

[0027]

[0028] Among them, when the weighted probability P > 1, it is determined to have a stroke risk; otherwise, there is no disease risk.

[0029] Furthermore, the comprehensive performance indicators include accuracy, sensitivity, specificity, F1 value, and the AUC values of the ROC curve and PR curve.

[0030] The present invention also provides a stroke early recognition system based on the fusion of sustained vowel and continuous speech, including a client and a cloud server. The client is used for user interaction, and the cloud server is used to obtain the recognition result by using the stroke early recognition method based on the fusion of sustained vowel and continuous speech and return it to the client.

[0031] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory. It is characterized in that the processor executes the computer program to implement the steps of the early stroke recognition method based on the fusion of sustained vowels and continuous speech.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] (1) The present invention integrates the bimodal analysis of sustained vowels and continuous speech. Among them, sustained vowels are suitable for a wide range of people because of their simple pronunciation, especially friendly to dialect users or people with strong accents, effectively overcoming the recognition obstacles brought by dialect and accent differences, and significantly improving the accuracy and universality of screening; while continuous speech is used to supplement the recognition of sustained vowels, providing supplementary evidence for diagnosis through richer acoustic features and language information. The combination of the two forms a complementary recognition system.

[0034] (2) Aiming at the environmental interference problem in the actual application scenario, the present invention innovatively introduces a noise suppression module, trains a voice enhancement model based on E-CDNN-SRU to suppress background noise, preprocesses the collected voice signals, eliminates environmental noise interference, and then integrates the bimodal recognition results through a weight fusion algorithm, and finally outputs stroke risk assessment suggestions to ensure stable recognition performance in non-ideal acoustic environments.

[0035] (3) Aiming at the problem that in the early stroke recognition, the pathological voice data of patients is scarce, resulting in limited accuracy of the current deep learning-based recognition methods, the present invention innovatively introduces a sustained vowel and continuous speech generation module based on VocGAN, and jointly trains the generated sustained vowels and continuous speech with the collected sustained vowels and continuous speech, thereby improving the recognition accuracy and generalization ability of the model, and further enhancing the innovation and practicality of the technical solution of the present invention.

[0036] (4) Through the collaborative design of cloud intelligent computing and lightweight clients, the present invention constructs an intelligent recognition system, which not only meets the requirements of hospital clinical diagnosis for professionalism, but also realizes the self-screening function in community and home scenarios. This technical solution not only expands the adaptability of the screening scenario, but also reduces the performance requirements for terminal devices through the centralized deployment of computing resources, which is conducive to large-scale promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a composition diagram of the early stroke recognition system based on the fusion of sustained vowels and continuous speech provided by the present invention.

[0038] Figure 2 is a technical solution diagram of the early stroke recognition based on the fusion of sustained vowels and continuous speech provided by the present invention.

[0039] Figure 3 is the early stroke recognition flowchart provided by the present invention based on the fusion of sustained vowels and continuous speech.

[0040] Figure 4 is the network architecture diagram of the speech generation model based on VocGAN provided by the present invention.

[0041] Figure 5 is the network architecture diagram of the speech enhancement E-CDNN-SRU provided by the present invention.

[0042] Figure 6 is the network architecture diagram of the continuous speech recognition StrokeConvNet provided by the present invention. Detailed implementation manners

[0043] The following will detail the technical solutions provided by the present invention in combination with specific embodiments. It should be understood that the following embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0044] Figure 1 Shown is the schematic diagram of the early stroke recognition system architecture provided by the present invention based on the fusion of sustained vowels and continuous speech. The system consists of a front-end signal acquisition device and a back-end intelligent analysis platform. In the front-end acquisition link, microphone 1-1 is responsible for collecting the speech signals of patients. In the system architecture design, a client-server (C / S) architecture based on the TCP / IP protocol is adopted to implement calculations. Among them, the cloud server is responsible for core model inference and algorithm processing, while the client realizes user access through a cross-platform interaction interface. Considering the requirements of different application scenarios, the system supports multi-terminal access methods, including but not limited to mobile terminals such as smartphones 1-2 and tablet computers 1-5, as well as PC devices such as laptop computers 1-3 and desktop computers 1-4. These devices establish a secure connection with the cloud through the standardized HTTPS protocol. The PC side can develop professional diagnostic software based on frameworks such as Qt, and the mobile side provides a convenient screening entrance through a small program, and is equipped with a signal acquisition module composed of a microphone to ensure the reliable acquisition of speech signals. The core processing unit of the system is deployed on the cloud server 1-6. The server integrates multi-level speech analysis algorithms and is responsible for core model inference and algorithm processing to ensure that the system can provide stable and accurate early stroke recognition capabilities on different devices. By integrating the complementary advantages of the two speech modalities, the system can improve the accuracy and applicability of early stroke recognition. The system realizes real-time data interaction through an optimized two-way communication mechanism. After the client initiates a request, the server can complete the calculation and return the diagnosis result by using the following early stroke recognition method based on the fusion of sustained vowels and continuous speech, which is applicable to the application requirements of multi-scenarios such as medical scenarios and home self-testing.

[0045] To implement the early stroke recognition method based on the fusion of sustained vowels and continuous speech provided by the present invention, first, an identification model based on sustained vowels, an identification model based on continuous speech, a speech generation method, a fusion decision method, and a noise suppression module should be established. The technical route is as Figure 2 shown.

[0046] Specifically, the process of establishing the identification model based on sustained vowels includes the following steps:

[0047] S1: Construct a sustained vowel dataset. Collect the sustained vowel pronunciation data of stroke patients and healthy controls through a microphone. The subjects are required to clearly read the single vowels / a / , / o / , / e / , / i / , / u / , and / ü / in Chinese pinyin. Each vowel pronunciation lasts about 1 second, and the speech signal is mono-recorded at a sampling rate of 16 kHz.

[0048] S2: Preprocess the sustained vowels. Perform pre-emphasis, frame windowing, and silent segment detection on the speech data in sequence. Pre-emphasis enhances the high-frequency components through a first-order high-pass filter. Frame windowing uses a Hamming window with a frame length of 25 ms and a frame shift of 10 ms. Silent segment detection removes invalid speech segments, and finally generates a preprocessed sustained vowel dataset.

[0049] S3: Extract high-dimensional features and fuse demographic features. Use the openSMILE tool to extract the acoustic features of the speech signal, and obtain features such as prosody, spectrum, and time-frequency domain based on the ComParE2016 standard feature set. At the same time, integrate the age and gender of the subjects as additional demographic features to form a feature vector containing 6375 dimensions, so as to comprehensively characterize the correlation between speech characteristics and individual physiological differences.

[0050] S4: Optimize feature dimensionality reduction. Use the mutual information method to rank the importance of the 6375-dimensional features, and select the 100-dimensional (the number of dimensions can be specified according to needs) features with the highest correlation with stroke recognition, eliminate redundant information, and alleviate the "curse of dimensionality" problem caused by high-dimensional features, improving the model training efficiency and generalization ability.

[0051] S5: Joint training of multiple classifiers. Construct six types of classifiers based on the K-nearest neighbor algorithm (KNN), support vector machine (SVM), decision tree (DT), random forest (RF), adaptive boosting algorithm (AdaBoost), and extreme gradient boosting algorithm (XGBoost). Use the grid search method to globally optimize the hyperparameters of each type of classifier, and randomly divide the dataset into five equal parts through a five-fold cross-validation strategy, where four parts are used as the training set and one part is used as the test set.

[0052] S6: Optimal classifier selection. Evaluate the accuracy, sensitivity, specificity, F1 value, AUC of the ROC curve, and AUC of the PR curve of each classifier on the test set, and select the optimal classifier for each type of vowel according to the comprehensive performance. By comparing the recognition effects of different vowel-classifier combinations, construct an optimal classifier mapping table for specific vowels.

[0053] S7: Weighted voting fusion decision of multiple classifiers. Calculate the weight ω based on the accuracy, sensitivity, specificity, F1 value, and mean AUC of each optimal classifier on the test set i , and use the weighted voting mechanism to fuse the prediction results of six types of classifiers. For each sustained vowel, there are six classifiers (denoted as C1, C2, …, C6), and each classifier will output a probability value belonging to a certain category. This patent aims to achieve binary classification of speech. For the sample x to be classified, its stroke risk probability is calculated by the following formula:

[0054]

[0055] where, when P > 0.5, it means that classifier C i predicts a stroke risk; otherwise, it is determined to be risk-free.

[0056] The process of establishing a recognition model based on continuous speech includes the following steps:

[0057] S8: Construction of continuous speech dataset. Collect the continuous speech data of stroke patients and healthy control groups through a microphone. Require the subjects to clearly read the standardized phrase "People's Republic of China" at a natural speed, and the speech signals are recorded in mono with a sampling rate of 16 kHz.

[0058] S9: Preprocessing of continuous speech signals. Perform pre-emphasis, frame windowing (25 ms Hamming window, 10 ms frame shift), and silent segment removal operations on the speech data in sequence to obtain the preprocessed continuous speech dataset.

[0059] S10: Standardized extraction of continuous speech features. Regularize the duration of the preprocessed speech, and unify each speech to a length of 2.5 seconds (corresponding to 40,000 sampling points) by padding zeros at the end. Use the short-time Fourier transform (STFT) to extract 64-channel Mel spectrum features, with a window length of 25 ms and a frame shift of 10 ms. After logarithmic compression and normalization, generate a time-frequency feature matrix with a dimension of 64×251 (frequency × time).

[0060] S11: Continuous speech data augmentation. At the time domain level, execute noise addition (NA), pitch shifting (PS), and time stretching (TS) strategies. At the time-frequency domain level, adopt the SpecAugment method to randomly mask frequency bands and time segments of the Mel spectrogram. The augmented data needs to re-execute the preprocessing and feature extraction processes of S9 - S10 to expand the scale of the training set.

[0061] S12: Train a continuous speech recognition model based on the StrokeConvNet network, using the extracted Mel spectrogram features as input, with an input dimension of B×1×64×251 (batch × channel × frequency × time). The architecture of the StrokeConvNet network is as Figure 6 shown and is divided into three parts:

[0062] Feature extraction module, which contains three parallel convolutional paths.

[0063] S13: Path 1 is used to extract spectral features. It uses a 9×1 convolutional kernel with a stride of 1 and a padding size of 4 in the frequency dimension and 0 in the time dimension. The output dimension remains B×32×64×251. This path retains the original frequency and time dimensions and extracts long-range dependencies in the spectral direction.

[0064] S14: Path 2 is used to extract time-domain features. It uses a 1×11 convolutional kernel with a stride of 1 and a padding of 5 in the time dimension and 0 in the frequency dimension. The output dimension remains B×32×64×251. This path captures local dynamic patterns in the time direction.

[0065] S15: Path 3 extracts joint time-frequency domain features. It uses a 3×3 convolutional kernel with a stride of 1 and a padding of 1 in both the frequency and time dimensions. The output dimension is B×32×64×251. This path jointly analyzes the short-range correlations in the time-frequency domain.

[0066] S16: Fuse the path features, merge the outputs of the three paths along the channel axis to generate fused features of B×96×64×251.

[0067] Feature learning module: This module further abstracts features through deep convolution and contains five consecutive CBRA blocks (convolution → batch normalization → ReLU → average pooling).

[0068] S17: CBRA block 1 uses a 3×3 convolutional kernel (padding = 1, stride = 1), followed by 2×2 average pooling (stride = 2, padding = 0), and the feature dimension is compressed to B×64×32×125.

[0069] S18: The CBRA block 2 uses a 3×3 convolutional kernel (padding = 1, stride = 1), followed by a 2×2 average pooling (stride = 2, padding = 0), and the feature dimension becomes B×128×16×62.

[0070] S19: The CBRA block 3 uses a 3×3 convolutional kernel (padding = 1), but the pooling layer is adjusted to 2×1 (stride = 2, padding = 0), the time dimension is halved, and the output dimension is B×256×16×31.

[0071] S20: The CBRA block 4 uses a 3×3 convolutional kernel (padding = 1), the pooling layer is 2×1 (stride = 2, padding = 0), the time dimension is halved, and the output dimension becomes B×256×16×15.

[0072] S21: The CBRA block 5 uses a 1×1 convolutional kernel (padding = 0), followed by a global average pooling layer, compresses the time-frequency dimension to 1×1, and the output dimension becomes B×256.

[0073] S22: Fuse the two demographic features of age and gender to generate a fused feature of B×258 dimensions.

[0074] Classification module: Implement the classification of health and stroke risk.

[0075] S23: Input the fused features into a fully connected layer (64 units), the output dimension is B×64, after Dropout, it is mapped to a binary classification probability (health / stroke risk) through a Softmax layer.

[0076] S24: The training adopts five-fold cross-validation, groups according to speakers, and introduces an early stopping mechanism to optimize the model according to the loss value of the test set. Calculate the accuracy, sensitivity, specificity, F1 value, ROC-AUC and PR-AUC metrics of the model on the test set, evaluate the performance of the model, and select the model with the best comprehensive metrics as the continuous speech recognition model.

[0077] In particular, considering the limitations in quantity and quality of the self-built stroke speech dataset, before training the specific recognition model, the present invention separately trains a speech generation model for sustained vowels and continuous speech based on VocGAN, and the specific steps are as follows:

[0078] S25: Speech preprocessing and feature extraction: Uniform the duration of the preprocessed sustained vowel speech in step S2 to 1 second, and uniform the duration of the preprocessed continuous speech in step S9 to 2.5 seconds. Subsequently, extract the Mel spectrogram as the input feature of the generation model.

[0079] S26: Construct a speech generation network based on VocGAN, and the network framework is as Figure 4As shown. A voice generation model is constructed. The generator is composed of stacked multi-scale upsampling modules. Each module restores the audio time resolution layer by layer through transposed convolution, and residual blocks are embedded in each layer to enhance the feature modeling ability. The discriminator adopts a multi-scale structure to discriminate between the generated audio and the real audio at different time resolutions, so as to guide the generator to generate more natural speech waveforms.

[0080] S27: Training of the generation model. Using the Mel spectrogram as the input of the generator and the corresponding real audio waveform as the training target, the generator and the discriminator are trained adversarially. During the training process, the generator outputs a speech waveform, and the discriminator simultaneously receives the real waveform and the generated waveform and classifies them as true or false. The training process uses three loss functions, namely the joint conditional and unconditional adversarial loss, the multi-resolution STFT loss, and the feature matching loss, to jointly optimize the parameters of the generator and the discriminator.

[0081] S28: Generation and augmentation of speech data. Stroke speech generation models based on sustained vowels and continuous speech are trained respectively. In the inference stage, by inputting the Mel spectrogram features of any segment of sustained vowels or continuous speech, speech samples with a specified number and similar to the original audio in the time domain but with certain random differences can be generated. Finally, the collected original speech data and the generated speech data are jointly constructed into the training set of the recognition model to achieve the expansion of the sample scale.

[0082] To achieve more reliable recognition, the fusion decision of sustained vowels and continuous speech is continued. The specific steps are as follows:

[0083] S29: Fusion decision based on the independent weights of the sustained vowel and continuous speech models. The weight ω of the sustained vowel vowel takes the mean value of the six-classifier metrics, and the weight ω of the continuous speech speech takes the average values of the accuracy, sensitivity, specificity, F1 value, AUC of the ROC curve, and AUC of the PR curve of the continuous speech model on the test set. For each sample x, the final stroke risk probability is calculated as follows:

[0084]

[0085] Among them, when the weighted probability P > 1, it is determined that there is a stroke risk; otherwise, there is no disease risk.

[0086] The noise suppression module is based on the E-CDNN-SRU (Extended Compact Deep Neural Network with SRU) network. Single-channel speech enhancement models are trained specifically for the two speech types of sustained vowels and continuous speech respectively. Its network framework is as Figure 5 shown, providing a clearer speech input for subsequent recognition tasks. The specific steps are as follows:

[0087] S30: Selection of the speech enhancement dataset. For the training data, open-source noise datasets and clean speech datasets are used. The noise data includes the DEMAND dataset and the Hospital Ambient Noise Dataset. The clean sustained vowel data is selected from the SVD, FEMH, and VOICED datasets, and the clean continuous speech data is selected from the Chinese open-source dataset THCHS-30. The test set uses the self-built stroke speech data. All audio data is uniformly resampled to a sampling rate of 16 kHz. Among them, the sustained vowel samples are cropped to a length of 1 second, the continuous speech samples are cropped to a length of 3 seconds, and the noise samples are correspondingly cropped according to the length of the speech samples to be superimposed to ensure that the duration of the noise matches that of the speech samples. Realistic noisy speech is simulated by superimposing noise at a certain signal-to-noise ratio.

[0088] S31: Preprocessing of the input data. The noisy speech is subjected to STFT to obtain the complex spectrum. |Y| is the amplitude spectrum corresponding to the frequency and time, and θ Y is the phase spectrum. The real and imaginary part spectra are separated as the two-channel input. Among them, the window uses the Hamming window, with a length of 20 ms (320 sampling points) and a frame shift of 10 ms (50% overlap); the number of FFT points is 320, and the frequency dimension is 161. Thus, the input shape of the sustained vowel is [Batch Size, 2, 100, 161], and the input shape of the continuous speech is [Batch Size, 2, 300, 161]. For the convenience of subsequent unified deduction, denote: the feature dimension is B×2×T×161 (B represents the batch, and T is the time dimension).

[0089] S32: Encoder feature extraction and compression. The input passes through five layers of causal convolution, with a convolution kernel of 1×3 and a stride of (1, 2). Each layer gradually compresses the frequency dimension (161→80→39→19→9→4) to extract high-frequency abstract features. At the same time, after each layer of convolution, an ELU activation function and batch normalization are connected to accelerate training and stabilize the gradient. The compressed feature dimension becomes B×128×T×4.

[0090] S33: Temporal modeling, that is, the SRU bottleneck layer. The output of the encoder is input to 2 stacked SRUs (512 units per layer, with a hidden state dimension of 512) after Reshape to model the time context (11-frame window: 10-frame historical frames + 1-frame current frame, covering 110 ms of context). The SRU efficiently captures long-term dependencies through a simplified gating mechanism, and the feature dimension is restored after Reshape, and the output is passed to the decoder. The feature dimension changes in sequence as B×128×T×4→B×T×512→B×128×T×4.

[0091] S34: Decoder spectrum reconstruction and attention fusion, restoring the frequency dimension layer by layer through deconvolution (4→9→19→39→80→161). Meanwhile, skip connections splice the low-level features output by each layer of the encoder with the high-level features of the corresponding decoder (along the channel dimension), and weight the important regions through attention gates, finally outputting the real part of the clean speech. And the imaginary part With the shape of B×2×T×161.

[0092] S35: Speech waveform reconstruction and post-processing, combining the output real and imaginary part spectra into a complex spectrum Converting it into a time-domain waveform through inverse STFT That is, the enhanced clean speech.

[0093] To enhance the adaptability of the model to different signal-to-noise ratio environments, the signal-to-noise ratios of the training data are set to 0dB, 5dB, 10dB, and 15dB, the signal-to-noise ratios of the validation set are set to 2.5dB, 7.5dB, 12.5dB, and 17.5dB, while the signal-to-noise ratios of the test set are uniformly sampled and distributed within the range of [0dB, 20dB] to ensure that the test set contains a wider range of noise conditions, so as to more comprehensively evaluate the actual noise reduction ability of the model.

[0094] In terms of dataset division, to ensure the independence between the training set, validation set, and test set, when dividing the clean speech samples, follow the principle of "the same speaker only appears in one set" to avoid data leakage problems and ensure the generalization ability of the model. For the division of noise samples, it is required that noise samples of the same category are evenly distributed in each dataset, thereby reducing the impact of data bias on the evaluation of model performance. Through the above strategies, the model can not only learn the deep features of speech, but also maintain a stable noise reduction effect in complex noise environments, improving the overall performance of speech enhancement.

[0095] Based on the above technologies, the early stroke recognition process based on the fusion of sustained vowels and continuous speech provided by the present invention (as Figure 3 shown) includes the following steps:

[0096] Step 1: Collect sustained vowel data and continuous speech data. Before this step, user information (ID, gender, age, etc.) also needs to be input, and the speech task can be selected to collect different speech data.

[0097] Step 2: First, perform speech enhancement on the sustained vowel data and continuous speech data respectively based on the trained speech enhancement model of E-CDNN-SRU. In this step, the speech enhancement model of E-CDNN-SRU is established and trained through the aforementioned noise suppression modules S30-S35 steps.

[0098] Step 3: Identify sustained vowels, preprocess the sustained vowels, extract openSMILE features, and combine two demographic features of age and gender. Use the mutual information method to reduce the feature dimension to 100 dimensions, and use the recognition model based on sustained vowels to obtain the prediction result. In this step, the recognition model based on sustained vowels is established by using the aforementioned S1 - S6 steps.

[0099] Step 5: Identify continuous speech, preprocess the continuous speech, and extract Mel - frequency cepstral features. Use the recognition model based on continuous speech to achieve the classification of health and stroke risk. In this step, the recognition model based on continuous speech is established by using the aforementioned S8 - S24 steps.

[0100] Step 6: Make a fusion decision for sustained vowels and continuous speech. Determine the overall weight of sustained vowels based on the mean value of the classifier weights for each category of sustained vowels, determine the weight of continuous speech based on the mean value of the comprehensive performance indicators of continuous speech on the test set, calculate the final fusion decision weight, and judge the stroke risk probability of each sample. Through this fusion strategy, the system fully utilizes the feature complementarity of sustained vowels and continuous speech at different levels, improving the accuracy and stability of early stroke recognition. In this step, the specific process of the fusion decision is implemented by using the aforementioned S7 and S29 steps.

[0101] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. An early stroke recognition method based on the fusion of sustained vowels and continuous speech, characterized in that, It includes the following steps: Step 1: Collect sustained vowel data and continuous speech data; Step 2: Perform speech generation based on the sustained vowel data and continuous speech data of VocGAN to expand the training dataset of the recognition model; Step 3: Train a speech enhancement model based on E-CDNN-SRU to perform speech enhancement on the sustained vowel data and continuous speech data respectively; Step 4: Recognize sustained vowels, preprocess the collected and generated sustained vowels, extract openSMILE features, and combine two demographic features of age and gender. Use the mutual information method to reduce the feature dimension. Based on a machine learning classifier, train a recognition model for sustained vowels; Step 5: Recognize continuous speech, preprocess the collected and generated continuous speech, and extract Mel spectrogram features. Train a recognition model for continuous speech based on StrokeConvNet; Step 6: Make a fusion decision according to the respective independent weights of the recognition model of sustained vowels and the recognition model of continuous speech to judge the stroke risk probability.

2. The early stroke recognition method based on the fusion of sustained vowels and continuous speech according to claim 1, wherein In the said Step 2, the establishment and training process of the generation model of sustained vowels and continuous speech data based on VocGAN is as follows: Select a stroke speech generation dataset, extract Mel spectrogram as the model input, construct a VocGAN model. The generator consists of a multi-scale upsampling module, restores the time resolution layer by layer through transposed convolution, and uses residual blocks to enhance feature modeling. The discriminator is a multi-scale structure, and discriminates the generated audio and the real audio at different resolutions respectively; During the training process, the generator inputs the Mel spectrogram and outputs waveform audio, and the discriminator inputs the real and generated audio and outputs the discrimination result; The loss function includes joint conditional and unconditional loss, multi-resolution STFT loss and feature matching loss, and jointly optimizes the parameters of the generator and the discriminator; after training, the model can generate speech samples that are similar but have certain differences in the time domain based on the Mel spectrogram of existing stroke speech.

3. The early stroke recognition method based on the fusion of sustained vowels and continuous speech according to claim 1, wherein In the said Step 3, the establishment and training process of the speech enhancement model based on E-CDNN-SRU is as follows: Select a speech enhancement dataset, preprocess the input data, perform STFT on the noisy speech to obtain a complex spectrum, separate the real and imaginary parts of the spectrum as dual-channel inputs, use an encoder for feature extraction and compression, and the output of the encoder is passed to the decoder after temporal modeling. The decoder restores the frequency dimension layer by layer through deconvolution, and at the same time, the skip connection splices the low-level features output by each layer of the encoder with the high-level features of the corresponding decoder, weights the important regions through an attention gate, and finally outputs the real and imaginary parts of the clean speech. Combine the output real and imaginary parts of the spectrum into a complex spectrum and convert it into a time-domain waveform through inverse STFT, that is, the enhanced clean speech.

4. The early stroke recognition method based on the fusion of sustained vowels and continuous speech according to claim 1, wherein In the said Step 4, the establishment process of the recognition model based on sustained vowels is as follows: Construct a sustained vowel dataset, preprocess the sustained vowels, extract openSMILE features, and combine two demographic features of age and gender; subsequently, use the mutual information method to reduce the feature dimension to N dimensions, train multiple machine learning classifiers respectively, select the optimal classifier for each type of sustained vowel according to the recognition performance of the test set, and take the mean of all metrics to determine the weight ω of each type of sustained vowel i , and calculate the stroke risk probability through the following formula: Among them, when P > 0.5, it indicates that classifier C i predicts a risk of stroke; otherwise, it is determined to be risk-free.

5. The early stroke recognition method based on the fusion of sustained vowels and continuous speech according to claim 4, characterized in that, The preprocessing of continuous vowels includes: pre-emphasis, framing and windowing, and silent segment detection. The multiple machine learning classifiers include: KNN, SVM, DT, RF, AdaBoost, and XGBoost. The recognition performance includes: accuracy, sensitivity, specificity, F1 value, and the AUC values of the ROC and PR curves.

6. The early stroke recognition method based on the fusion of sustained vowels and continuous speech according to claim 1, characterized in that In step 5, the establishment process of the recognition model based on continuous speech is as follows: Construct a continuous speech dataset, preprocess the continuous speech, extract Mel spectrum features, and then input them into the StrokeConvNet network for training to obtain the optimal recognition model.

7. The early stroke recognition method based on the fusion of sustained vowels and continuous speech according to claim 6, characterized in that, During the training process of the StrokeConvNet network, two demographic features of age and gender are also incorporated into the fully connected layer.

8. The early stroke recognition method based on the fusion of sustained vowels and continuous speech according to claim 1, wherein, In step 6, the weight ω of the recognition model based on sustained vowels vowel is the mean of the classifier weights of various categories of sustained vowels, and the weight ω of the recognition model based on continuous speech speech is the mean of the comprehensive performance indicators of continuous speech on the test set. For fusion decision-making, for each sample x, the final stroke risk probability is calculated as follows: Among them, when the weighted probability P > 1, it is determined that there is a risk of stroke; otherwise, there is no risk of disease.

9. The early stroke recognition method based on the fusion of sustained vowels and continuous speech according to claim 8, characterized in that, The comprehensive performance indicators include accuracy, sensitivity, specificity, F1 value, and the AUC values of the ROC curve and PR curve.

10. A stroke early recognition system based on the fusion of sustained vowels and continuous speech, including a client and a cloud server, where the client is used for user interaction, and is characterized in that, The cloud server is used to obtain the recognition result by using the early stroke recognition method based on the fusion of continuous vowels and continuous speech described in any one of claims 1-9 and return it to the client.

Citation Information

Patent Citations

  • Cerebral apoplexy risk prediction method based on deep voice features

    CN109559761A

  • Stroke risk assessment device and equipment

    CN110415824B

  • Cerebral stroke dysarthria risk prediction method based on ResNet and LSTM network

    CN110600053A

  • Cerebral stroke early screening method combining comparative learning and multi-modal fusion

    CN118538394A