Multi-scale grouping feature fusion AD screening method based on voice features

Through a multi-scale grouping feature fusion method based on speech features, the high cost and invasiveness problems of early screening for Alzheimer's disease are solved, and high-sensitivity, low-cost diagnostic accuracy and efficiency are achieved.

CN120708661APending Publication Date: 2025-09-26GUANGDONG MEDICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510845649.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies have problems with high cost, invasive procedures and strong subjectivity in early screening of Alzheimer's disease. Traditional feature extraction methods are difficult to effectively process high-dimensional nonlinear data, resulting in insufficient ability to identify early biomarkers.

Method used

A multi-scale grouping feature fusion method based on speech features is adopted. The speech signal is processed by a frequency reduction algorithm, the Mel-frequency cepstral coefficients are extracted, and feature optimization is performed using residual blocks and attention mechanisms. A multi-scale feature pyramid is constructed to generate globally optimized features, which are then input into a four-classification neural network for diagnosis.

Benefits of technology

It has achieved high-sensitivity, low-cost early screening for Alzheimer's disease, improved diagnostic accuracy and efficiency, and reduced dependence on professional equipment and personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708661A_ABST
    Figure CN120708661A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-scale grouping feature fusion AD screening method based on voice features, and the method comprises the steps: obtaining an original voice signal from a subject, and carrying out the preprocessing of the original voice signal, so as to generate standardized voice data; frequency domain features are extracted according to the standardized voice data, and a multi-dimensional feature mapping matrix is constructed; performing grouping processing on the multi-dimensional feature mapping matrix and generating deep speech feature representation through a residual connection mechanism; calculating a feature difference weight by adopting an attention mechanism and carrying out feature fusion to generate a local optimization feature vector and a global optimization feature; splicing the local optimization feature vector and the global optimization feature, and processing through an activation function to obtain a joint feature vector; and processing the joint feature vector through a pre-trained classification model, and outputting a classification result of the cognitive state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information technology, and in particular to an AD screening method based on multi-scale grouping feature fusion of speech features. Background Art

[0002] Currently, screening for Alzheimer's disease (AD) and its early stages (such as subjective cognitive decline (SCD) and mild cognitive impairment (MCI)) mainly relies on clinical scale assessments, neuroimaging examinations (such as MRI and PET), and biomarker testing. Although these methods play an important role in diagnosis, they also have some limitations. For example, clinical scale assessments are highly subjective and easily influenced by patients' emotions and environmental factors; neuroimaging examinations are expensive, complex, and difficult to implement large-scale screening; biomarker testing requires invasive procedures and has low patient acceptance.

[0003] The main reasons for the above problems are: due to technical limitations, traditional feature extraction methods cannot effectively process high-dimensional nonlinear data such as speech and behavior, resulting in insufficient recognition of early biomarkers; due to cost and resource constraints, high-precision detection technology relies on specialized equipment and professionals, with poor economic and accessibility; subjective interference: lack of objective and quantitative physiological indicators, diagnostic results are easily affected by human factors;

[0004] There is an urgent need to develop an early screening method that is highly sensitive, low-cost, and non-invasive. In recent years, speech signal analysis has become a potential breakthrough due to its advantages of being non-invasive, easy to obtain, and containing rich cognitive features. However, how to stably extract multi-scale features related to AD from speech and achieve high-precision classification remains a technical difficulty to be solved. Summary of the Invention

[0005] In order to solve the problems existing in the above-mentioned prior art, the purpose of this application is to provide an AD screening method based on multi-scale grouping feature fusion of speech features.

[0006] The multi-scale grouping feature fusion AD screening method based on speech features described in this application includes:

[0007] Step S101: After collecting the original speech signal of the subject, the audio sampling rate is uniformly converted through a frequency reduction algorithm, and continuous speech segments are automatically intercepted according to the speech quality evaluation index. If the background noise intensity exceeds a preset threshold, the segment is filtered to obtain standardized speech data;

[0008] Step S102: extracting Mel-frequency cepstral coefficients from the standardized speech data using an Fbank filter bank, generating a multidimensional frequency domain feature mapping matrix through time-frequency transformation, and obtaining energy distribution characteristics of the speech signal at different frequency bandwidths;

[0009] Step S103: Grouping is performed according to the channel dimension of the multi-dimensional frequency domain feature mapping matrix. Each group of feature data is input into the corresponding residual block for nonlinear transformation. The feature gradient propagation is maintained through the residual connection to obtain a preliminary deep speech feature representation. The residual block includes a neural network unit with three convolutional layers and skip connections.

[0010] Step S104: The Block_diff_AFF module is used to calculate the difference attention weights between the output features of adjacent residual blocks. If the difference between adjacent feature blocks is higher than a set value, the corresponding weight coefficient is increased, and a local optimized feature vector is generated through a weighted fusion mechanism. The Block_diff_AFF module is an attention mechanism unit that calculates feature difference based on Euclidean distance.

[0011] Step S105: Perform a two-dimensional convolution operation on the local optimized feature vector and perform downsampling processing, gradually increasing the number of feature channels to construct a multi-scale feature pyramid, and obtaining abstract speech feature representations at different levels. The multi-scale feature pyramid is a hierarchical structure composed of convolution feature maps with decreasing resolution.

[0012] Step S106, calculating the correlation weight distribution between cross-level features through the AFF attention fusion module, adaptively modulating the multi-scale features according to the weight distribution, and generating a global optimized feature containing global semantic information, wherein the AFF attention fusion module: a fusion unit that calculates the correlation of cross-level features based on cosine similarity;

[0013] Step S107: Concatenate the local optimized feature vector and the global optimized feature according to the feature dimension. The concatenated high-dimensional features are nonlinearly mapped using the SiLU activation function to obtain a joint feature vector that integrates local details and global semantics.

[0014] Step S108: Use the AdamW optimizer to train a four-classification neural network model, use the joint feature vector as the network input, and map it to the probability distribution of the four categories through the fully connected layer, and output the classification results of normal, subjective cognitive decline, mild cognitive impairment, and Alzheimer's disease.

[0015] Preferably, in step S101, obtaining an original speech signal from a subject and preprocessing the original speech signal to generate standardized speech data includes:

[0016] Collecting the original speech signal through a multi-channel device and calculating the signal spectrum using a spectrum analysis method;

[0017] Adjusting the sampling rate to a preset standard according to the signal spectrum to generate adjusted spectrum data;

[0018] extracting continuous speech segments from the adjusted spectral data using a voice activity detection algorithm;

[0019] Calculating the background noise intensity for the continuous speech segment, and if the background noise intensity exceeds a preset threshold, purifying the noise by a filtering method to obtain purified speech data;

[0020] Extracting speech features from the purified speech data and performing quality assessment, and if the quality assessment score is lower than a preset threshold, removing the corresponding segment to generate the standardized speech data;

[0021] The standardized speech data is normalized and the feature amplitude is adjusted to obtain final preprocessed data.

[0022] Preferably, in step S102, extracting frequency domain features based on the standardized speech data and constructing a multi-dimensional feature mapping matrix includes:

[0023] Processing the standardized speech data using a filter bank to extract frequency features to obtain an initial feature set;

[0024] Processing the initial feature set through time-frequency transformation to generate the multidimensional feature mapping matrix;

[0025] If the dimension of the multidimensional feature mapping matrix exceeds a preset threshold, a dimensionality reduction algorithm is used to obtain a reduced-dimensional feature matrix;

[0026] Calculate the energy distribution characteristics of each frequency bandwidth according to the feature matrix after dimension reduction to generate an energy distribution vector;

[0027] If the variance of the energy distribution vector is lower than a preset threshold, processing the reduced-dimensional feature matrix by an edge enhancement method to obtain an enhanced feature matrix;

[0028] The enhanced feature matrix is ​​classified to generate frequency bandwidth feature clusters and determine the final feature distribution.

[0029] Preferably, in step S103, grouping the multidimensional feature mapping matrix and generating a deep speech feature representation through a residual connection mechanism includes:

[0030] Grouping is performed according to the channel dimension of the multidimensional feature mapping matrix to obtain grouped feature data;

[0031] If the distribution of the grouped feature data does not meet the preset balance condition, the distribution is adjusted by a resampling method to obtain an adjusted feature grouping;

[0032] Performing a nonlinear transformation on the adjusted feature grouping through a residual block, maintaining feature gradients using a residual connection mechanism, and generating a transformed feature vector;

[0033] Based on the transformed feature vector, a multi-layer feature extraction method is used to determine whether the extracted features meet the preset accuracy;

[0034] If the preset accuracy is not met, the residual block parameters are adjusted through the feedback mechanism to obtain the optimized feature representation;

[0035] Extract key information from the optimized feature representation to determine the final deep speech features.

[0036] Preferably, in step S104, the step of using the attention mechanism to calculate feature difference weights and perform feature fusion to generate a local optimized feature vector includes:

[0037] Obtain output features from the residual block, and determine the difference between the output features of adjacent residual blocks through the difference calculation module;

[0038] If the difference is higher than a preset threshold, the attention weight is adjusted through a weight enhancement mechanism to obtain an enhanced weight coefficient;

[0039] Performing weighted fusion on the output features of adjacent residual blocks according to the enhanced weight coefficients to generate a preliminary local optimized feature vector;

[0040] Processing the preliminary local optimized feature vector by a deep extraction method to obtain a deep feature representation;

[0041] Extracting local features from the deep feature representation and performing weighted processing, and if the feature difference after weighted processing is lower than a preset threshold, adjusting the features through residual connection to obtain an adjusted feature vector;

[0042] The adjusted feature vector is processed by a feature integration method to generate a final local optimized feature vector.

[0043] Preferably, in step S106, the use of the attention mechanism to calculate feature difference weights and perform feature fusion to generate global optimization features includes:

[0044] Obtain the initial feature set from the multi-scale feature data, calculate the correlation weights of cross-level features through the attention fusion module, and obtain the weight distribution matrix;

[0045] Adaptively modulating the multi-scale feature data according to the weight distribution matrix to generate a modulated feature vector;

[0046] If the correlation of some features in the modulated feature vector is lower than a preset threshold, a secondary fusion process is performed on the features to obtain an enhanced feature vector;

[0047] fusing global semantic information according to the enhanced feature vector to generate semantic enhancement features;

[0048] Optimizing the semantic enhancement features by a smoothing method to obtain a final global optimization feature;

[0049] Extract key semantic information from the final global optimization features and determine feature expressions.

[0050] Preferably, in step S107, the concatenation of the local optimization feature vector and the global optimization feature and processing through an activation function to obtain a joint feature vector includes:

[0051] Extract feature information from the local optimized feature vector and the global optimized feature, and if the dimensions of the two are inconsistent, adjust them to the same dimension through a dimension alignment method to obtain a feature vector with consistent dimension;

[0052] The feature vectors with the same dimension are concatenated according to the feature dimension to generate a high-dimensional feature vector;

[0053] Performing nonlinear mapping on the high-dimensional feature vector through an activation function to obtain a preliminary fused feature vector;

[0054] Processing the preliminary fused feature vector using a normalization method to obtain a normalized feature vector;

[0055] Extracting key features from the normalized feature vector and performing dimensionality reduction processing to obtain a low-dimensional joint feature vector;

[0056] The low-dimensional joint feature vector is processed by a classification model to generate a final joint feature vector.

[0057] Preferably, in step S108, the processing of the joint feature vector by a pre-trained classification model and outputting a classification result of the cognitive state includes:

[0058] performing normalization processing on the joint feature vector to obtain normalized feature data;

[0059] Initialize the classification model parameters through the optimizer, and use the normalized feature data for training to obtain a preliminary training model;

[0060] The output of the preliminary training model is converted into a probability distribution by a mapping method, and the probability value of each classification is calculated;

[0061] If the maximum value of the probability value is lower than a preset threshold, performing enhancement processing on the joint feature vector to obtain enhanced feature data;

[0062] Retraining the classification model using the enhanced feature data, adjusting parameters through an optimizer to obtain an optimized model;

[0063] The final classification result is outputted according to the optimized model, and a corresponding category label is generated to determine the cognitive state classification.

[0064] The multi-scale grouping feature fusion AD screening method based on speech features described in the present application has the advantages that, first, the subject's speech is standardized and feature extracted to obtain a multi-dimensional frequency domain feature mapping matrix, then the features are locally optimized through residual blocks and attention mechanisms, and a multi-scale feature pyramid is constructed to obtain abstract representations at different levels, and then the attention fusion module is used to realize adaptive modulation of cross-level features to generate optimized features containing global semantic information, and finally the local and global features are fused and input into a four-classification neural network model to realize automatic diagnosis of normal, subjective cognitive decline, mild cognitive impairment and Alzheimer's disease. The present invention effectively captures cognitive impairment-related features in speech signals through multi-level feature extraction and fusion, thereby improving the diagnostic accuracy and efficiency of cognitive impairment. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is the process of the multi-scale grouping feature fusion AD screening method based on speech features described in this application Figure 1 ;

[0066] Figure 2 This is the process of the multi-scale grouping feature fusion AD screening method based on speech features described in this application Figure 2 . DETAILED DESCRIPTION

[0067] like Figure 1-Figure 2 As shown, the multi-scale grouping feature fusion AD screening method based on speech features described in this application includes:

[0068] Step S101: After collecting the original speech signal of the subject, the audio sampling rate is uniformly converted through a frequency reduction algorithm, and continuous speech segments are automatically intercepted according to the speech quality evaluation index. If the background noise intensity exceeds a preset threshold, the segment is filtered to obtain standardized speech data;

[0069] Step S102: extracting Mel-frequency cepstral coefficients from the standardized speech data using an Fbank filter bank, generating a multidimensional frequency domain feature mapping matrix through time-frequency transformation, and obtaining energy distribution characteristics of the speech signal at different frequency bandwidths;

[0070] Step S103: Grouping is performed according to the channel dimension of the multidimensional frequency domain feature mapping matrix. Each group of feature data is input into the corresponding residual block for nonlinear transformation. The feature gradient propagation is maintained through the residual connection to obtain a preliminary deep speech feature representation.

[0071] Step S104: The Block_diff_AFF module is used to calculate the difference attention weights between the output features of adjacent residual blocks. If the difference between adjacent feature blocks is higher than a set value, the corresponding weight coefficient is enhanced, and a local optimized feature vector is generated through a weighted fusion mechanism.

[0072] Step S105: performing a two-dimensional convolution operation on the locally optimized feature vector and performing downsampling processing, increasing the number of feature channels layer by layer to construct a multi-scale feature pyramid, and obtaining abstract speech feature representations at different levels;

[0073] Step S106: Calculate the correlation weight distribution between cross-level features through the AFF attention fusion module, adaptively modulate the multi-scale features according to the weight distribution, and generate a global optimized feature containing global semantic information;

[0074] Step S107: Concatenate the local optimized feature vector and the global optimized feature according to the feature dimension. The concatenated high-dimensional features are nonlinearly mapped using the SiLU activation function to obtain a joint feature vector that integrates local details and global semantics.

[0075] Step S108: Use the AdamW optimizer to train a four-classification neural network model, use the joint feature vector as the network input, and map it to the probability distribution of the four categories through the fully connected layer, and output the classification results of normal, subjective cognitive decline, mild cognitive impairment, and Alzheimer's disease.

[0076] like Figure 1-Figure 2 As shown, in step S101, after collecting the original speech signal of the subject, the audio sampling rate is uniformly converted through a down-conversion algorithm, and continuous speech segments are automatically intercepted according to the speech quality evaluation index. If the background noise intensity is detected to exceed a preset threshold, the segment is filtered to obtain standardized speech data.

[0077] Furthermore, in step S101, the original speech signal of the subject is obtained, and multi-channel speech data is collected through a microphone array to obtain initial multi-channel speech signal data;

[0078] Calculating the signal spectrum using short-time Fourier transform according to the initial multi-channel speech signal data to obtain first spectrum data;

[0079] Using the first spectrum data, a down-conversion algorithm is used to adjust the sampling rate, and a linear interpolation method is used to unify the audio sampling rate to a preset standard to obtain second spectrum data;

[0080] Extracting speech-active regions using a voice activity detection algorithm based on the second spectrum data; if the duration of a non-speech region detected exceeds a preset threshold, removing the region to obtain a continuous speech segment;

[0081] For the continuous speech segment, the power spectral density of the background noise intensity is calculated. If the power spectral density exceeds a preset threshold, the background noise is filtered out by spectral subtraction to obtain first purified speech data;

[0082] By first purifying the speech data, extracting speech features using Mel-frequency cepstral coefficients, generating feature vectors, and obtaining speech feature data;

[0083] Based on the speech feature data, the speech quality is evaluated using a Gaussian mixture model. If the evaluation score is lower than the preset threshold, the corresponding segment is removed to obtain standardized speech data;

[0084] For the standardized speech data, normalization processing is used to adjust the amplitude of the feature vector to obtain the final speech data;

[0085] The final speech data is standardized by combining it with the pre-established feature extraction module to obtain normalized feature data.

[0086] Specifically, in step S101, when obtaining the original speech signal of the subject, it is collected through an array composed of four microphones, each microphone records data at a sampling rate of 44.1kHz, and obtains initial multi-channel speech signal data;

[0087] Based on these data, a 256-point short-time Fourier transform is used to calculate the signal spectrum, with a window length of 20 milliseconds and an overlap rate of 50%, to obtain first spectrum data;

[0088] Using the first spectrum data, a downsampling algorithm is used to reduce the sampling rate from 44.1kHz to 16kHz, and a linear interpolation method is used to unify the audio sampling rate to a preset standard to obtain second spectrum data;

[0089] Extracting speech-active regions using a voice activity detection algorithm based on an energy threshold based on the second spectrum data; if a non-speech region is detected to be longer than 1 second, the region is removed to obtain a continuous speech segment;

[0090] For the continuous speech segment, the power spectral density of the background noise intensity is calculated. If the power spectral density exceeds -50dB, the background noise is filtered out by spectral subtraction to obtain the first purified speech data;

[0091] By first purifying the voice data, extracting voice features using Mel-frequency cepstral coefficients, generating a feature vector containing 13 dimensions, and obtaining voice feature data;

[0092] Based on the speech feature data, the speech quality is evaluated using a Gaussian mixture model. If the evaluation score is lower than 0.7, the corresponding segment is removed to obtain standardized speech data.

[0093] For the standardized speech data, Z-score normalization is used to adjust the feature vector amplitude so that the feature vector mean is 0 and the standard deviation is 1, and the final speech data is obtained;

[0094] The final speech data is normalized by combining it with the pre-established feature extraction module, and a normalization factor of 0.1 is used to obtain normalized feature data.

[0095] like Figure 1-Figure 2 As shown, in step S102, the Mel-frequency cepstral coefficients are extracted from the standardized speech data using the Fbank filter bank, and a multi-dimensional frequency domain feature mapping matrix is ​​generated through time-frequency transformation to obtain the energy distribution characteristics of the speech signal in different frequency bandwidths.

[0096] Furthermore, in step S102, the standardized speech data is processed using an Fbank filter bank to extract Mel-frequency cepstral coefficients to obtain an initial feature set;

[0097] Perform time-frequency transformation on the initial feature set through short-time Fourier transform to generate a multi-dimensional frequency domain feature mapping matrix and determine the frequency domain feature distribution;

[0098] If the dimension of the frequency domain feature mapping matrix exceeds the preset threshold, the principal component analysis algorithm is used to reduce the dimension of the matrix to obtain the reduced dimension feature matrix;

[0099] According to the feature matrix after dimensionality reduction, the energy distribution characteristics of each frequency bandwidth are calculated to obtain the energy distribution vector;

[0100] If the variance of the energy distribution vector is lower than a preset threshold, the edge of the feature matrix is ​​enhanced by the Laplace operator to obtain an enhanced feature matrix;

[0101] The K-means clustering algorithm is used to classify the enhanced feature matrix to obtain feature clusters with different frequency bandwidths;

[0102] By performing statistical analysis on the feature clusters, the frequency bandwidth energy distribution characteristics of the speech signal are generated and the final feature representation is determined;

[0103] Based on the final feature representation, the preset signal processing tools are used to clean the data to obtain a standardized feature set;

[0104] The normalized feature set is grouped according to the channel dimension through the matrix decomposition method to determine the feature distribution after grouping.

[0105] Specifically, in step S102, the standardized speech data is processed using an Fbank filter bank, specifically setting the number of filter banks to 40 and the sampling frequency to 16 kHz, extracting Mel-frequency cepstral coefficients to obtain an initial feature set containing 39-dimensional features;

[0106] The initial feature set is transformed into a time-frequency matrix by short-time Fourier transform, with a window length of 25ms and a window shift of 10ms. A 256-dimensional multidimensional frequency domain feature mapping matrix is ​​generated to determine the frequency domain feature distribution.

[0107] If the dimension of the frequency domain feature map matrix exceeds the preset threshold of 256 dimensions, the principal component analysis algorithm is used to reduce the matrix dimension, retaining the first 80 principal components to obtain the reduced dimension feature matrix;

[0108] According to the feature matrix after dimensionality reduction, the energy distribution characteristics of each frequency bandwidth are calculated. The bandwidth is set to 200 Hz, and an energy distribution vector containing 20 elements is obtained.

[0109] If the variance of the energy distribution vector is lower than the preset threshold of 0.01, the edge of the feature matrix is ​​enhanced by the Laplacian operator, and the convolution kernel size is set to 3x3 to obtain the enhanced feature matrix;

[0110] The K-means clustering algorithm is used to classify the enhanced feature matrix, and the number of clusters is set to 10 to obtain feature clusters with different frequency bandwidths;

[0111] By performing statistical analysis on the feature clusters, calculating the average energy and standard deviation of each cluster, the frequency bandwidth energy distribution characteristics of the speech signal are generated and the final feature representation is determined;

[0112] Based on the final feature representation, the data was cleaned using a preset signal processing tool, and the noise threshold was set to -30dB to obtain a normalized feature set.

[0113] The normalized feature set is grouped according to the channel dimension through the matrix decomposition method. Singular value decomposition is used, the number of channels is set to 5, and the feature distribution after grouping is determined.

[0114] like Figure 1-Figure 2 As shown, in step S103, group processing is performed according to the channel dimension of the multi-dimensional frequency domain feature mapping matrix, and each group of feature data is input into the corresponding residual block for nonlinear transformation. The feature gradient propagation is maintained through the residual connection to obtain a preliminary deep speech feature representation.

[0115] Furthermore, in step S103, the feature data is grouped according to the channel dimension of the multi-dimensional frequency domain feature mapping matrix to obtain a grouped feature data set;

[0116] By using the matrix decomposition method to structure the data sets after grouping, the feature distribution after grouping is determined;

[0117] If the feature distribution after grouping meets the preset balance condition, each group of feature data is input into the corresponding residual block for processing;

[0118] If the equilibrium condition is not met, the distribution is adjusted by resampling the data to obtain the adjusted feature grouping;

[0119] According to the adjusted feature grouping, the residual block is used to perform nonlinear transformation on each group of data to obtain the transformed feature vector;

[0120] Through the residual connection mechanism, the feature gradient in the transformation process is maintained and the distribution of the maintained feature gradient is determined;

[0121] According to the maintained feature gradient distribution, a convolutional neural network is used to perform multi-layer feature extraction on the transformed feature vector to obtain the extracted feature set;

[0122] If the extracted feature set meets the preset representation accuracy, a preliminary deep speech feature representation is output; if not, the parameters of the residual block are adjusted through a feedback mechanism to obtain an optimized feature representation;

[0123] Based on the optimized feature representation, the propagation and maintenance of the feature gradient are continuously monitored to obtain the feature stability data after monitoring;

[0124] By using the monitored feature stability data and combining it with statistical analysis methods, key speech information is extracted to determine the final deep speech feature output.

[0125] Specifically, in step S103, the feature data is divided into 8 groups according to the channel dimension of the multi-dimensional frequency domain feature mapping matrix, each group contains 128 feature vectors;

[0126] The singular value decomposition method is used to structure each group of data and determine the characteristic distribution. If the variance of the characteristic distribution after grouping is less than 0.05, it meets the preset equilibrium condition;

[0127] If the conditions are not met, the feature vector is adjusted to 2 times the sampling rate through data resampling to obtain a new feature grouping;

[0128] The adjusted feature groups are input into the residual block, and the ReLU activation function is used for nonlinear transformation to obtain the transformed feature vector;

[0129] Through the residual connection mechanism, the input features are added to the output features to maintain the feature gradient distribution and ensure that the gradient does not decay;

[0130] Based on the maintained feature gradient distribution, a convolutional neural network is used to set up three convolution layers. Each layer uses 64 3x3 convolution kernels to perform multi-layer extraction on the feature vector to obtain the extracted feature set.

[0131] If the mean square error of the extracted feature set is less than 0.01 and meets the preset representation accuracy, the preliminary deep speech feature representation is output;

[0132] When it is not satisfied, the weight of the residual block is adjusted through the back-propagation algorithm to optimize the feature representation;

[0133] Based on the optimized feature representation, use the gradient monitoring tool to record feature gradient changes in real time and obtain feature stability data;

[0134] Combined with statistical analysis methods, the mean and standard deviation of feature stability data are calculated to extract key speech information and determine the final deep speech feature output;

[0135] The Fbank filter bank is used to process the speech signal, extract 40 Mel-frequency cepstral coefficients, and generate the initial feature set;

[0136] The initial feature set is converted into a time-frequency matrix through short-time Fourier transform to determine the frequency domain feature distribution;

[0137] If the matrix dimension exceeds 256, the principal component analysis algorithm is used to reduce the dimension to 128;

[0138] According to the feature matrix after dimensionality reduction, the energy distribution of each frequency bandwidth is calculated to obtain the energy distribution vector;

[0139] If the variance of the energy distribution vector is less than 0.02, the feature edge is enhanced by the Laplace operator to obtain the enhanced feature matrix;

[0140] The K-means clustering algorithm is used to set 5 cluster centers to classify the enhanced feature matrix and obtain feature clusters with different frequency bandwidths.

[0141] By statistically analyzing the feature clusters, the frequency bandwidth energy distribution characteristics of the speech signal are generated and the final feature representation is determined.

[0142] like Figure 1-Figure 2 As shown, in step S104, the Block_diff_AFF module is used to calculate the difference attention weights between the output features of adjacent residual blocks. If the difference between adjacent feature blocks is higher than the set value, the corresponding weight coefficient is enhanced, and a local optimized feature vector is generated through a weighted fusion mechanism.

[0143] Furthermore, in step S104, output features are obtained from the residual blocks, and the difference between the output features of adjacent residual blocks is calculated using the Block_diff_AFF module to obtain feature difference;

[0144] If the feature difference is higher than the preset threshold, the attention weight is adjusted through the weight enhancement mechanism to determine the enhanced weight coefficient;

[0145] According to the enhanced weight coefficient, the output features of adjacent residual blocks are fused using a weighted fusion mechanism to generate a preliminary local optimized feature vector;

[0146] Perform feature extraction on the preliminary local optimized feature vector through convolutional neural network to obtain deep feature representation;

[0147] Extract local optimized features from deep feature representation, use attention mechanism to weight feature vectors, and obtain optimized feature vectors;

[0148] If the difference of the optimized eigenvector is lower than the preset threshold, the eigenvector is adjusted through the residual connection mechanism to determine the adjusted eigenvector;

[0149] According to the adjusted feature vector, a fully connected layer is used to integrate features to obtain the final local optimized feature vector;

[0150] Through the pre-established feature extraction framework, multi-scale feature data is obtained from the final local optimized feature vector to form an initial feature set;

[0151] According to the initial feature set, the attention fusion module is used to perform correlation analysis on cross-level features, calculate the correlation weights between features at each level, and obtain the weight distribution matrix.

[0152] Specifically, in step S104, output features are extracted from the residual block. Assuming that the output features of residual blocks A and B are vectors A and B respectively, the Euclidean distance between vectors A and B is calculated as the difference through the Block_diff_AFF module. The threshold is set to 0.5. If the calculated difference is 0.6, which is higher than the threshold, the corresponding weight coefficient is increased from 0.8 to 1.0 through the weight enhancement mechanism.

[0153] According to the enhanced weight coefficient, a weighted fusion mechanism is used to perform weighted averaging on vector A and vector B to generate a preliminary local optimized feature vector C;

[0154] Use a convolutional neural network, such as ResNet-50, to extract features from vector C and obtain a deep feature representation D.

[0155] Extract local optimized features from the deep feature representation D, use an attention mechanism such as SENet to weight the feature vector and obtain the optimized feature vector E;

[0156] If the calculated difference of the eigenvector E is 0.4, which is lower than the preset threshold of 0.5, the vector E is residually connected with the original vector A through the residual connection mechanism to obtain the adjusted eigenvector F;

[0157] According to the adjusted feature vector F, a fully connected layer is used for feature integration. Assuming that the fully connected layer contains 256 neurons, the final local optimized feature vector G is output;

[0158] Through a pre-established feature extraction framework, such as VGG-16, multi-scale feature data is obtained from the vector G to form an initial feature set H;

[0159] Based on the initial feature set H, an attention fusion module, such as CBAM, is used to perform correlation analysis on cross-level features and calculate the correlation weights between features at each level. Assuming that the correlation weights of level 1 and level 2 are 0.7 and 0.3 respectively, the weight distribution matrix I is obtained.

[0160] Through the weight distribution matrix I, the multi-scale features are adaptively modulated to generate the modulated feature vector J;

[0161] If the correlation of some features in the modulated feature vector J is lower than the preset threshold of 0.2, the features are subjected to secondary fusion processing, and a convolutional neural network such as InceptionV3 is used for deep extraction to determine the enhanced feature vector K;

[0162] According to the enhanced feature vector K, global semantic information is obtained, and the semantic information is fused with the modulation feature to form the semantic enhancement feature L;

[0163] The feature L is smoothed by information optimization means, such as smoothing filtering, to obtain the final optimized feature data M.

[0164] like Figure 1-Figure 2 As shown, in step S105, a two-dimensional convolution operation is performed on the local optimized feature vector and down-sampling is performed, the number of feature channels is expanded layer by layer to construct a multi-scale feature pyramid, and abstract speech feature representations at different levels are obtained.

[0165] Furthermore, in step S105, the input speech signal is obtained, and frequency domain feature extraction is performed on the original speech data using a preset signal processing tool to obtain multi-dimensional feature data;

[0166] According to the multi-dimensional feature data, the data cleaning method is used to normalize the features to obtain a normalized feature set;

[0167] Through the normalized feature set, the feature data is grouped according to the channel dimension, and the matrix decomposition method is used to structure the data of each group to obtain the feature distribution after grouping;

[0168] If the feature distribution after grouping meets the preset balance condition, the feature data of each group is directly input into the residual block for processing;

[0169] If the equilibrium condition is not met, the distribution is adjusted by resampling the data to obtain the adjusted feature grouping;

[0170] According to the adjusted feature grouping, the residual connection mechanism is used to process the nonlinear transformation in the residual block to obtain the transformed feature vector;

[0171] The transformed feature vector is used to extract local features by a two-dimensional convolution operation, and a convolution feature map is obtained using a preset convolution kernel size.

[0172] If the resolution of the convolution feature map is higher than the preset threshold, downsampling is performed through the pooling operation to reduce the resolution of the feature map and obtain a downsampled feature map;

[0173] If the resolution is lower than or equal to the preset threshold, the convolution feature map is directly used to obtain the downsampled feature map;

[0174] According to the downsampled feature map, the number of feature channels is expanded by stacking convolution layers layer by layer, and a multi-scale feature pyramid is constructed to obtain a multi-scale feature set;

[0175] Through the multi-scale feature set, the global average pooling operation is used to extract the abstract speech features of each level, the multi-level feature information is fused, and the final speech feature representation is obtained through full connection layer mapping.

[0176] Specifically, in step S105, the input speech signal is obtained, and the frequency domain feature extraction is performed on the original speech data using a preset FFT (Fast Fourier Transform) tool to obtain multidimensional feature data containing 256 frequency domain components;

[0177] According to the multidimensional feature data, the normalization method is used to normalize the features, and the amplitude of each frequency domain component is scaled to the interval [0,1] to obtain a normalized feature set;

[0178] Through the normalized feature set, the feature data is divided into 8 groups according to the channel dimension, each group contains 32 frequency domain components, and the PCA (principal component analysis) method is used to structure the data of each group, retaining the first 80% of the principal components to obtain the feature distribution after grouping;

[0179] If the variance ratio of the grouped feature distribution is greater than 0.9, each group of feature data is directly input into the residual block for processing;

[0180] If the variance ratio is less than 0.9, the distribution is adjusted by random resampling to ensure that the variance ratio of each group of data is close to 1, and the adjusted feature grouping is obtained;

[0181] According to the adjusted feature grouping, the residual connection mechanism is used to process the nonlinear transformation in the residual block, and the ReLU activation function and Batch Normalization layer are used to obtain the transformed feature vector;

[0182] Through the transformed feature vector, a 3x3 convolution kernel is used to perform a two-dimensional convolution operation on the feature vector to extract local features and obtain a convolution feature map with a resolution of 128x128;

[0183] If the resolution of the convolution feature map is higher than the preset 64x64 threshold, it is downsampled through a 2x2 maximum pooling operation to reduce the feature map resolution to 64x64 to obtain a downsampled feature map;

[0184] If the resolution is lower than or equal to the preset threshold, the convolution feature map is directly used to obtain the downsampled feature map;

[0185] According to the downsampled feature map, by stacking convolution layers layer by layer, each convolution layer adds 32 feature channels, gradually expanding from 64 channels to 256 channels, constructing a multi-scale feature pyramid, and obtaining a multi-scale feature set containing feature maps of different scales;

[0186] Through the multi-scale feature set, the global average pooling operation is used to extract the abstract speech features of each level, the feature vectors of each level are spliced, the multi-level feature information is fused, and the final speech feature representation is obtained through mapping through a fully connected layer containing 1024 neurons.

[0187] like Figure 1-Figure 2 As shown, in step S106, the correlation weight distribution between cross-level features is calculated through the AFF attention fusion module, and the multi-scale features are adaptively modulated according to the weight distribution to generate a global optimization feature containing global semantic information.

[0188] Furthermore, in step S106, multi-scale feature data is extracted from the input data through a pre-built feature extraction framework to form an initial feature set;

[0189] Based on the initial feature set, the AFF attention fusion module is used to perform correlation analysis on cross-level features, calculate the correlation weights between features at each level, and obtain the weight distribution matrix;

[0190] Through the weight distribution matrix, adaptive modulation processing is performed on the multi-scale features to generate the modulated feature vector;

[0191] If the correlation of some features in the modulated feature vector is lower than the preset threshold, the features are subjected to secondary fusion processing, and a convolutional neural network is used to deeply extract the features to determine the enhanced feature vector;

[0192] According to the enhanced feature vector, global semantic information is extracted and fused with the modulation feature to obtain the semantic enhancement feature;

[0193] By enhancing the semantic features, information optimization is used to smooth the features and obtain the final global optimized feature data;

[0194] If the distribution of the global optimization feature data in a specific dimension does not meet the preset equilibrium condition, the distribution is adjusted by data resampling to obtain the adjusted feature distribution;

[0195] According to the adjusted feature distribution, the residual connection mechanism is used to perform nonlinear transformation on the features, maintain the feature gradient during the transformation process, and determine the feature representation after the transformation;

[0196] Through the transformed feature representation, a multi-layer convolutional neural network is used to perform deep feature extraction, determine whether the extracted features meet the preset representation accuracy, and obtain the final optimized feature output.

[0197] Specifically, in step S106, multi-scale feature data is extracted from the input data through a pre-built feature extraction framework to form an initial feature set;

[0198] For example, the input data is a 3-second speech clip. After short-time Fourier transform, a 256×128 time-frequency spectrum matrix is ​​generated. The first three layers of the VGG16 network are used to extract feature maps of sizes 64x64, 32x32, and 16x16, respectively, to form the initial feature set.

[0199] Based on the initial feature set, the AFF attention fusion module is used to perform correlation analysis on cross-level features, calculate the correlation weights between features at each level, and obtain the weight distribution matrix;

[0200] Specifically, the attention mechanism is used to calculate the correlation weight between the 64x64 feature map and the 32x32 feature map, with the weight value ranging from 0 to 1, forming an 8x8 weight distribution matrix;

[0201] Through the weight distribution matrix, adaptive modulation processing is performed on the multi-scale features to generate the modulated feature vector;

[0202] For example, the weight distribution matrix is ​​multiplied element-by-element with the 32x32 feature map to obtain the modulated feature vector;

[0203] If the correlation of some features in the modulated feature vector is lower than the preset threshold of 0.2, the features are subjected to secondary fusion processing, and a convolutional neural network is used to deeply extract the features to determine the enhanced feature vector;

[0204] Specifically, the low-correlation features are extracted twice using the 3x3 convolution kernel and the ReLU activation function to obtain the enhanced feature vector;

[0205] According to the enhanced feature vector, global semantic information is extracted and fused with the modulation feature to obtain the semantic enhancement feature;

[0206] For example, global average pooling is used to extract global semantic information and concatenate it with the modulated feature vector to obtain semantically enhanced features;

[0207] By enhancing the semantic features, information optimization is used to smooth the features and obtain the final global optimized feature data;

[0208] Specifically, Batch Normalization is used to smooth the semantic enhancement features to ensure that the mean of the features is 0 and the variance is 1;

[0209] If the distribution of the global optimization feature data in a specific dimension does not meet the preset equilibrium condition, that is, the standard deviation is greater than 0.5, the distribution is adjusted by resampling the data to obtain the adjusted feature distribution;

[0210] For example, the feature data is resampled using the random resampling method to reduce the standard deviation to 0.3;

[0211] According to the adjusted feature distribution, the residual connection mechanism is used to perform nonlinear transformation on the features, maintain the feature gradient during the transformation process, and determine the feature representation after the transformation;

[0212] Specifically, the ResNet residual block is used to perform nonlinear transformation on the features, and the feature gradient is maintained through the jump connection;

[0213] Through the transformed feature representation, a multi-layer convolutional neural network is used to perform deep feature extraction to determine whether the extracted features meet the preset representation accuracy, that is, the feature entropy is less than 0.1, and obtain the final optimized feature output;

[0214] For example, a three-layer convolutional neural network is used to extract features, with each layer using a 3x3 convolution kernel and a ReLU activation function, and the final output is optimized features that meet the preset accuracy.

[0215] like Figure 1-Figure 2 As shown, in step S107, the local optimized feature vector and the global optimized feature are spliced ​​according to the feature dimension, and the spliced ​​high-dimensional features are nonlinearly mapped through the SiLU activation function to obtain a joint feature vector that integrates local details and global semantics.

[0216] Furthermore, in step S107, local optimized feature vectors and global optimized feature vectors are extracted from the input data, and processed using a preset feature extraction model to obtain local detail information and global semantic information;

[0217] If the dimensions of the local optimized feature vector and the global optimized feature vector are inconsistent, the dimension alignment operation is used to adjust the two to the same dimension to obtain feature vectors with the same dimension.

[0218] For feature vectors with the same dimension, concatenate them according to the feature dimension to generate a high-dimensional feature vector;

[0219] The high-dimensional feature vector is nonlinearly mapped through the SiLU activation function to obtain the preliminary fused feature vector;

[0220] The batch normalization method is used to normalize the initial fused feature vector to obtain the normalized feature vector;

[0221] Extract key features from the normalized feature vector and use principal component analysis algorithm to perform dimensionality reduction to obtain a low-dimensional joint feature vector;

[0222] According to the low-dimensional joint feature vector, the multi-scale feature data is obtained by combining the pre-established feature extraction framework to obtain the initial feature set;

[0223] Through the initial feature set, the attention fusion module is used to perform correlation analysis on cross-level features, calculate the correlation weights between features at each level, and obtain the weight distribution matrix;

[0224] According to the weight distribution matrix, adaptive modulation processing is performed on the multi-scale features to generate a modulated feature vector.

[0225] Specifically, in step S107, local optimized feature vectors and global optimized feature vectors are extracted from the input data, and processed using a preset feature extraction model, such as using a convolutional neural network (CNN) to extract local detail features of the image, and using a global average pooling layer to extract global semantic features, thereby obtaining local detail information and global semantic information;

[0226] If the dimensions of the local optimized feature vector and the global optimized feature vector are inconsistent, for example, the local feature dimension is 256 and the global feature dimension is 128, then a dimension alignment operation is performed, such as using a linear transformation layer to expand the global feature to 256 dimensions, to obtain a feature vector with the same dimension;

[0227] For feature vectors with the same dimension, concatenation is performed according to the feature dimension to generate a high-dimensional feature vector. For example, a 256-dimensional local feature and a 256-dimensional global feature are concatenated according to the dimension to form a 512-dimensional high-dimensional feature vector.

[0228] Perform nonlinear mapping on high-dimensional feature vectors through SiLU activation function. For example, use SiLU function to activate 512-dimensional feature vectors to obtain preliminary fused feature vectors.

[0229] The batch normalization method is used to normalize the initially fused feature vectors, such as calculating the mean and variance of the feature vectors and performing normalization operations to obtain normalized feature vectors.

[0230] Extract key features from the normalized feature vector and use the principal component analysis algorithm to reduce the dimension. For example, select the first 50 principal components and reduce the dimension of the 512-dimensional feature vector to 50 dimensions to obtain a low-dimensional joint feature vector.

[0231] Based on the low-dimensional joint feature vector, multi-scale feature data is obtained in combination with a pre-established feature extraction framework. For example, a multi-scale convolutional layer is used to extract features of different scales to obtain an initial feature set.

[0232] Based on the initial feature set, the attention fusion module is used to perform correlation analysis on cross-level features and calculate the correlation weights between features at each level. For example, the self-attention mechanism is used to calculate the similarity between features and obtain a weight distribution matrix.

[0233] According to the weight distribution matrix, adaptive modulation processing is performed on multi-scale features. For example, weighted summation of features is performed according to weights to generate a modulated feature vector, ensuring that the information in the feature vector is effectively integrated, providing richer feature representation for subsequent classification or recognition tasks.

[0234] like Figure 1-Figure 2 As shown, in step S108, the AdamW optimizer is used to train the four-classification neural network model, and the joint feature vector is used as the network input to be mapped to the probability distribution of the four categories through the fully connected layer, and the classification results of normal, subjective cognitive decline, mild cognitive impairment, and Alzheimer's disease are output.

[0235] Furthermore, in step S108, the original speech signal is preliminarily cleaned by a preset signal processing tool to obtain multi-dimensional feature data and obtain a normalized feature set;

[0236] Based on the normalized feature set, the feature data is grouped by channel dimension, and the matrix decomposition method is used for structured arrangement to determine the feature distribution after grouping.

[0237] If the feature distribution after grouping meets the preset balance condition, each group of feature data is input into the corresponding residual block for processing;

[0238] If the equilibrium condition is not met, the distribution is adjusted by resampling the data to obtain the adjusted feature grouping;

[0239] The adjusted feature grouping is nonlinearly transformed through the residual block, and the residual connection mechanism is used to maintain the feature gradient during the transformation process to obtain the transformed feature vector;

[0240] Based on the transformed feature vector, a convolutional neural network is used to perform multi-layer feature extraction to determine whether the extracted features meet the preset representation accuracy and obtain deep speech feature representation;

[0241] If the deep speech feature representation meets the preset representation accuracy, it is normalized as a joint feature vector to obtain normalized feature data;

[0242] If it is not satisfied, the residual block parameters are adjusted through the feedback mechanism to obtain the optimized feature representation;

[0243] The AdamW optimizer is used to initialize the parameters of the four-classification neural network model, and the normalized feature data is used for training to obtain the preliminarily trained neural network.

[0244] The output of the preliminarily trained neural network is mapped to a probability distribution through a fully connected layer, and the probability of each sample belonging to normal, subjective cognitive decline, mild cognitive impairment, and Alzheimer's disease is calculated to obtain the initial classification probability;

[0245] If the maximum value of the initial classification probability is lower than the preset threshold, the joint feature vector is enhanced, and the four-classification neural network model is retrained using the enhanced feature data. The parameters are adjusted through the AdamW optimizer to obtain the optimized neural network and determine the final classification result.

[0246] Specifically, in step S108, the original speech signal is preliminarily cleaned by using a preset signal processing tool, such as converting the time domain signal into a frequency domain signal using Fourier transform and removing noise using a bandpass filter to obtain multi-dimensional feature data including fundamental frequency, formant, etc., and obtain a normalized feature set;

[0247] Based on the normalized feature set, the feature data is grouped and processed by channel dimension. For example, the feature data is divided into fundamental frequency channel, formant channel, etc. The singular value decomposition method is used to structure each group of data and determine the feature distribution after grouping.

[0248] If the feature distribution after grouping meets the preset balance condition, for example, the variance of the feature values ​​of each channel is less than the preset threshold value of 0.05, then each group of feature data is input into the corresponding residual block for processing;

[0249] If the equilibrium condition is not met, the distribution is adjusted by resampling the data. For example, the SMOTE algorithm is used to oversample the minority class samples to obtain the adjusted feature grouping.

[0250] Perform nonlinear transformation on the adjusted feature grouping through residual blocks, such as using ReLU activation function and batch normalization layer, and adopting residual connection mechanism to maintain the feature gradient during the transformation process to obtain the transformed feature vector;

[0251] Based on the transformed feature vector, a convolutional neural network is used to perform multi-layer feature extraction. For example, three convolutional layers are set, each using 64 3x3 convolution kernels. The extracted features are judged to see if they meet the preset representation accuracy, for example, if the feature entropy value is greater than 0.8, to obtain the deep speech feature representation.

[0252] If the deep speech feature representation meets the preset representation accuracy, it is normalized as a joint feature vector, for example, using a Z-score normalization method to obtain normalized feature data;

[0253] If it is not satisfied, the residual block parameters are adjusted through the feedback mechanism, for example, the learning rate is adjusted to 0.001 to obtain the optimized feature representation;

[0254] The AdamW optimizer is used to initialize the parameters of the four-classification neural network model. For example, the initial learning rate is set to 0.01 and the weight decay is set to 0.001. The normalized feature data is used for training to obtain a preliminarily trained neural network.

[0255] Map the output of the preliminarily trained neural network to a probability distribution through a fully connected layer. For example, set the fully connected layer to contain 128 neurons and use the Softmax function to calculate the probability of each sample belonging to normal, subjective cognitive decline, mild cognitive impairment, or Alzheimer's disease to obtain the initial classification probability;

[0256] If the maximum value of the initial classification probability is lower than a preset threshold, such as 0.7, the joint feature vector is enhanced, for example, by using data enhancement techniques such as random noise addition, and the enhanced feature data is used to retrain the four-classification neural network model. The parameters are adjusted through the AdamW optimizer, for example, the learning rate is adjusted to 0.005, to obtain the optimized neural network and determine the final classification result.

[0257] Those skilled in the art can make various other corresponding changes and deformations based on the technical solutions and concepts described above, and all of these changes and deformations should fall within the scope of protection of the claims of this application.

Claims

1. A multi-scale grouping feature fusion AD screening method based on speech features, characterized by: include: S101, collecting the original speech signal of the subject, unifying the sampling rate through a frequency reduction algorithm, intercepting continuous speech segments based on speech quality assessment indicators, and generating standardized speech data; S102, extracting Mel-frequency cepstral coefficients from the standardized speech data, generating a multi-dimensional frequency domain feature mapping matrix, and obtaining energy distribution characteristics of different frequency bandwidths; S103, grouping the multidimensional frequency domain feature mapping matrix according to the channel dimension, performing nonlinear transformation on the input residual blocks of each group, maintaining the feature gradient through residual connection, and outputting deep speech feature representation; S104, calculating the difference of output features of adjacent residual blocks, increasing the weight coefficient when the difference is higher than a set value, and weighted fusion to generate a local optimized feature vector; S105, performing a two-dimensional convolution operation on the locally optimized feature vector and performing downsampling processing, expanding the number of feature channels by increasing the number of convolution kernels layer by layer, constructing a feature pyramid, and extracting hierarchical abstract speech features; S106, calculating the cross-level feature correlation weights of the feature pyramid, and adaptively modulating to generate global optimization features containing global semantic information; S107, splicing the local optimized feature vector and the global optimized feature, performing nonlinear mapping through SiLU activation function, and outputting a joint feature vector; S108. Input the joint feature vector into a four-category neural network model trained by the AdamW optimizer, map it to a four-category probability distribution through a fully connected layer, and output classification results of normal, subjective cognitive decline, mild cognitive impairment, and Alzheimer's disease.

2. The multi-scale grouping feature fusion AD screening method based on speech features according to claim 1 is characterized in that: Collect voice signals through multi-channel equipment and use spectrum analysis to adjust the sampling rate to the preset standard; Extracting continuous speech segments from the adjusted spectral data based on a voice activity detection algorithm; When the background noise power spectral density of the continuous speech segment exceeds -50dB, spectral subtraction is used to reduce noise and generate purified speech data; Mel-frequency cepstral coefficients are extracted from the purified speech data, and after quality evaluation, segments with scores lower than 0.7 are removed to generate the standardized speech data.

3. The multi-scale grouping feature fusion AD screening method based on speech features according to claim 2 is characterized in that: 40 groups of Fbank filters are used to extract 39 Mel-frequency cepstral coefficients from the standardized speech data; A 256-dimensional feature map matrix is ​​generated by short-time Fourier transform, with a window length of 25ms and a window shift of 10ms; When the dimension of the feature map matrix exceeds 256, principal component analysis is used to reduce the dimension to 128, where 128 dimensions refers to 128 eigenvalues; For the feature matrix whose energy distribution vector variance after dimensionality reduction is less than 0.02, 3×3 Laplacian operator edge enhancement is performed.

4. The multi-scale grouping feature fusion AD screening method based on speech features according to claim 1 is characterized in that: The feature map matrix after dimensionality reduction or edge enhancement is divided into 8 groups by channel, with 128 feature vectors in each group; When the variance ratio of the within-group characteristic distribution is less than 0.9, 2-fold frequency resampling is used to adjust the distribution; Among them, variance ratio: calculate the variance of each group of eigenvectors, and calculate the ratio of this variance to the total variance of all eigenvectors, which is defined as the intra-group feature distribution variance ratio, 2x frequency resampling: if the intra-group feature distribution variance ratio is less than 0.9, then the group of eigenvectors is upsampled by 2 along the time dimension, or replicated and expanded along the feature dimension. The adjusted features are grouped and input into the residual block, which performs ReLU activation and batch normalization in sequence, and maintains the gradient through the residual connection; A three-layer convolutional network is used to extract features, with 64 3×3 convolution kernels in each layer. The mean square error between the output features and the expected labels is calculated through the back propagation algorithm. If the error is less than 0.01, it is judged to be a deep speech feature representation.

5. The multi-scale grouping feature fusion AD screening method based on speech features according to claim 4 is characterized in that: Calculate the Euclidean distance between the output features of adjacent residual blocks as the difference; When the difference exceeds 0.5, the weight coefficient is increased from 0.8 to 1.0; The weighted average fusion feature is used and then input into ResNet-50 to extract the deep representation; When the weighted feature difference is lower than 0.5, the feature vector is adjusted through residual connection to generate a locally optimized feature vector.

6. The multi-scale grouping feature fusion AD screening method based on speech features according to claim 5 is characterized in that: Performing a 3×3 two-dimensional convolution on the locally optimized feature vector to generate a 128×128 feature map; When the feature map resolution is higher than 64×64, 2×2 maximum pooling is used to downsample to 64×64; The number of channels is increased from 64 to 256 by stacking convolutional layers layer by layer to construct a multi-scale feature pyramid. Abstract speech features at each level are extracted through global average pooling.

7. The multi-scale grouping feature fusion AD screening method based on speech features according to claim 6 is characterized in that: The AFF module is used to calculate the cross-level correlation weights of the multi-scale feature pyramid; When the weight value is lower than 0.2, a 3×3 convolution secondary fusion is performed on the corresponding features; The semantic information extracted by global average pooling is integrated and then normalized and smoothed to generate global optimized features. When the feature standard deviation exceeds 0.5, random resampling is used to adjust the distribution.

8. The multi-scale grouping feature fusion AD screening method based on speech features according to claim 7 is characterized in that: When the local optimized feature vector is inconsistent with the global optimized feature dimension, the dimensions are aligned through linear transformation; Generate a high-dimensional vector by splicing according to the feature dimension, and perform batch normalization after activation by the SiLU function; Principal component analysis is used to reduce the dimension of the normalized features to 50 dimensions to generate a joint feature vector.

9. The multi-scale grouping feature fusion AD screening method based on speech features according to claim 8, characterized in that: Performing Z-score normalization on the joint feature vector; The AdamW optimizer was used to initialize the four-classification model, with an initial learning rate of 0.01 and a weight decay of 0.001; The 128-neuron fully connected layer outputs the four-category probabilities, and the Softmax function calculates the probability distribution; When the maximum probability value is lower than 0.7, random noise is added to the feature and the model is retrained.