Classification system and method for Alzheimer's disease language detection based on feature purification network

Through the method of feature purification network, the Transformer model and G-Net and P-Net subnets eliminate suboptimal representations are used to solve the problem of insufficient accuracy in Alzheimer's disease detection, and more efficient early identification and large-scale screening are achieved.

CN113961700BActive Publication Date: 2025-08-19QUANZHOU NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111143501.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-28
Publication Date
2025-08-19
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

The existing Alzheimer's disease detection methods have problems such as insufficient accuracy, strong invasiveness and high cost, especially in early identification and large-scale screening, and deep learning models have suboptimal representation interference in characterization learning, which affects the classification effect.

Method used

Using a method based on feature purification network, semantic feature vectors are extracted through the Transformer model, and the G-Net and P-Net subnets are used to extract common features and purification features, combining gradient inversion layer (GRL) and classifiers to eliminate suboptimal representations to improve the discrimination of features.

Benefits of technology

It significantly improves the classification performance of Alzheimer's disease language detection, reduces interference with suboptimal representation, and enhances the generalization ability and classification accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113961700B_ABST
    Figure CN113961700B_ABST
Patent Text Reader

Abstract

The present invention discloses a classification system and method for Alzheimer's disease language detection based on a feature purification network. The system obtains audio information for language detection and transcribes it to obtain text data. The transcribed text data is passed through a Transformer to obtain a semantic feature vector #imgabs0# corresponding to a G-Net model and a semantic feature vector #imgabs1# corresponding to a P-Net model. The semantic feature vector #imgabs2# is processed by a G-Net to extract common features #imgabs3# between the datasets. The common features #imgabs4# are then output to a G-Net classifier Cc for processing and outputting the final common features #imgabs5#. The P-Net model removes the common features #imgabs7# from the semantic feature vector #imgabs6# to obtain a purified feature vector #imgabs8#. The P-Net classifier Cp then obtains a classification result after the purified features #imgabs9#. The present invention combines a Transformer-based model with a feature purification network, significantly improving classification performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computing data processing technology, and in particular to a classification system and method for Alzheimer's disease language detection based on a feature purification network. Background Art

[0002] Alzheimer's disease (AD) is an irreversible neurodegenerative disease with an insidious onset, making it difficult to detect at all stages. AD can affect patients' daily and social abilities and may even lead to disability.1,2 Currently, AD detection methods mainly include clinical diagnosis, pathological examination, magnetic resonance imaging (MRI), positron emission tomography (PET), and AD biomarkers (such as cerebrospinal fluid, blood, and urine tests). Clinical diagnosis lacks a standardized approach and often requires a combination of expert experience and multiple test results. These methods each have advantages and disadvantages in terms of safety, cost, ease of use, accuracy, and specificity. For example, clinical diagnosis is a rapid screening method, but it requires a face-to-face consultation between a physician and patient, using tests such as the Mini-Mental State Examination (MMSE), the Montreal Cognitive Assessment (MOCA), or the Clinical Dementia Rating (CDR). This process takes approximately one hour, and the diagnostic accuracy is acceptable. Pathological examination is an accurate and reliable method, but it can only be performed after death. MRI and PET examinations are relatively expensive. Reliable AD biomarkers would undoubtedly lead to early identification of the disease, but existing biomarkers (such as cerebrospinal fluid testing and amyloid ligand imaging) are invasive, time-consuming, and expensive. Therefore, they are currently unsuitable for routine or large-scale AD screening. 3-4 Researchers have found that in addition to profoundly affecting patients' mood, attention, memory, and movement, AD also profoundly affects their language function. Language, as a representation of mental activity, clearly reflects the close relationship between language, cognition, and communication. Language disturbances are a common manifestation in patients with Alzheimer's disease and may even precede orientation and memory impairments. 8-9 The picture description task, adapted from the Boston Aphasia Diagnostic Test, 10 has been shown to be sensitive to subtle cognitive deficits and can therefore provide valuable clinical information from spontaneous speech for identifying AD. Natural language processing (NLP) techniques can effectively detect AD based on speech text. Existing research on the diagnosis of spontaneous speech disorders focuses primarily on two approaches. One is manual feature extraction from features, such as acoustic features, 12-15 linguistic features, 16-19 or a combination of these. 20-21 The number of acoustic and linguistic features varies from 4 to 920, and a review 22 reported the most informative features among acoustic and linguistic features, respectively, based on existing research. Studies have also shown that combined features outperform single features 20-21. However, this approach is subjective and requires a lot of expertise. It requires designing different features for different tasks, which are not universal. Another approach is deep learning, which can automatically extract deep semantic features. Based on its powerful representation learning capabilities, deep learning generally outperforms manual methods. In addition, deep learning architectures eliminate the constraints imposed by manual features, improve the generalizability of classifiers, and can be further used in different clinical settings.Representation learning based on deep learning can represent different tasks in the same vector space, leading to cross-task, cross-language and even cross-modal transfer, making artificial intelligence a more common development direction.

[0003] Because the original text is a participant's description of a picture, it should be comprehensive and complete for a healthy person. This means that discriminative words and sentences should be included, while meaningless or ambiguous words should be reduced. For example, accurate descriptive words such as "mother," "tap," and "the stool is tilted" are generally better cognitive signals. Words and sentences such as "I don't know," "hmm," and "pause" are likely indicators of poor cognitive status and are discriminative in AD recognition. However, ambiguous, irrelevant, or even irrelevant words and sentences, such as "Isn't that enough?" and "Great, there might be a little breeze," are not helpful for classification. They interfere with deep learning representation learning by producing suboptimal representations, where words have little or no effect, potentially interfering with the final classification results. While attention mechanisms can mitigate this impact by assigning higher or lower weights to relevant words, their inaccuracies and data peculiarities make them ineffective. Currently, machine learning methods for AD diagnosis have advanced to the level of deep learning, but research on improving deep learning representation learning in this field is limited. This paper uses a novel feature purification method to improve representation learning to obtain more discriminative feature vectors for detecting AD. Although deep learning models have made great progress in generating more discriminative features through powerful representation learning, to the best of our knowledge, no research has yet been conducted on identifying AD from spontaneous speech by purifying deep learned representations. Summary of the Invention

[0004] The purpose of the present invention is to provide a classification system and method for Alzheimer's disease language detection based on a feature purification network.

[0005] The technical solution adopted in the present invention is:

[0006] A classification method for Alzheimer's disease language detection based on a feature purification network includes the following steps:

[0007] Step 1: Obtain audio information for language detection and transcribe it to obtain text data;

[0008] Step 2: Pass the transcribed text data through Transformer to obtain the semantic feature vector f corresponding to the G-Net model c And the semantic feature vector f corresponding to the P-Net model p ;

[0009] Step 3, semantic feature vector f cAfter G-Net processing to extract common features f between datasets c ′, and then the common feature f c ′ is output to the classifier Cc of G-Net for processing and outputs the final common feature f c ';

[0010] Step 4: The P-Net model obtains the final common feature f output by the classifier Cc c ′, and from the semantic feature vector f p Remove common features f c ′Get the purified feature vector f w , the purified feature vector f w The classification result Y after the purified features is obtained by the classifier Cp provided to P-Net Re sult .

[0011] Furthermore, the Transformer encoder in step 2 includes a self-attention layer, a first normalization layer, a feedforward layer, and a second normalization layer arranged in sequence;

[0012] The self-attention layer receives two input vectors x1 and x2 and obtains vectors z1 and z2, which are passed to the first normalization layer respectively. The first normalization layer normalizes the two output vectors z1 and z2 of the self-attention layer and outputs them to a feedforward layer respectively. The normalized vectors z1 and z2 are output to the second normalization layer through the corresponding feedforward layer respectively, and are normalized by the second normalization layer to obtain vectors r1 and r2. r1 and r2 replace x1 and x2 and are input to the encoder again. The cycle is repeated six times to obtain the final semantic feature vector f. c and f p .

[0013] Furthermore, the maximum length of the speech text data does not exceed 500, and the length of the word embeddings x1 and x2 is set to 500.

[0014] Furthermore, the output function expression of the first normalization layer is as follows:

[0015] z1=LN(x1+z1), (1)

[0016] z2=LN(x2+z2) (3)

[0017] Among them, x1 represents the output of the Embedding layer, and the x1 layer passes through the self-attention layer to obtain z1; the x1+z1 vector is normalized to obtain the new z1 layer; x2 represents the output of the Embedding layer, and the x2 layer passes through the self-attention layer to obtain z2; the x2+z2 vector is normalized to obtain the new z2 layer; LN means layer normalization;

[0018] Furthermore, the output function expression of the second normalization layer is as follows:

[0019] r1=LN(z1+FeedForward(z1)) (2)

[0020] r2=LN(z2+FeedForward(z2)) (4)

[0021] Among them, z1 passes through the feedforward neural network layer, passes through the residual structure and is added to itself, and also passes through the LN layer to finally output the vector r1; z2 passes through the feedforward neural network layer, passes through the residual structure and is added to itself, and also passes through the LN layer to finally output the vector r2.

[0022] Furthermore, the semantic feature vector f in step 3 c GRL is used to train the common features close to the real ones. GRL acts as an identity transformation during the forward propagation process, and then takes the gradient from the next layer and changes its value during the backward propagation process, that is, multiplying it by λ; and then passes it to the next layer.

[0023] Specifically, without loss of generality, the GRL processing function can regard the gradient reversal layer as a "pseudo-function" defined by two incompatible equations describing its forward and backward propagation behaviors as follows:

[0024] G λ (x)=x (7)

[0025]

[0026] Among them, λ is a hyperparameter, x is the input vector of GRL, that is, the semantic feature vector f c , to get G λ (f c )=f c ′.

[0027] Furthermore, the function expression of the classifier Cc in step 3 is:

[0028] Y GRL =soft max(W c *f c ′+b c ) (9)

[0029] The corresponding loss function is:

[0030] Loss c =CrossEntropy(Y True , Y GRL ) (10)

[0031] Among them, Y GRLis the output vector of the G-Net neural network, which is the f that is finally provided to P-net c ′; Wc, bc are the weights and biases of the classifier Cc, f c ′ is the extracted common feature; Y True is an ideal vector; common features of different classes are extracted by optimizing the Lossc feature extractor Fc.

[0032] Furthermore, G-Net and P-Net have the same input text x, and the two sub-networks G-Net and P-Net have the same structure but do not share parameters.

[0033] Furthermore, the function expression of the classifier Cp in step 4 is:

[0034] Y Re sult =soft max(W p *f w +b p ) (12)

[0035] The corresponding loss function is:

[0036] Loss p =CrossEntropy(Y True , Y Re sult ) (13)

[0037] Among them, Y Re sult is the output vector of the P-Net neural network, W p 、b p For the classifier C p The weights and biases of w is the final purified feature vector, CrossEntropy is the cross entropy loss function, Loss p Measure a neural network output vector Y Re sult and the ideal vector Y True The degree of closeness; by optimizing Loss p Feature Extractor F p Extract common features of different classes.

[0038] A classification system for Alzheimer's disease language detection based on a feature purification network applies the classification method for Alzheimer's disease language detection based on a feature purification network. The GP-Net network includes two sub-networks, G-Net and P-Net. G-Net and P-Net have the same input text x. The two sub-networks G-Net and P-Net have the same structure but do not share parameters. G-Net includes an input layer X, a feature extractor Fc, a gradient reverse layer, and a classification layer Cc. The feature extractor Fc extracts common features, and the classification layer Cc classifies the common features to obtain a classification result of the common features and provides it to P-Net. P-Net includes an input layer X, a feature extractor Fp, and a classification layer Cp. The common features extracted from the feature extractor Fc are deleted from the feature vector extracted by the feature extractor Fp to obtain a purified feature. The classification layer Cp classifies the purified features to obtain a purified feature classification result.

[0039] Furthermore, the feature extractor Fp and the feature extractor Fc together form the Transformer encoder of the GP-Net network; the Transformer encoder includes a self-attention layer, a first normalization layer, a feedforward layer and a second normalization layer arranged in sequence; the self-attention layer receives two input vectors x1 and x2 and obtains vectors z1 and z2, which are passed to the first normalization layer respectively; the first normalization layer normalizes the two output vectors z1 and z2 of the self-attention layer and outputs them to a feedforward layer respectively; the normalized vectors z1 and z2 are output to the second normalization layer respectively through the feedforward layer, and are normalized by the second normalization layer to obtain vectors r1 and r2, and r1 and r2 replace x1 and x2 and are input into the encoder again, and the cycle is repeated six times to obtain the final semantic feature vector f c and f p .

[0040] The present invention adopts the above technical solution, combines the Transformer-based model with the feature purification network, and greatly improves the classification performance. The present invention pre-trains the Transformer and fine-tunes the model on a new dataset so that the learned knowledge can be transferred to the text classification task of the present invention. The work of the present invention is significantly different from previous research on AD recognition through NLP technology, because previous research has not improved the representation learning of deep learning. Transformer is still the most widely used deep learning algorithm, but the time complexity of self-attention is 0(n2), which hinders the development of the model, so it is of great significance to improve the efficiency of Transformer in the future. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments;

[0042] Figure 1 Schematic diagram of the network structure of GP-Net of the present invention;

[0043] Figure 2 It is a data flow chart of the encoder of the present invention;

[0044] Figure 3 Schematic diagram of the two-dimensional mathematical principle of the purification feature of the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0046] like Figures 1 to 3 As shown in FIG1 , the present invention discloses a classification method for Alzheimer's disease language detection based on a feature purification network, which includes the following steps:

[0047] Step 1: Obtain audio information for language detection and transcribe it to obtain text data;

[0048] Step 2: Pass the transcribed text data through Transformer to obtain the semantic feature vector f corresponding to the G-Net model c And the semantic feature vector f corresponding to the P-Net model p ;

[0049] Step 3, semantic feature vector f c After G-Net processing to extract common features f between datasets c ′, and then the common feature f c ′ is output to the classifier Cc of G-Net for processing and outputs the final common feature f c ';

[0050] Step 4: The P-Net model obtains the final common feature f output by the classifier Cc c ′, and from the semantic feature vector f p Remove common features f c ′Get the purified feature vector f w , the purified feature vector f w The classification result Y after the purified features is obtained by the classifier Cp provided to P-Net Re sult .

[0051] Furthermore, in step 2, the Transformer encoder includes a self-attention layer, a first normalization layer, a feedforward layer, and a second normalization layer arranged in sequence;

[0052] The self-attention layer receives two input vectors x1 and x2 and obtains vectors z1 and z2, which are passed to the first normalization layer respectively. The first normalization layer normalizes the two output vectors z1 and z2 of the self-attention layer and outputs them to a feedforward layer respectively. The normalized vectors z1 and z2 are output to the second normalization layer through the corresponding feedforward layer respectively, and are normalized by the second normalization layer to obtain vectors r1 and r2. r1 and r2 replace x1 and x2 and are input to the encoder again. The cycle is repeated six times to obtain the final semantic feature vector f. c and f p .

[0053] Specifically, if Figure 2 As shown, the self-attention layer receives the n-gram encoding of two input vectors x1 and x2 (dark gray) to obtain the output of the word embedding layer. After adding the position encoding, the final feature vector x1 (light gray) input to the Transformer encoder is obtained. x1 then passes through the Self-Attention layer to obtain the vector z1 (dark gray). x1 and z1 are added together and layer normalization (LN) is performed to obtain the new z1 and z2 (light gray). The new z1 and z2 (light gray) are added to themselves through the FeedForward layer and the residual structure, and also pass through the LN layer to obtain the final output vectors r1 and r2. At the same time, r1 and r2 will also serve as the input of the next layer of encoder, and the cycle continues until the encoder output of the last layer. Here, the dimensions of x1, z1, r1 (x2, z2, r2) are all 512. The formula is:

[0054] z1=LN(x1+z1), (1)

[0055] r1=LN(z1+FeedForward(z1)) (2)

[0056] z2=LN(x2+z2) (3)

[0057] r2=LN(z2+FeedForward(z2)) (4)

[0058] Here, x1 represents the output of the Embedding layer. The x1 layer passes through the Self-Attention layer to obtain z1. The x1+z1 vector is normalized to obtain the new z1 layer. x2 represents the output of the Embedding layer. The x2 layer passes through the Self-Attention layer to obtain z2. The x2+z2 vector is normalized to obtain the new z2 layer. LN means layer normalization.

[0059] z1 passes through the feedforward neural network layer, passes through the residual structure and is added to itself, and also passes through the LN layer to finally output the vector r1; z2 passes through the feedforward neural network layer, passes through the residual structure and is added to itself, and also passes through the LN layer to finally output the vector r2.

[0060] Furthermore, the maximum length of the speech text data does not exceed 500, and the length of the word embeddings x1 and x2 is set to 500.

[0061] Furthermore, the semantic feature vector f in step 3 c GRL is used to train the common features close to the real ones. GRL acts as an identity transformation during the forward propagation process, and then takes the gradient from the next layer and changes its value during the backward propagation process, that is, multiplying it by λ; and then passes it to the next layer.

[0062] Specifically, without loss of generality, the GRL processing function can regard the gradient reversal layer as a "pseudo-function" defined by two incompatible equations describing its forward and backward propagation behaviors as follows:

[0063] G λ (x)=x (7)

[0064]

[0065] Among them, λ is a hyperparameter, x is the input vector of GRL, that is, the semantic feature vector f c , to get G λ (f c )=f c ′.

[0066] Furthermore, the function expression of the classifier Cc in step 3 is:

[0067] Y GRL =soft max(W c *f c ′+b c ) (9)

[0068] The corresponding loss function is:

[0069] Loss c =CrossEntropy(Y True , Y GRL ) (10)

[0070] Among them, Y GRL is the output vector of the G-Net neural network, which is the f that is finally provided to P-net c ′; Wc, bc are the weights and biases of the classifier Cc, f c ′ is the extracted common feature; Y Trueis an ideal vector; common features of different classes are extracted by optimizing the Lossc feature extractor Fc.

[0071] Furthermore, G-Net and P-Net have the same input text x, and the two sub-networks G-Net and P-Net have the same structure but do not share parameters.

[0072] Furthermore, the function expression of the classifier Cp in step 4 is:

[0073] Y Result =soft max(W p *f w +b p ) (12)

[0074] The corresponding loss function is:

[0075] Loss p =CrossEntropy(Y True , Y Result ) (13)

[0076] Among them, Y Re sult is the output vector of the P-Net neural network, W p 、b p For the classifier C p The weights and biases of w is the final purified feature vector, CrossEntropy is the cross entropy loss function, Loss p Measure a neural network output vector Y Re sult and the ideal vector Y True The degree of closeness; by optimizing Loss p Feature Extractor F p Extract common features of different classes.

[0077] A classification system for Alzheimer's disease language detection based on a feature purification network applies the classification method for Alzheimer's disease language detection based on a feature purification network. The GP-Net network includes two sub-networks, G-Net and P-Net. G-Net and P-Net have the same input text x. The two sub-networks G-Net and P-Net have the same structure but do not share parameters. G-Net includes an input layer X, a feature extractor Fc, a gradient reverse layer, and a classification layer Cc. The feature extractor Fc extracts common features, and the classification layer Cc classifies the common features to obtain a classification result of the common features and provides it to P-Net. P-Net includes an input layer X, a feature extractor Fp, and a classification layer Cp. The common features extracted from the feature extractor Fc are deleted from the feature vector extracted by the feature extractor Fp to obtain a purified feature. The classification layer Cp classifies the purified features to obtain a purified feature classification result.

[0078] Furthermore, the feature extractor Fp and the feature extractor Fc together form the Transformer encoder of the GP-Net network; the Transformer encoder includes a self-attention layer, a first normalization layer, a feedforward layer and a second normalization layer arranged in sequence; the self-attention layer receives two input vectors x1 and x2 and obtains vectors z1 and z2, which are passed to the first normalization layer respectively; the first normalization layer normalizes the two output vectors z1 and z2 of the self-attention layer and outputs them to a feedforward layer respectively; the normalized vectors z1 and z2 are output to the second normalization layer respectively through the feedforward layer, and are normalized by the second normalization layer to obtain vectors r1 and r2, and r1 and r2 replace x1 and x2 and are input into the encoder again, and the cycle is repeated six times to obtain the final semantic feature vector f c and f p .

[0079] The specific principle of the present invention is described in detail below:

[0080] like Figure 1 As shown, the entire GP-NET network consists of two parts, G-net and P-net. The purpose of G-Net is to extract common features, because these common features are common to all classes and make no difference to classification. The purpose of P-Net is to further purify features by removing common features in traditional features. G-Net consists of 4 parts, input layer X, feature extractor Fc, gradient reverse layer (GRL) and classification layer Cc. P-Net also consists of 4 parts, input layer X, feature extractor Fp (the features extracted by Fc and Fp have no shared parameters). The main idea of this network is: the feature vector extracted by feature extractor Fp deletes the common features extracted from feature extractor Fc to obtain more discriminative purified features for final classification.

[0081] like Figure 2Figure 2 shows the data flow diagram of the Transformer encoder. LN stands for layer normalization. The dark gray x1 and x2 are the outputs of the embedding layer. The embedding layer adds positional embeddings, resulting in the feature vectors input to the encoder (light gray x1 and x2), with a dimension of 512. Therefore, x1 and x2, represented in light gray, are the word embeddings for the words "Mother" and "is." A sequence of vectors x1 is encoded into another sequence of vectors z1 (light gray) after passing through the self-attention layer. Then, after Eqs. 1, 2, 3, and 4, we obtain the new dark gray vectors z1 and z2, and the light green vectors r1 and r2.

[0082] z1=LN(x1+z1), (1)

[0083] r1=LN(z1+FeedForward(z1)) (2)

[0084] z2=LN(x2+z2) (3)

[0085] r2=LN(z2+FeedForward(z2)) (4)

[0086] At the same time, r1 and r2 are used as the input of the next layer of Encoder to replace x1 and x2. The present invention loops this process six times to obtain the final semantic feature vector. The dimension of the vector xzr is 512. The transcripts in this study are partial transcripts with a maximum length of no more than 500. Therefore, in order to extract the complete semantic information of the sentence, the present invention sets the length of the word embeddings x1 and x2 to 500. GP-Net includes two subnetworks, namely G-Net and P-Net, which both have the same input text xi. The two subnetworks have the same structure, but the parameters are not shared. The feature extractors of G-Net and P-Net are Fc and Fp, respectively. The high-level features fc and fp are obtained from the feature extractors fc and fp after the Transformer layer, respectively. Therefore, the feature vectors fc and fp are recorded as Formula 5 and Formula 6, respectively.

[0087] f c =Transformer c (X) (5)

[0088] f p =Transformer p (X) (6)

[0089] The G-Net module extracts common features: The main goal of the G-Net module is to extract common features between datasets, which are indistinguishable from the classification task. Since common features are features shared by all classes, the classifier cannot effectively use them to distinguish different classes. To obtain common features, GRL is added between the feature extractor Fc and the classifier to reverse the gradient direction. By training the module, common features between different classes are obtained. gλ can be regarded as two incompatible equations that describe the forward and backward propagation behaviors:

[0090] G λ (x)=x (7)

[0091]

[0092] Where λ is a hyperparameter. We process the feature vector f through GRL c , we get f c , such as G λ (f c )=f c ′. In order to make f c ′ is closer to the true common features. GRL acts as an identity transformation during the forward propagation process, and then takes the gradient from the next layer and changes its value (i.e., multiplying it by λ) during the backward propagation process, and then passes it to the next layer. This ensures that the feature distribution is similar and is as indistinguishable as possible to the classifier. In this way, we can obtain common features shared between classes. Finally, f c is sent to the classifier C c .

[0093] Y GRL =soft max(W c *f c ′+b c ), (9)

[0094] Loss c =CrossEntropy(Y True , Y GRL ), (10)

[0095] Among them, W c 、b c are the weights and biases of the classifier Cc. By optimizing Lossc, the feature extractor Fc can extract common features of different classes.

[0096] Using the P-Net model to calculate the purified features: The main purpose of the P-Net model is to extract semantic information from the input instance and classify the features. This paper introduces the mathematical principles in two-dimensional space. Figure 3As shown, the method of the present invention is to delete the common features in the Transformer extractor Fn, which not only retains the distinctive features of the class but also removes the common features of the class. These common features are not helpful for the classification task and may even cause confusion. Mathematically, the formula is shown in Equation 11:

[0097] f w =f p -f c ′ (11)

[0098] Among them, f p is the traditional characteristic obtained in Eq.6, f c ′ is a common feature, f w is the final purified feature vector. Finally, the purified feature vector fw is provided to the classifier Cp.

[0099] Y Re sult =soft max(W p *f w +b p ), (12)

[0100] Loss p =CrossEntropy(Y True , Y Re sult ), (13)

[0101] Where Wp, bp are the weight and bias of the classifier Cp. By optimizing Lossp, the feature extractor f p Features can be purified. Lossc and Lossp are trained simultaneously, but they use different optimizers. Lossc uses Moment SGD as the optimizer, while Lossp uses Adam optimizer. Although the two losses are opposite to the optimization goal, the present invention finds a balance so that the extracted features f w Closer to real common features.

[0102] The algorithm description of the entire training process is as follows:

[0103]

[0104] Setting of feature parameters: During the training phase of the GP-Net module, the momentum uses a stochastic gradient of 0.9, and the annealing learning rate can be calculated as follows:

[0105]

[0106] Where l0 = 0.01, α = 10, and β = 0.75 are the training progresses that change linearly from 0 to 1. In Equation 8, the value of λ is set to [0.05, 0.1, 0.2, 0.4, 0.8, 1.0].

[0107] It is well known that manual feature extraction methods based on machine learning do not generalize well because this method requires a lot of specialized knowledge and annotations to extract features. Due to the high cost of manual annotation, it is not feasible to obtain large annotated data sets for most clinical NLP tasks. Deep learning does not require annotation and can be completed automatically. The present invention combines a Transformer-based model with a feature purification network to greatly improve the classification performance. The present invention pre-trains the Transformer and fine-tunes the model on a new dataset so that the learned knowledge can be transferred to the text classification task of the present invention. The work of the present invention is significantly different from previous studies on AD identification through NLP technology because previous studies have not improved the representation learning of deep learning. Transformer is still the most widely used deep learning algorithm, but the time complexity of self-attention is O(n2), which hinders the development of the model, so it is of great significance to improve the efficiency of Transformer in the future.

[0108] Transformer, as the feature extractor used by the present invention in this study, can also be replaced by other deep learning algorithms, such as Bert, RNN, CNN, etc., and the present invention will further improve the work next. The present invention also believes that the feature purification method of the present invention can effectively classify other diseases related to language and cognitive disorders, such as Parkinson's disease, aphasia and autism spectrum disorder. Among them, aphasia may be more obvious because aphasia mainly affects language. The method of the present invention provides a feasible solution for treating AD patients at home. The feature purification method of deep learning is considered by the present invention to be a very promising exploration direction in the future.

[0109] Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

Claims

1. A classification method for Alzheimer's disease language detection based on a feature purification network, characterized by: It includes the following steps: Step 1: Obtain audio information for language detection and transcribe it into text data; Step 2: Pass the transcribed text data through the Transformer encoder of the GP-Net network to obtain the semantic feature vector f corresponding to the G-Net model c And the semantic feature vector f corresponding to the P-Net model p ; Among them, the GP-Net network includes two sub-networks, G-Net and P-Net. G-Net and P-Net have the same input text x. The two sub-networks G-Net and P-Net have the same structure but do not share parameters. The Transformer encoder includes a self-attention layer, a first normalization layer, a feedforward layer, and a second normalization layer arranged in sequence; The self-attention layer receives two input vectors x1 and x2 and generates corresponding vectors z1 and z2, which are passed to the first normalization layer respectively. The first normalization layer normalizes the two output vectors z1 and z2 of the self-attention layer and outputs them to a feedforward layer respectively. The normalized vectors z1 and z2 are output to the second normalization layer through the corresponding feedforward layer, and are normalized by the second normalization layer to obtain vectors r1 and r2; r1 and r2 replace x1 and x2 and are input into the encoder again, and the final semantic feature vector f is obtained after six cycles. c and f p ; Step 3, semantic feature vector f c After G-Net processing to extract common features f between datasets c ′, and then the common feature f c ′ is output to the classifier Cc of G-Net and outputs the final common feature f c "; Step 4: The P-Net model obtains the final common feature f output by the classifier Cc c ", and from the semantic feature vector f p The final common feature f is removed c "Get the purified feature vector f w , the purified feature vector f w The classification result Y after the purified features is obtained by the classifier Cp provided to P-Net Result .

2. The classification method for Alzheimer's disease language detection based on feature purification network according to claim 1 is characterized by: The maximum length of speech-to-text data does not exceed 500, so the length of word embeddings x1 and x2 is set to 500.

3. The classification method for Alzheimer's disease language detection based on feature purification network according to claim 1, characterized in that: The output function expression of the first normalization layer is as follows: z1=LN(x1+z1), (1) z2=LN(x2+z2) (3) Among them, x1 represents the output of the Embedding layer, and the x1 layer passes through the self-attention layer to obtain z1; the x1+z1 vector is normalized to obtain the new z1 layer; x2 represents the output of the Embedding layer, and the x2 layer passes through the self-attention layer to obtain z2; the x2+z2 vector is normalized to obtain the new z2 layer; LN means layer normalization; Furthermore, the output function expression of the second normalization layer is as follows: r1=LN(z1+FeedForward(z1)) (2) r2=LN(z2+FeedForward(z2)) (4) Among them, z1 passes through the feedforward neural network layer, passes through the residual structure and is added to itself, and also passes through the LN layer to finally output the vector r1; z2 passes through the feedforward neural network layer, passes through the residual structure and is added to itself, and also passes through the LN layer to finally output the vector r2.

4. The classification method for Alzheimer's disease language detection based on feature purification network according to claim 1, characterized in that: The semantic feature vector f in step 3 c GRL processing is used to train and obtain common features close to the real ones; GRL acts as an identity transform during forward propagation, and then takes the gradient from the next layer and changes its value during backward propagation, that is, multiplying it by λ; then passes it to the next layer; specifically, the function expression processed by GRL is as follows: G λ (x)=x′ (7) Among them, λ is a hyperparameter, x is the input vector of GRL, that is, the semantic feature vector f c , to get G λ (f c )=f c ′.

5. The classification method for Alzheimer's disease language detection based on feature purification network according to claim 1, characterized in that: The function expression of the classifier Cc in step 3 is: Y GRL =soft max(W c *f c ′+b c ) (9) The corresponding loss function is: Loss c =CrossEntropy(Y True ,Y GRL ) (10) Among them, Y GRL is the output vector of the G-Net neural network, which is the f that is finally provided to P-net c ′; Wc, bc are the weight and bias of classifier Cc, f c ′ is the extracted common feature; Y True is an ideal vector; common features of different classes are extracted by optimizing the Lossc feature extractor Fc.

6. The classification method for Alzheimer's disease language detection based on feature purification network according to claim 1, characterized in that: The function expression of the classifier Cp in step 4 is: Y Re sult =soft max(W p *f w +b p ) (12) The corresponding loss function is: Loss p =CrossEntropy(Y True ,Y Re sult ) (13) Among them, Y Result is the output vector of the P-Net neural network, W p 、b p For the classifier C p The weights and biases of w is the final purified feature vector, CrossEntropy is the cross entropy loss function, Loss p Measure a neural network output vector Y Result and the ideal vector Y True The degree of closeness; by optimizing Loss p Feature Extractor F p Extract common features of different classes.

7. A classification system for Alzheimer's disease language detection based on a feature purification network, applying the classification method for Alzheimer's disease language detection based on a feature purification network according to any one of claims 1 to 6, characterized in that: The GP-Net network includes two sub-networks, G-Net and P-Net. G-Net and P-Net have the same input text x. The two sub-networks G-Net and P-Net have the same structure but do not share parameters. G-Net includes an input layer X, a feature extractor Fc, a gradient reversal layer, and a classification layer Cc. The feature extractor Fc extracts common features, and the classification layer Cc classifies the common features to obtain the classification results of the common features and provides them to P-Net. P-Net includes an input layer X, a feature extractor Fp, and a classification layer Cp. The feature vector extracted by the feature extractor Fp removes the common features extracted from the feature extractor Fc to obtain purified features. The classification layer Cp classifies the purified features to obtain the purified feature classification results.

8. The classification system for Alzheimer's disease language detection based on feature purification network according to claim 7, characterized in that: The feature extractor Fp and the feature extractor Fc together form the Transformer encoder of the GP-Net network; the Transformer encoder includes a self-attention layer, a first normalization layer, a feedforward layer, and a second normalization layer arranged in sequence; the self-attention layer receives two input vectors x1 and x2 and obtains corresponding vectors z1 and z2, which are passed to the first normalization layer respectively; the first normalization layer normalizes the two output vectors z1 and z2 of the self-attention layer and outputs them to a feedforward layer respectively; The normalized vectors z1 and z2 are respectively output to the second normalization layer through the feedforward layer, and are normalized by the second normalization layer to obtain vectors r1 and r2. r1 and r2 replace x1 and x2 and are input into the encoder again. The final semantic feature vector f is obtained after six cycles. c and f p .

Citation Information

Patent Citations

  • Early Alzheimer's disease diagnosis system based on artificial intelligence

    CN109285604A

  • Classification model training method, Alzheimer disease classification method and device

    CN113392938A