Cerebral stroke NIHSS scoring method and system based on improved BERT model and feature fusion

By improving the BERT model and feature fusion method, a stroke NIHSS scoring model was constructed, which solved the problem of time-consuming and inaccurate existing scoring methods, and realized automated and accurate NIHSS scoring, assisting clinicians to quickly identify the severity of stroke and ensure timely treatment of patients.

CN120260913APending Publication Date: 2025-07-04DALIAN NEUSOFT UNIV OF INFORMATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510334663.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing NIHSS scoring methods are time-consuming and inaccurate, causing stroke patients to miss the optimal treatment time.

Method used

The improved BERT model and feature fusion method are used to construct a stroke NIHSS scoring model, including input embedding module, position coding module, Transformer encoder, bidirectional GRU module, feature fusion module and output layer. Feature fusion is performed through self-attention mechanism and ridge regression model to improve the accuracy of the score.

Benefits of technology

It realizes automatic, efficient and accurate NIHSS scores, assisting clinicians to quickly identify stroke severity and key pathological information, ensure that patients receive timely intervention, and improve prognosis and treatment effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260913A_ABST
    Figure CN120260913A_ABST
Patent Text Reader

Abstract

The invention discloses a cerebral apoplexy NIHSS scoring method and system based on an improved BERT model and feature fusion. The method comprises the steps that S1, a cerebral apoplexy NIHSS scoring data set is acquired; s2, a stroke NIHSS scoring model is obtained through construction, the stroke NIHSS scoring model is trained, and the trained stroke NIHSS scoring model is obtained; the cerebral apoplexy NIHSS scoring model comprises an input embedding module, a position coding module, a Transform encoder, a bidirectional GRU module, a feature fusion module and an output layer; and S3, carrying out cerebral apoplexy NIHSS scoring based on the trained cerebral apoplexy NIHSS scoring model. According to the method, the existing BERT model is improved to construct the cerebral apoplexy NIHSS scoring model, the cerebral apoplexy NIHSS scoring model is utilized to automatically, efficiently and accurately process the text data (such as chief complaint, medical history and the like) of the patient, the efficiency and effect of analyzing and processing the cerebral apoplexy NIHSS scoring data are improved, namely, the prediction accuracy of the NIHSS score of the patient is improved, and the accuracy of predicting the NIHSS score of the patient is improved. By optimizing the NIHSS score data analysis process, a clinician is assisted in quickly identifying the severity of the cerebral apoplexy and key pathological information, it is ensured that a patient is intervened in time, and prognosis and treatment effects are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of NIHSS scoring for stroke, and in particular, to a method and system for NIHSS scoring of stroke based on an improved BERT model and feature fusion. Background Art

[0002] Stroke, also known as cerebrovascular accident or stroke, is a common clinical disease in which blood circulation in the brain is impaired due to sudden rupture or blockage of blood vessels in the brain, leading to brain tissue damage. It is characterized by high incidence, high disability rate and high mortality, and is one of the main causes of long-term disability and death worldwide. The clinical manifestations of stroke are diverse, including disturbance of consciousness, slurred speech, limb weakness, etc., which seriously affect the quality of life and social function of patients. In addition to the above characteristics, the condition of stroke develops rapidly. In acute stroke events, quickly and accurately assessing the degree of neurological deficit in patients is of great significance for guiding emergency treatment decisions and improving the prognosis of patients.

[0003] In order to accurately evaluate the degree of neurological deficit in stroke patients and guide clinical treatment and prognosis judgment, NIHSS is widely used in clinical practice. NIHSS is a standardized neurological examination tool that quantifies the neurological status of patients through a series of item scores, including levels of consciousness, eye movement, visual field, facial movement, limb movement, sensation, language and articulation. It can not only help doctors quickly judge the severity of stroke in patients, but also serve as an important indicator for evaluating treatment effects and prognosis. However, the existing NIHSS scoring methods mainly rely on a large amount of experience of doctors for manual evaluation, which is not only time-consuming but also may have subjective differences. At the same time, the detailed medical history and clinical manifestation data of patients cannot be fully utilized, resulting in inaccurate scoring and causing patients to miss the best treatment time. Summary of the Invention

[0004] The present invention provides a method and system for NIHSS scoring of stroke based on an improved BERT model and feature fusion to overcome the technical problems of the existing NIHSS scoring methods being time-consuming, inaccurate and likely to cause patients to miss the best treatment time.

[0005] To achieve the above object, the technical solution of the present invention is:

[0006] A method for NIHSS scoring of stroke based on an improved BERT model and feature fusion, the specific steps include:

[0007] S1: Obtain a stroke NIHSS scoring data set, where the stroke NIHSS scoring data set includes text data, and the text data includes patient complaints and feature information;

[0008] S2: Improve the BERT model to construct a stroke NIHSS scoring model, and train the stroke NIHSS scoring model based on the stroke NIHSS scoring dataset to obtain a trained stroke NIHSS scoring model;

[0009] The stroke NIHSS scoring model includes an input embedding module, a position encoding module, a Transformer encoder, a bidirectional GRU module, a feature fusion module, and an output layer;

[0010] The input embedding module is used to convert the text data into word embedding vectors and transmit them to the position embedding module;

[0011] The position encoding module is used to capture the sequential information and relative position relationship between the word embedding vectors to obtain position encoding vectors and transmit them to the Transformer encoder;

[0012] The Transformer encoder is used to extract semantic and position information from the position encoding vectors based on the self-attention mechanism to obtain sequence data containing semantic and position information and transmit it to the bidirectional GRU module;

[0013] The bidirectional GRU module is used to capture bidirectional information in the sequence data and transmit it to the feature fusion module;

[0014] The feature fusion module is used to perform feature fusion on the bidirectional information to obtain fused feature vectors and transmit them to the output layer;

[0015] The output layer is used to output the NIHSS score based on the fused feature vectors;

[0016] S3: Perform stroke NIHSS scoring based on the trained stroke NIHSS scoring model.

[0017] Further, the position encoding module captures the sequential information and relative position relationship between the vectors, and obtaining the position encoding vectors includes:

[0018] The position encoding module includes absolute position encoding and relative position encoding;

[0019] The absolute position encoding is used to encode the word embedding vectors to capture the sequential information between the vectors to obtain absolute position encoding vectors;

[0020] The relative position encoding is used to encode the word embedding vectors to capture the relative position relationship between the vectors to obtain relative position encoding vectors;

[0021] Add the absolute position encoding vector and the relative position encoding vector to obtain a position encoding vector.

[0022] Further, the input embedding module converts the text data into word embedding vectors, including:

[0023] Use jieba to segment the text data and count the word frequencies to construct a Chinese vocabulary;

[0024] Adjust the dimension of the embedding matrix according to the size of the Chinese vocabulary;

[0025] Map each word in the Chinese vocabulary to the corresponding embedding vector in the embedding matrix to obtain the word embedding vector. The mapping expression is:

[0026]

[0027] where, e im is the embedding vector of the i-th word in the text data; e' wim is the embedding vector in the embedding matrix; Vocab is the Chinese vocabulary; e'{[UNK]} represents the set unknown word embedding vector.

[0028] Further, the Transformer encoder extracts semantic and position information in the position encoding vector based on the self-attention mechanism, including:

[0029] The attention weight calculation formula in the self-attention mechanism is:

[0030]

[0031] where, B is the bias matrix; d k is the dimension of the key vector; Q is the query vector; K is the key vector; V is the value vector.

[0032] Further, the feature fusion module includes a pooling layer, a ridge regression model, and a splicing layer;

[0033] The pooling layer is used to perform a pooling operation on the bidirectional information to obtain a pooled vector and transmit it to the ridge regression model;

[0034] The ridge regression model is used to make predictions based on the pooled vector to obtain the NIHSS score of each sample in the text data and transmit it to the splicing layer;

[0035] The splicing layer is used to splice and fuse the output of the ridge regression model and the pooled vector to obtain a fused feature vector.

[0036] Further, during the training process of the improved BERT model based on the stroke NIHSS score dataset, the loss function used is the mean squared error loss function, expressed as:

[0037]

[0038] where N is the number of samples, and y i are the predicted score and the actual score of the i-th sample respectively.

[0039] A stroke NIHSS scoring system based on an improved BERT model and feature fusion, used to implement a stroke NIHSS scoring method based on an improved BERT model and feature fusion, includes: a data acquisition module, used to obtain a stroke NIHSS score dataset, where the stroke NIHSS score dataset includes text data, and the text data includes patient chief complaints and feature information;

[0040] A stroke NIHSS scoring model establishment module, used to improve the BERT model to construct a stroke NIHSS scoring model, and train it based on the stroke NIHSS score dataset to obtain a trained stroke NIHSS scoring model;

[0041] The stroke NIHSS scoring model includes an input embedding module, a position encoding module, a Transformer encoder, a bidirectional GRU module, a feature fusion module, and an output layer;

[0042] The input embedding module is used to convert the text data into word embedding vectors and transmit them to the position embedding module;

[0043] The position encoding module is used to capture the sequential information and relative position relationship between the word embedding vectors to obtain position encoding vectors and transmit them to the Transformer encoder;

[0044] The Transformer encoder is used to extract semantic and position information in the position encoding vectors based on the self-attention mechanism to obtain sequence data containing semantic and position information and transmit it to the bidirectional GRU module;

[0045] The bidirectional GRU module is used to capture bidirectional information in the sequence data and transmit it to the feature fusion module;

[0046] The feature fusion module is used to perform feature fusion on the bidirectional information to obtain a fused feature vector and transmit it to the output layer;

[0047] The output layer is used to output the NIHSS score based on the fused feature vector.

[0048] Beneficial effects: The present invention improves the existing BERT model to construct a stroke NIHSS scoring model. The stroke NIHSS scoring model includes an input embedding module, a position encoding module, a Transformer encoder, a bidirectional GRU module, a feature fusion module, and an output layer. Using the stroke NIHSS scoring model, it is possible to automatically, efficiently, and accurately process the text data of patients (such as chief complaints, medical histories, etc.). By means of the input embedding module, the position encoding module, the Transformer encoder, and the bidirectional GRU module, the ability to process Chinese texts is improved, more position information is captured, and the ability to capture long-term dependencies is enhanced, generating an NIHSS score highly consistent with the evaluation of clinicians. Ultimately, the efficiency and effect of data analysis and processing of the stroke NIHSS score are improved, that is, the prediction accuracy of the NIHSS score of patients is improved. The present invention optimizes the data analysis process of the NIHSS score to assist clinicians in quickly identifying the severity of stroke and key pathological information, so as to make timely treatment decisions subsequently, ensure that patients receive timely intervention, and guarantee the prognosis and treatment effect. Brief Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flowchart of a stroke NIHSS scoring method based on an improved BERT model and feature fusion in the present invention. Detailed Embodiments

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0052] This embodiment provides a stroke NIHSS scoring method based on an improved BERT model and feature fusion, as Figure 1 shown. The specific steps include:

[0053] S1: Acquire a stroke NIHSS scoring dataset, wherein the stroke NIHSS scoring dataset includes text data, and the text data includes patient complaints and feature information; specifically, the feature information includes medical history, life characteristics, etc.

[0054] S2: improving the BERT model to construct a stroke NIHSS scoring model, training the stroke NIHSS scoring model based on the stroke NIHSS scoring dataset to obtain a trained stroke NIHSS scoring model;

[0055] The stroke NIHSS scoring model includes an input embedding module, a position encoding module, a Transformer encoder, a bidirectional GRU (Gated Recurrent Unit) module, a feature fusion module and an output layer;

[0056] Specifically, the existing BERT model is built based on the Transformer encoder, and its structure includes an input embedding module, a position encoding module, a Transformer encoder, and an output layer. In order to improve the accuracy of NIHSS scoring of stroke patients, this embodiment improves the existing BERT model, including optimizing the existing input embedding module, position encoding module, and self-attention mechanism in the Transformer encoder, and adding a bidirectional GRU module and a feature fusion module.

[0057] The input embedding module is used to convert the text data into word embedding vectors and transmit them to the position embedding module;

[0058] Specifically, in this embodiment, the jieba word segmentation is used to replace the original WordPiece word segmenter. The jieba word segmentation is a commonly used word segmentation tool in Chinese natural language processing. It is based on statistical and rule-based methods and can handle the Chinese vocabulary segmentation problem well, thereby constructing a Chinese vocabulary and adjusting the embedding matrix to adapt to the new Chinese vocabulary. The words in the input text are mapped to the embedding matrix to obtain the word embedding vector.

[0059] The input embedding module converts the text data into a word embedding vector including:

[0060] (1) Using Jieba word segmentation to segment the text data and count the word frequencies to construct a Chinese vocabulary;

[0061] Specifically, let the text data be P. For each sentence s in P, i , use jieba word segmentation to segment, and get the word sequence w i1 ,w i2 ,...,w in, count the frequency \(f(w)\) of the word \(w\) in the word sequence, where \(m = 1, 2, \ldots, n\), in the text data. That is the word frequency. According to the word frequency and actual requirements, construct the Chinese stroke chief complaint vocabulary, namely the Chinese vocabulary Vocab, \(Vocab = \{w | f(w), Vocab > V'\}\), where \(V'\) is the set threshold for the Chinese vocabulary. im , \(m = 1, 2, \ldots, n\), the frequency \(f(w\) im ) of the word in the text data, i.e., the word frequency. According to the word frequency and actual requirements, construct the Chinese stroke chief complaint vocabulary, namely the Chinese vocabulary Vocab, \(Vocab = w\) im | \(f(w\) im ) where \(Vocab > V'\), and \(V'\) is the set threshold for the Chinese vocabulary.

[0062] (2) When applying jieba word segmentation to the improved BERT model, the embedding matrix needs to be adjusted accordingly. Therefore, adjust the dimension of the embedding matrix according to the size of the Chinese vocabulary.

[0063] Specifically, let the new embedding matrix obtained after adjusting the size of the embedding matrix be \(E'\), \(E' \in R\) |Vocab|×d , where \(|Vocab|\) is the size of the Chinese vocabulary; \(d\) is the dimension of the embedding vector. For the word \(w\) in the Chinese vocabulary im , assign an embedding vector \(e'\) wim to it through random initialization or pre-training, \(e'\) wim \(\in R\) d ;

[0064] Specifically, during the pre-training process, these embedding vectors will be updated as model parameters.

[0065] (3) Map the word \(w\) im to the corresponding embedding vector \(e'\) wim in the embedding matrix \(E'\) to obtain the embedding vector of the word. The mapping expression is:

[0066]

[0067] where \(e\) im is the embedding vector of the \(i\)-th word in the text data; \(Vocab\) is the Chinese vocabulary; \(e'_{[UNK]}\) represents the set unknown word embedding vector. If a word is not in the Chinese vocabulary (i.e., an OOV word), it can be processed with a random vector or a zero vector.

[0068] The position encoding module is used to capture the sequential information and relative position relationship between the embedding vectors of the words, obtain the position encoding vector, and transmit it to the Transformer encoder;

[0069] In a specific embodiment, the position encoding module captures the sequential information and relative position relationship between the vectors to obtain the position encoding vector, including:

[0070] The position encoding module includes absolute position encoding and relative position encoding;

[0071] The absolute position encoding is used to encode the embedding vectors of the words to capture the sequential information between the vectors, obtaining absolute position encoding vectors;

[0072] Specifically, in this embodiment, to further enhance the position encoding mechanism of the existing BERT model, a method of combining absolute position encoding and relative position encoding is adopted. The absolute position encoding provides a unique encoding for each position, which helps the model understand the sequential information in the sequence. In addition to the absolute position encoding, in order to enable the improved BERT model to capture the relative position relationship between words, relative position encoding is introduced in this embodiment to represent the relative distance between different positions. The relative position encoding is used to encode the embedding vectors of the words to capture the relative position relationship between the vectors, obtaining relative position encoding vectors;

[0073] Specifically, the relative position encoding is implemented through a learnable relative position embedding matrix R i’j’ R i’j’ represents the relative position embedding between position i' and position j'. To reduce the number of parameters of the improved BERT model, a parameter sharing strategy is adopted for the relative position embedding matrix. The parameter sharing strategy includes: First, considering the maximum interval between relative positions, a maximum relative distance D is determined. The relative position embedding matrix contains relative position embeddings from -D to D. Then, a random initialization method is used to initialize each element in R i’j’ in it.

[0074] Adding the absolute position encoding vector and the relative position encoding vector together, the position encoding vector is obtained, expressed as:

[0075] X = E + P abs + P rel (2)

[0076] where E is the embedding vector of the input word, P abs is the absolute position encoding vector, P rel is the relative position encoding vector, and X is the output position encoding vector.

[0077] The Transformer encoder is used to extract the semantic and position information in the position encoding vector based on the self-attention mechanism, obtaining sequence data containing semantic and position information, and transmitting it to the bidirectional GRU module;

[0078] Specifically, the Transformer encoder extracts features through the Self-Attention Mechanism and the Feed-Forward Neural Network, enabling the improved BERT model to focus on different positions in the sequence.

[0079] Specifically, the Transformer encoder extracts semantic and positional information from the positional encoding vectors based on the self-attention mechanism, including:

[0080] The self-attention mechanism is a core component of the BERT model, which allows the model to focus on information at different positions when processing the sequence. Since the positional encoding module uses relative positional encoding, in order to consider both absolute and relative positional encoding information when calculating self-attention, this embodiment modifies the original attention calculation method, including:

[0081] In the existing BERT model, the self-attention weights are calculated through linear transformations of queries (Q), keys (K), and values (V). The formula for calculating the self-attention weights is:

[0082]

[0083] where d k is the dimension of the key vector, which is used to scale the dot-product attention to avoid gradient vanishing or explosion.

[0084] After introducing relative positional encoding information, in order to consider relative positional information when calculating attention, it is necessary to add the relative positional encoding matrix to the dot product of the query and the key, and the formula for calculating the attention weights is:

[0085]

[0086] where Q pos is the absolute positional encoding corresponding to the query vector, and R is the relative positional encoding matrix.

[0087] In this way, the model can consider both the semantic relationship and the relative positional relationship between words when calculating the attention weights. However, formula (4) may increase the computational complexity. To more efficiently implement relative positional encoding, this embodiment adopts a more concise method, that is, directly introducing an additional bias term B when calculating the attention weights. This bias term is generated by the relative positional encoding, that is, the formula for calculating the attention weights in the self-attention mechanism is:

[0088]

[0089] where B is the bias matrix, and its elements are generated by the relative positional encoding.

[0090] Specifically, the relative position embedding matrix R i’j’ is converted into a bias vector b ij , and then all bias vectors are combined into a bias matrix B; Q is the query vector; K is the key vector; V is the value vector.

[0091] Specifically, the attention weight calculation formula (5) can introduce relative position information without significantly increasing the computational complexity. At the same time, since the bias matrix B is directly added to the original attention weights inside the softmax function, it can more directly affect the attention distribution of the model.

[0092] Specifically, the features output by the Transformer encoder already contain rich semantic and position information. Let the output of the Transformer encoder be H = h1, h2,..., h T , where T is the sequence length, and h t is the vector representation of the t-th word.

[0093] In this embodiment, to further improve the performance of the BERT model, a bidirectional GRU module is added. The bidirectional GRU module is used to capture bidirectional information in the sequence data, that is, the long-term dependencies in the sequence data, and transmit them to the feature fusion module; specifically, the long-term dependencies in the sequence are captured by further processing through the bidirectional GRU module.

[0094] Specifically, the bidirectional GRU module capturing bidirectional information in the sequence data includes:

[0095] Specifically, the bidirectional GRU module includes a forward GRU and a backward GRU. The forward GRU processes the information from the start to the end of the sequence, and the backward GRU processes the information from the end to the start of the sequence; the output of the bidirectional GRU module is the concatenation of the forward and backward representations at each time step (i.e., each word), expressed as:

[0096]

[0097] In the formula, and correspond to the outputs of the forward and backward GRUs respectively.

[0098] The feature fusion module is used to perform feature fusion based on the bidirectional information to obtain a fused feature vector and transmit it to the output layer;

[0099] Specifically, the feature fusion module includes a pooling layer, a ridge regression model, and a concatenation layer;

[0100] The pooling layer is used to perform a pooling operation on the bidirectional information to obtain a pooled vector and transmit it to the ridge regression model;

[0101] Specifically, the pooling operation can be average pooling or max pooling to obtain a fixed-length vector representation v, where v = pool(G);

[0102] where G is all the outputs of the bidirectional GRU module, G = {g1, g2,..., g T}.

[0103] The ridge regression model is used to make predictions based on the pooled vector to obtain the NIHSS score for each sample in the text data and transmit it to the concatenation layer;

[0104] Specifically, let the prediction result of the ridge regression model be R, R = {r1, r2,..., r N},

[0105] where N is the number of samples, and r i is the predicted value of the i-th sample.

[0106] Specifically, ridge regression is an existing improved linear regression model. In this embodiment, the ridge regression model is used to make predictions based on the pooled vector, and the specific process will not be elaborated here.

[0107] The concatenation layer is used to concatenate and fuse the output of the ridge regression model and the pooled vector to obtain a fused feature vector F, F = {f1, f2,..., f N}, where f i = [v i ; r i is the fused feature vector of the i-th sample.

[0108] The output layer is used to output the NIHSS score based on the fused feature vector;

[0109] Specifically, the output layer is a fully connected layer. To enable the model to predict continuous score values, the fused feature vector F is input into the fully connected layer. The fully connected layer maps the fused feature vector F to a continuous NIHSS score. The output of the fully connected layer is expressed as:

[0110]

[0111] where W f and b f are the weights and bias terms of the fully connected layer respectively, is the finally predicted NIHSS score value,

[0112] S3: Perform the NIHSS score for stroke based on the trained NIHSS score model for stroke.

[0113] In the stroke task, the trained BERT model can deeply understand the semantics of the patient's chief complaint and perform excellently in the task. The BERT model can also adapt to specific needs by changing the size of the output layer, demonstrating strong flexibility and scalability.

[0114] In this embodiment, during the process of training the improved BERT model, the mean squared error (MSE) is used as the loss function, which measures the difference between the NIHSS score predicted by the model and the actual score. The mean squared error loss function can be defined as:

[0115]

[0116] where N is the number of samples, and y i are the predicted score and the actual score of the i-th sample, respectively.

[0117] Specifically, optimization algorithms such as gradient descent are used to minimize the above loss function. When the loss function converges, the trained BERT model is obtained.

[0118] This embodiment also provides a NIHSS score system for stroke based on the improved BERT model and feature fusion, which is used to implement the NIHSS score method for stroke, including: a data acquisition module, which is used to obtain the NIHSS score dataset for stroke, and the NIHSS score dataset for stroke includes text data, and the text data includes the patient's chief complaint and feature information;

[0119] a NIHSS score model establishment module, which is used to improve the BERT model to construct a NIHSS score model for stroke, and perform training based on the NIHSS score dataset for stroke to obtain the trained NIHSS score model for stroke;

[0120] The NIHSS score model for stroke includes an input embedding module, a position encoding module, a Transformer encoder, a bidirectional GRU module, a feature fusion module, and an output layer;

[0121] The input embedding module is used to convert the text data into word embedding vectors and transmit them to the position embedding module;

[0122] The position encoding module is used to capture the sequential information and relative position relationship between the word embedding vectors to obtain position encoding vectors and transmit them to the Transformer encoder;

[0123] The Transformer encoder is used to extract semantic and positional information in the positional encoding vector based on the self-attention mechanism, obtain sequence data containing semantic and positional information, and transmit it to the bidirectional GRU module;

[0124] The bidirectional GRU module is used to capture bidirectional information in the sequence data and transmit it to the feature fusion module;

[0125] The feature fusion module is used to perform feature fusion on the bidirectional information, obtain a fused feature vector, and transmit it to the output layer;

[0126] The output layer is used to output the NIHSS score based on the fused feature vector.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for NIHSS scoring of stroke based on an improved BERT model and feature fusion, characterized in that, The specific steps include: S1: Obtain a stroke NIHSS score dataset, where the stroke NIHSS score dataset includes text data, and the text data includes patient chief complaints and feature information; S2: Improve the BERT model to construct a stroke NIHSS score model, and train the stroke NIHSS score model based on the stroke NIHSS score dataset to obtain a trained stroke NIHSS score model; The stroke NIHSS score model includes an input embedding module, a position encoding module, a Transformer encoder, a bidirectional GRU module, a feature fusion module, and an output layer; The input embedding module is used to convert the text data into word embedding vectors and transmit them to the position embedding module; The position encoding module is used to capture the sequential information and relative position relationship between the word embedding vectors to obtain position encoding vectors and transmit them to the Transformer encoder; The Transformer encoder is used to extract semantic and position information in the position encoding vectors based on the self-attention mechanism to obtain sequence data containing semantic and position information and transmit it to the bidirectional GRU module; The bidirectional GRU module is used to capture bidirectional information in the sequence data and transmit it to the feature fusion module; The feature fusion module is used to perform feature fusion on the bidirectional information to obtain fused feature vectors and transmit them to the output layer; The output layer is used to output the NIHSS score based on the fused feature vectors; S3: Perform stroke NIHSS scoring based on the trained stroke NIHSS score model.

2. The stroke NIHSS scoring method based on the improved BERT model and feature fusion according to claim 1, wherein The position encoding module captures the sequential information and relative position relationship between the vectors to obtain position encoding vectors, including: The position encoding module includes absolute position encoding and relative position encoding; The absolute position encoding is used to encode the word embedding vectors to capture the sequential information between the vectors to obtain absolute position encoding vectors; The relative position encoding is used to encode the word embedding vectors to capture the relative position relationship between the vectors to obtain relative position encoding vectors; Add the absolute position encoding vectors and the relative position encoding vectors to obtain position encoding vectors.

3. The stroke NIHSS scoring method based on the improved BERT model and feature fusion according to claim 1, wherein, The input embedding module converts the text data into word embedding vectors, including: Use jieba word segmentation to segment the text data and count the word frequencies to construct a Chinese vocabulary; Adjust the dimension of the embedding matrix according to the size of the Chinese vocabulary; Map each word in the Chinese vocabulary to the corresponding embedding vector in the embedding matrix to obtain word embedding vectors. The mapping expression is: where, e im is the embedding vector of the i-th word in the text data; e' wim is the embedding vector in the embedding matrix; Vocab is the Chinese vocabulary; e'{[UNK]} represents the set unknown word embedding vector.

4. The stroke NIHSS scoring method based on the improved BERT model and feature fusion according to claim 1, wherein, The Transformer encoder extracts semantic and position information in the position encoding vectors based on the self-attention mechanism, including: The attention weight calculation formula in the self-attention mechanism is: where B is a bias matrix; d k is the dimension of the key vector; Q is the query vector; K is the key vector; and V is the value vector.

5. The stroke NIHSS scoring method based on the improved BERT model and feature fusion according to claim 1, wherein The feature fusion module includes a pooling layer, a ridge regression model, and a splicing layer; The pooling layer is used to perform a pooling operation on the bidirectional information to obtain a pooled vector and transmit it to the ridge regression model; The ridge regression model is used to make predictions based on the pooled vector to obtain the NIHSS score of each sample in the text data and transmit it to the splicing layer; The splicing layer is used to splice and fuse the output of the ridge regression model and the pooled vector to obtain a fused feature vector.

6. The NIHSS scoring method for stroke based on the improved BERT model and feature fusion according to claim 1, wherein During the training process of the improved BERT model based on the stroke NIHSS score dataset, the loss function used is the mean squared error loss function, which is expressed as: where N is the number of samples, and yi are the predicted score and the actual score of the i-th sample respectively.

7. A stroke NIHSS scoring system based on an improved BERT model and feature fusion, for implementing the method described in claim 1, characterized in that, Including: A data acquisition module for obtaining a stroke NIHSS score dataset, where the stroke NIHSS score dataset includes text data, and the text data includes patient complaints and feature information; A stroke NIHSS score model establishment module for improving the BERT model to construct a stroke NIHSS score model and training it based on the stroke NIHSS score dataset to obtain a trained stroke NIHSS score model; The stroke NIHSS score model includes an input embedding module, a position encoding module, a Transformer encoder, a bidirectional GRU module, a feature fusion module, and an output layer; The input embedding module is used to convert the text data into word embedding vectors and transmit them to the position embedding module; The position encoding module is used to capture the sequential information and relative position relationship between the word embedding vectors to obtain position encoding vectors and transmit them to the Transformer encoder; The Transformer encoder is used to extract semantic and position information in the position encoding vectors based on the self-attention mechanism to obtain sequence data containing semantic and position information and transmit it to the bidirectional GRU module; The bidirectional GRU module is used to capture bidirectional information in the sequence data and transmit it to the feature fusion module; The feature fusion module is used to perform feature fusion on the bidirectional information to obtain a fused feature vector and transmit it to the output layer; The output layer is used to output the NIHSS score based on the fused feature vector.