Multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis

Through the dynamic interactive fusion method of multimodal confidence, the problem of insufficient multimodal information integration in stroke rehabilitation diagnosis is solved, the diagnostic accuracy and resource utilization efficiency are improved, and personalized rehabilitation assessment is provided.

CN120473125BActive Publication Date: 2025-09-19QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510968897.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-19
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing stroke rehabilitation diagnosis methods mostly rely on single-modality data and are unable to fully integrate multimodal information, resulting in limited evaluation accuracy. They also fail to effectively capture the high-level correlations between different modalities in the confidence space, leading to confidence conflicts between modalities and affecting diagnostic effectiveness.

Method used

By constructing a multimodal confidence dynamic interactive fusion method, including a multimodal confidence feature joint module and a multimodal confidence collaborative correction module, and utilizing the confidence feature correction matrix and the confidence feature association matrix, multimodal information is dynamically interactively fused to resolve confidence conflicts and improve diagnostic accuracy.

Benefits of technology

It achieves more efficient and accurate stroke rehabilitation diagnosis, conforms to the actual diagnostic process of professional doctors, alleviates the shortage of medical resources, and provides more intelligent and personalized rehabilitation assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120473125B_ABST
    Figure CN120473125B_ABST
Patent Text Reader

Abstract

The present invention relates to a multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis, belonging to the technical field of stroke rehabilitation intelligent diagnosis. The method comprises the following steps: obtaining multimodal data of stroke patient rehabilitation training and constructing a data set, inputting the data into a confidence generation network after preprocessing to obtain different modal confidence features; inputting the features into a multimodal confidence feature joint module to construct a full feature matrix, calculating the Euclidean distance to add spatial information, then extracting cross-modal correlation features through two layers of convolution, and outputting the final full feature matrix; at the same time, in a multimodal confidence collaborative correction module, combining the correct label confidence to calculate the correction matrix, and using cosine similarity, linear projection features, etc. to obtain the final correction matrix; after extending and splicing the outputs of the two modules, inputting them into a fully connected network to obtain the diagnosis result, and calculating the loss through a loss function and the true result, and performing training optimization. The present invention can improve the accuracy of stroke rehabilitation diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent diagnosis of stroke rehabilitation, and particularly relates to a multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis. Background Art

[0002] As the second leading cause of death and leading cause of disability worldwide, stroke, characterized by high morbidity, disability, mortality, and recurrence rates, imposes a heavy burden on society. Up to 70% of survivors suffer from severe functional impairments such as hemiplegia and aphasia, making scientifically accurate rehabilitation assessment and training central to patients' functional recovery. However, traditional rehabilitation assessments rely heavily on subjective judgments by clinicians (e.g., muscle strength and tone), resulting in inconsistent standards and inefficiencies. The aging population and younger patient population are further exacerbating the strain on medical resources, necessitating an urgent need for efficient and objective tools. The rise of artificial intelligence technologies (e.g., machine learning and CNNs) has opened up the possibility of intelligent stroke diagnosis. By leveraging pathology databases to aid assessment, it is expected to alleviate the burden on healthcare. However, existing methods often rely on single-modality data (e.g., imaging or scales alone), failing to fully integrate the multimodal information essential for clinical diagnosis, including speech, motor performance, and electronic medical records. This makes it difficult to capture the complex and comprehensive picture of a patient's recovery status, limiting assessment accuracy.

[0003] In recent years, multimodal fusion technology has attracted considerable attention due to its ability to integrate complementary information such as text, images, and audio, making it more suitable for comprehensive clinical decision-making. However, in the specific context of stroke rehabilitation diagnosis, multimodal learning faces significant challenges: First, there is a lack of clearly annotated multimodal clinical datasets covering the entire patient rehabilitation phase; second, existing fusion methods fail to effectively capture the high-level correlations between different modalities in the confidence space. Multimodal data from stroke patients (such as ambiguous speech, abnormal motor pattern images, and contradictory symptom descriptions) is complex and prone to intermodal confidence conflicts—that is, significant differences or even contradictions in the confidence of different modalities regarding the same rehabilitation status. Previous methods have inadequately addressed these conflicts, failing to fully explore and utilize the deep correlations between cross-modal confidence features, resulting in suboptimal fusion diagnostic results. Therefore, designing a new method that can dynamically and interactively fuse multimodal information, effectively correlate high-level cross-modal confidence features, and resolve conflicts using a confidence feature correction matrix and a confidence feature correlation matrix to improve the accuracy of stroke rehabilitation diagnosis has become a key technical challenge that urgently needs to be overcome. Summary of the Invention

[0004] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions:

[0005] The present invention provides a multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis, comprising the following steps:

[0006] S1. Acquire multimodal data from stroke patients during rehabilitation training to construct a stroke dataset; preprocess the data in the stroke dataset; input the preprocessed data into the confidence generation network of the corresponding modality to obtain confidence features of different modalities;

[0007] S2. Input the confidence features of different modalities into the multimodal confidence feature joint module, construct a confidence full feature matrix using the maximum and minimum confidence of each category, calculate the Euclidean distance between all equally likely confidence features in the confidence full feature matrix, add high-level spatial information to the confidence full feature matrix, extract cross-modal confidence correlation features through two layers of convolution, and output the final confidence full feature matrix.

[0008] S3. Input the confidence features of different modalities into the multimodal confidence collaborative correction module. Calculate the confidence of the correct label to obtain the confidence feature correction matrix. Then, based on the cosine similarity feature and supplemented by the linear projection feature, the cross-attention and convolution modules are combined to calculate the final confidence feature correction matrix.

[0009] S4. Performing a tensor extension operation on the final confidence full feature matrix output by the multimodal confidence feature joint module and the final confidence feature correction matrix output by the multimodal confidence collaborative correction module, respectively. Then, performing dimension concatenation on the extended confidence feature tensors, the concatenated confidence features are input into a fully connected network to obtain the final stroke rehabilitation diagnosis result.

[0010] S5. The above process is trained and optimized through the loss function, and the loss is calculated between the obtained stroke rehabilitation diagnosis results and the actual results of professional doctors' diagnosis.

[0011] Furthermore, in step S1, audio, video and text data of the stroke patient during rehabilitation training are collected to obtain multimodal data. ,in, represents the audio modality data of stroke patients, represents the video modality data of stroke patients, Representing text modality data of stroke patients; selection and multimodal data Corresponding professional doctor's rehabilitation diagnosis results , build a stroke dataset , represents the amount of multimodal data, represents the lth multimodal data, Represents the professional doctor's rehabilitation diagnosis result corresponding to the lth multimodal data.

[0012] Furthermore, in step S1, the video modality data in the stroke dataset is enhanced, the audio data is converted into a frequency domain graph, and the text data is converted into a tensor, which are input into the corresponding confidence generation network to generate the video modality confidence feature. , audio modal confidence features , text modal confidence features , the formula is as follows:

[0013] ,

[0014] ,

[0015] ,

[0016] in, represents the video modality confidence generation network; represents the audio modality confidence generation network; Represents a text modality confidence generation network.

[0017] Furthermore, in step S2, the video modality confidence feature , audio modal confidence features , text modal confidence features Input into the multimodal confidence feature joint module to calculate the maximum confidence feature and the minimum confidence feature, and obtain the maximum confidence feature and the minimum confidence feature of the category. The formula is as follows:

[0018] ,

[0019] ,

[0020] in, Indicates taking the maximum value; Indicates taking the minimum value; Indicates the number of categories; represents the i-th category; represents the jth category; Represents the maximum confidence feature of the category, Represents the minimum confidence feature of the category.

[0021] Furthermore, in step S2, a confidence full feature matrix is ​​constructed based on the maximum confidence feature of the category and the minimum confidence feature of the category. By calculating the Euclidean distance between all equally likely confidence features in the confidence full feature matrix, a confidence full feature matrix with high-level spatial information is obtained. The formula is as follows:

[0022] ,

[0023] ,

[0024] in, Represents the full permutation calculation of confidence features; represents the initial confidence full feature matrix; Indicates Euclidean distance calculation; represents the first matrix of the initial confidence full feature matrix OK; represents the first matrix of the initial confidence full feature matrix OK; Represents the confidence full feature matrix with high-level spatial information; Indicates that Line and A real matrix with columns.

[0025] Furthermore, in step S2, the confidence full feature matrix with high-level spatial information is subjected to two-layer convolution calculation to extract cross-modal confidence correlation features to obtain the final confidence full feature matrix; the formula is as follows:

[0026] ,

[0027] in, Represents a convolution operation with a convolution kernel size of 3; represents a nonlinear activation function; Represents the final confidence full feature matrix.

[0028] Furthermore, step S3 specifically includes:

[0029] S31. Video modality confidence features , audio modal confidence features , text modal confidence features Input into the multimodal confidence collaborative correction module and calculate with the confidence of each category in the one-hot encoding to obtain the confidence feature correction matrix; the formula is as follows:

[0030] ,

[0031] in, Represents the confidence feature of a single modality; Represents the one-hot encoding of the correct label vectors of different categories OK; Represents the confidence feature correction matrix;

[0032] S32. Perform cosine similarity processing and linear projection calculation on the confidence feature correction matrix to obtain a confidence feature similarity matrix and a confidence feature projection matrix, which are expressed as follows:

[0033] ,

[0034] ,

[0035] in, represents the cosine similarity calculation, represents the linear projection calculation, represents the confidence feature similarity matrix, Represents the confidence feature projection matrix; Indicates that Line and A real matrix of columns, Indicates the number of modes;

[0036] S33. The confidence feature projection matrix After processing by the convolution block, it is similar to the confidence feature matrix Perform cross attention calculation to obtain the main matrix of confidence features ; The confidence feature similarity matrix After processing by the convolution block and the confidence feature projection matrix Perform cross attention calculation to obtain the confidence feature secondary matrix ; The confidence feature secondary matrix After processing by the convolution block and the main matrix of confidence features Perform cross-attention calculation to obtain the final confidence feature correction matrix, which is expressed as follows:

[0037] ,

[0038] ,

[0039] ,

[0040] in, Indicates the calculation of the convolution module with a convolution kernel size of 3×3. represents the cross attention calculation; Represents the final confidence feature correction matrix.

[0041] Furthermore, in step S4, the final confidence full feature matrix And the final confidence feature correction matrix Perform the extension operation to obtain the extension results of the confidence feature joint matrix and the confidence feature correction matrix. The formula is as follows:

[0042] ,

[0043] ,

[0044] in, Represents the extension operation of a tensor; Represents the result of extending the confidence feature joint matrix; Represents the extended result of the confidence feature correction matrix.

[0045] Furthermore, in step S4, the reliability feature joint matrix is ​​extended And the confidence feature correction matrix extension result Perform tensor splicing calculations on the dimensions, and then input the spliced ​​tensor into the fully connected network to obtain the final stroke rehabilitation diagnosis result; the formula is as follows:

[0046] ,

[0047] in, Represents the concatenation operation of tensors; represents a fully connected network; Indicates the final stroke rehabilitation diagnosis result.

[0048] Furthermore, the loss function in step S5 adopts cross entropy loss, which is expressed as follows:

[0049] ,

[0050] in, represents the probability distribution predicted by the model, represents the true label.

[0051] The advantages of the present invention are: by collecting multiple modal data as the basis for diagnosing stroke patients, the present invention is not only more in line with the actual diagnostic process of professional doctors, but also has a higher accuracy rate than the previous intelligent stroke diagnosis method relying on a single modality. In addition, the method proposed in the present invention solves the confidence conflict problem existing in the previous multimodal stroke rehabilitation diagnosis method through a multimodal confidence feature joint module and a multimodal confidence collaborative correction module, further alleviating the problem of medical resource shortage caused by traditional manual stroke diagnosis, and providing stroke patients with a more intelligent and personalized rehabilitation diagnosis evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0053] Figure 1 is a flow chart of the steps of the method of the present invention;

[0054] Figure 2 This is a working principle diagram of the multimodal confidence collaborative correction module of the method of the present invention;

[0055] Figure 3 is the confusion matrix of the MCDIF method of the present invention on the stroke test set 1;

[0056] Figure 4 This is the confusion matrix of the HEALNet model on the stroke test set 1;

[0057] Figure 5 This is the confusion matrix of the TFN model on the stroke test set 1;

[0058] Figure 6 This is the confusion matrix of the LMF model on the stroke test set 1. DETAILED DESCRIPTION

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0060] Example 1

[0061] In this embodiment, Figure 1 As shown, the present invention provides a multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis, which specifically includes the following steps:

[0062] S1. Obtain multimodal data from stroke patients during rehabilitation training and construct a stroke dataset. Preprocess the data in the stroke dataset and input the preprocessed data into the confidence generation network of the corresponding modality to obtain confidence features of different modalities.

[0063] Specifically, audio, video and text data of stroke patients during rehabilitation training are collected to obtain multimodal data. ,in, represents the audio modality data of stroke patients, represents the video modality data of stroke patients, Representing text modality data of stroke patients; selection and multimodal data Corresponding professional doctor's rehabilitation diagnosis results , build a stroke dataset , represents the amount of multimodal data, represents the lth multimodal data, represents the rehabilitation diagnosis result of the professional doctor corresponding to the lth multimodal data;

[0064] The video modality data in the stroke dataset is subjected to data enhancement operations such as horizontal flipping and rotation, the audio data is converted into a frequency domain graph, and the text data is converted into a tensor, which are input into the corresponding confidence generation network to generate video modality confidence features. , audio modal confidence features , text modal confidence features , the formula is as follows:

[0065] ,

[0066] ,

[0067] ,

[0068] in, represents the video modality confidence generation network; represents the audio modality confidence generation network; Represents a text modality confidence generation network.

[0069] S2. Input the confidence features of different modalities into the multimodal confidence feature joint module, construct a confidence full feature matrix using the maximum and minimum confidence of each category, calculate the Euclidean distance between all equally likely confidence features in the confidence full feature matrix, add high-level spatial information to the confidence full feature matrix, extract cross-modal confidence correlation features through two layers of convolution, and output the final confidence full feature matrix.

[0070] Specifically, S21. Video modality confidence features , audio modal confidence features , text modal confidence features Input into the multimodal confidence feature joint module to calculate the maximum confidence feature and the minimum confidence feature, and obtain the maximum confidence feature and the minimum confidence feature of the category. The formula is as follows:

[0071] ,

[0072] ,

[0073] in, Indicates taking the maximum value; Indicates taking the minimum value; Indicates the number of categories; represents the i-th category; represents the jth category; Represents the maximum confidence feature of the category, Represents the minimum confidence feature of the category;

[0074] S22. Construct a confidence full feature matrix based on the maximum confidence feature and the minimum confidence feature of the category. By calculating the Euclidean distance between all equally likely confidence features in the confidence full feature matrix, a confidence full feature matrix with high-level spatial information is obtained. The formula is as follows:

[0075] ,

[0076] ,

[0077] in, Represents the full permutation calculation of confidence features; represents the initial confidence full feature matrix; Indicates Euclidean distance calculation; represents the first matrix of the initial confidence full feature matrix OK; represents the first matrix of the initial confidence full feature matrix OK; Represents the confidence full feature matrix with high-level spatial information; Indicates that Line and A real matrix of columns;

[0078] S23. Perform two-layer convolution calculations on the confidence full feature matrix with high-level spatial information to extract cross-modal confidence correlation features and obtain the final confidence full feature matrix; the formula is as follows:

[0079] ,

[0080] in, Represents a convolution operation with a convolution kernel size of 3; represents a nonlinear activation function; Represents the final confidence full feature matrix.

[0081] S3. Figure 2 As shown in the figure, the confidence features of different modalities are input into the multimodal confidence collaborative correction module, and the confidence feature correction matrix is ​​obtained by calculating the correct label confidence. Then, the cosine similarity feature is used as the main feature and the linear projection feature is used as the auxiliary feature. The cross attention and convolution modules are combined to calculate the final confidence feature correction matrix.

[0082] Specifically, S31. Video modality confidence features , audio modal confidence features , text modal confidence features Input into the multimodal confidence collaborative correction module and calculate with the confidence of each category in the one-hot encoding to obtain the confidence feature correction matrix; the formula is as follows:

[0083] ,

[0084] in, Represents the confidence feature of a single modality; Represents the one-hot encoding of the correct label vectors of different categories OK; Represents the confidence feature correction matrix;

[0085] S32. Perform cosine similarity processing and linear projection calculation on the confidence feature correction matrix to obtain a confidence feature similarity matrix and a confidence feature projection matrix, which are expressed as follows:

[0086] ,

[0087] ,

[0088] in, represents the cosine similarity calculation, represents the linear projection calculation, represents the confidence feature similarity matrix, Represents the confidence feature projection matrix; Indicates that Line and A real matrix of columns, Indicates the number of modes;

[0089] S33. The confidence feature projection matrix After processing by the convolution block, it is similar to the confidence feature matrix Perform cross attention calculation to obtain the main matrix of confidence features ; The confidence feature similarity matrix After processing by the convolution block and the confidence feature projection matrix Perform cross attention calculation to obtain the confidence feature secondary matrix ; The confidence feature secondary matrix After processing by the convolution block and the main matrix of confidence features Perform cross-attention calculation to obtain the final confidence feature correction matrix, which is expressed as follows:

[0090] ,

[0091] ,

[0092] ,

[0093] in, Indicates the calculation of the convolution module with a convolution kernel size of 3×3. represents the cross attention calculation; Represents the final confidence feature correction matrix.

[0094] S4. Perform tensor extension operations on the final confidence full feature matrix output by the multimodal confidence feature joint module and the final confidence feature correction matrix output by the multimodal confidence collaborative correction module, respectively. Then, perform dimension splicing on the extended confidence feature tensors. Finally, input the spliced ​​confidence features into the fully connected network to obtain the final stroke rehabilitation diagnosis results.

[0095] Specifically, S41. The final confidence full feature matrix And the final confidence feature correction matrix Perform the extension operation to obtain the extension results of the confidence feature joint matrix and the confidence feature correction matrix. The formula is as follows:

[0096] ,

[0097] ,

[0098] in, Represents the extension operation of a tensor; Represents the result of extending the confidence feature joint matrix; Represents the extended result of the confidence feature correction matrix;

[0099] S42. Extend the reliability feature joint matrix And the confidence feature correction matrix extension result Perform tensor splicing calculations on the dimensions, and then input the spliced ​​tensor into the fully connected network to obtain the final stroke rehabilitation diagnosis result; the formula is as follows:

[0100] ,

[0101] in, Represents the concatenation operation of tensors; represents a fully connected network; Indicates the final stroke rehabilitation diagnosis result.

[0102] S5. The above process is trained and optimized through the loss function, and the loss is calculated between the obtained stroke rehabilitation diagnosis results and the actual results of professional doctors' diagnosis.

[0103] Specifically, the present invention processes a series of multimodal stroke rehabilitation clinical datasets and divides them into four parts: a training set, a validation set, and two test sets. The training set contains rehabilitation audio, video, and text data of 72 stroke patients, while the validation set and the two test sets contain rehabilitation audio, video, and text data of 16 stroke patients respectively.

[0104] The video data primarily includes videos of stroke patients undergoing Bobas handshake rehabilitation training under the guidance of professional physicians. The audio data primarily includes voice conversations and rehabilitation movement instructions between medical personnel and stroke patients during Bobas handshake rehabilitation training. The text data primarily includes basic pathological information about the stroke patients, such as age, gender, muscle strength, muscle tone, and stroke type. Muscle strength is graded from 0 to 5, as assessed by professional physicians. A higher grade indicates better muscle strength. Muscle tone is assessed using the Ashworth scale, which ranges from 0 to 4, with higher grades indicating more severe muscle tone. The Brunnstrom scale is used to assess the diagnostic results of stroke patients. The Brunnstrom scale ranges from 1 to 6, with higher grades indicating better recovery.

[0105] The present invention sequentially inputs the training set into the multimodal confidence dynamic interactive fusion network model according to steps S1-S4, and selects the cross entropy loss as the loss function to measure the absolute difference between the model output value and the true value. The formula is as follows:

[0106] ,

[0107] in, represents the probability distribution predicted by the model, represents the true label.

[0108] Example 2

[0109] In this embodiment, the method of the present invention and the existing method were experimentally compared on the data set constructed by the present invention; as shown in Table 1, MCDIF has better performance than the single-modality stroke diagnosis method. Except that the performance of MCDIF on the stroke test set 1 is basically the same as that of the audio modality stroke rehabilitation diagnosis network AST, the accuracy of MCDIF compared with other single-modality stroke intelligent diagnosis methods has been improved to a certain extent. Specifically, the accuracy of MCDIF on the validation set and two test sets is 0.3121, 0.1041 and 0.4821 higher than that of the audio modality stroke diagnosis network AST, the text modality stroke diagnosis network ERNIE and the video modality stroke diagnosis network VideoSwin, respectively. Therefore, multimodal stroke intelligent diagnosis is not only more in line with the actual application of clinical professional doctors, but also the diagnostic results are more accurate than single-modality stroke intelligent diagnosis methods, further illustrating that multimodal stroke intelligent diagnosis has certain potential in the rehabilitation diagnosis of stroke in the future.

[0110] Table 1 Performance comparison of MCDIF and other single-modality intelligent stroke diagnosis methods on the stroke dataset

[0111]

[0112] As shown in Table 2, MCDIF outperforms other multimodal fusion methods on the stroke dataset. Specifically, MCDIF achieves an average improvement in diagnostic accuracy of 0.2315, 0.2815, 0.1468, 0.1454, 0.1458, and 0.0398 on the validation set and two test sets compared to TFN, LMF, MFM, HEALNet, BBFN, and MMIM. MCDIF not only outperforms other multimodal fusion methods in diagnostic accuracy, but also demonstrates superior performance in weighted F1 scores and macro-average F1 scores. Therefore, with its efficient and precise performance, MCDIF can effectively alleviate the workload of physicians in actual diagnosis and provide more intelligent rehabilitation guidance for stroke patients.

[0113] Table 2 Performance comparison of MCDIF and other multimodal stroke intelligent diagnosis methods on the stroke dataset

[0114]

[0115] Figure 3 、 Figure 4 、 Figure 5 and Figure 6This is the confusion matrix of MCDIF and some other multimodal fusion methods on the stroke test set 1. Brunnstrom Ⅰ, Brunnstrom Ⅱ, Brunnstrom Ⅲ, Brunnstrom Ⅳ, Brunnstrom Ⅴ, and Brunnstrom Ⅵ represent the six stages of Brunnstrom stroke hemiplegia recovery; through the analysis of multiple sets of confusion matrices, it can be found that MCDIF shows significant and stable performance advantages in Brunnstrom-based stroke rehabilitation diagnosis. Specifically, for Brunnstrom III, a key rehabilitation stage for stroke patients, other multimodal fusion methods generally have errors in the diagnosis of this stage. For example, HEALNet made one diagnostic error in the stroke test set 1, TFN made two diagnostic errors in the stroke test set 1, and LMF made all diagnostic errors in the stroke test set 1. The MCDIF proposed in the present invention made all correct diagnoses in the two test sets. Therefore, it further demonstrates that our proposed MCDIF exhibits excellent consistency and stability in diagnostic results in Brunnstrom-based stroke rehabilitation diagnosis.

[0116] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis, characterized by: The following steps are involved: S1. Obtain multimodal data from stroke patients during rehabilitation training and construct a stroke dataset; And preprocess the data in the stroke dataset; Input the preprocessed data into the confidence generation network of the corresponding modality to obtain the confidence features of different modalities; S2. Input the confidence features of different modalities into the multimodal confidence feature joint module, construct a confidence full feature matrix using the maximum and minimum confidence of each category, calculate the Euclidean distance between all equally likely confidence features in the confidence full feature matrix, add high-level spatial information to the confidence full feature matrix, extract cross-modal confidence correlation features through two layers of convolution, and output the final confidence full feature matrix. S3. Input the confidence features of different modalities into the multimodal confidence collaborative correction module. Calculate the confidence of the correct label to obtain the confidence feature correction matrix. Then, based on the cosine similarity feature and supplemented by the linear projection feature, the cross-attention and convolution modules are combined to calculate the final confidence feature correction matrix. S4. Performing a tensor extension operation on the final confidence full feature matrix output by the multimodal confidence feature joint module and the final confidence feature correction matrix output by the multimodal confidence collaborative correction module, respectively. Then, performing dimension concatenation on the extended confidence feature tensors, the concatenated confidence features are input into a fully connected network to obtain the final stroke rehabilitation diagnosis result. S5. The above process is trained and optimized through the loss function, and the loss is calculated between the obtained stroke rehabilitation diagnosis results and the actual results of professional doctors' diagnosis.

2. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 1 is characterized in that: In step S1, audio, video and text data of stroke patients during rehabilitation training are collected to obtain multimodal data. ,in, represents the audio modality data of stroke patients, represents the video modality data of stroke patients, Representing text modality data of stroke patients; selection and multimodal data Corresponding professional doctor's rehabilitation diagnosis results , build a stroke dataset , represents the amount of multimodal data, represents the lth multimodal data, Represents the professional doctor's rehabilitation diagnosis result corresponding to the lth multimodal data.

3. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 2 is characterized in that: In step S1, the video modality data in the stroke dataset is enhanced, the audio data is converted into a frequency domain graph, and the text data is converted into a tensor, which are input into the corresponding confidence generation network to generate the video modality confidence feature. , audio modal confidence features , text modal confidence features , the formula is as follows: , , , in, represents the video modality confidence generation network; represents the audio modality confidence generation network; Represents a text modality confidence generation network.

4. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 3 is characterized in that: In step S2, the video modality confidence feature , audio modal confidence features , text modal confidence features Input into the multimodal confidence feature joint module to calculate the maximum confidence feature and the minimum confidence feature, and obtain the maximum confidence feature and the minimum confidence feature of the category. The formula is as follows: , , in, Indicates taking the maximum value; Indicates taking the minimum value; Indicates the number of categories; represents the i-th category; represents the jth category; Represents the maximum confidence feature of the category, Represents the minimum confidence feature of the category.

5. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 4 is characterized in that: In step S2, a confidence full feature matrix is ​​constructed based on the maximum confidence feature of the category and the minimum confidence feature of the category. By calculating the Euclidean distance between all equally likely confidence features in the confidence full feature matrix, a confidence full feature matrix with high-level spatial information is obtained. The formula is as follows: , , in, Represents the full permutation calculation of confidence features; represents the initial confidence full feature matrix; Indicates Euclidean distance calculation; represents the first matrix of the initial confidence full feature matrix OK; represents the first matrix of the initial confidence full feature matrix OK; Represents the confidence full feature matrix with high-level spatial information; Indicates that Line and A real matrix with columns.

6. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 5, characterized in that: In step S2, the confidence full feature matrix with high-level spatial information is subjected to two-layer convolution calculation to extract cross-modal confidence correlation features to obtain the final confidence full feature matrix; the formula is as follows: , in, Represents a convolution operation with a convolution kernel size of 3; represents a nonlinear activation function; Represents the final confidence full feature matrix.

7. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 6 is characterized in that: Step S3 specifically includes: S31. Video modality confidence features , audio modal confidence features , text modal confidence features Input into the multimodal confidence collaborative correction module and calculate with the confidence of each category in the one-hot encoding to obtain the confidence feature correction matrix; the formula is as follows: , in, Represents the confidence feature of a single modality; Represents the one-hot encoding of the correct label vectors of different categories OK; Represents the confidence feature correction matrix; S32. Perform cosine similarity processing and linear projection calculation on the confidence feature correction matrix to obtain a confidence feature similarity matrix and a confidence feature projection matrix, which are expressed as follows: , , in, represents the cosine similarity calculation, represents the linear projection calculation, represents the confidence feature similarity matrix, Represents the confidence feature projection matrix; Indicates that Line and A real matrix of columns, Indicates the number of modes; S33. The confidence feature projection matrix After processing by the convolution block, it is similar to the confidence feature matrix Perform cross attention calculation to obtain the main matrix of confidence features ; The confidence feature similarity matrix After processing by the convolution block and the confidence feature projection matrix Perform cross attention calculation to obtain the confidence feature secondary matrix ; The confidence feature secondary matrix After processing by the convolution block and the main matrix of confidence features Perform cross-attention calculation to obtain the final confidence feature correction matrix, which is expressed as follows: , , , in, Indicates the calculation of the convolution module with a convolution kernel size of 3×3. represents the cross attention calculation; Represents the final confidence feature correction matrix.

8. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 7, characterized in that: In step S4, the final confidence full feature matrix And the final confidence feature correction matrix Perform the extension operation to obtain the extension results of the confidence feature joint matrix and the confidence feature correction matrix. The formula is as follows: , , in, Represents the extension operation of a tensor; Represents the result of extending the confidence feature joint matrix; Represents the extended result of the confidence feature correction matrix.

9. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 8, characterized in that: In step S4, the reliability feature joint matrix is ​​extended And the confidence feature correction matrix extension result Perform tensor splicing calculations on the dimensions, and then input the spliced ​​tensor into the fully connected network to obtain the final stroke rehabilitation diagnosis result; the formula is as follows: , in, Represents the concatenation operation of tensors; represents a fully connected network; Indicates the final stroke rehabilitation diagnosis result.

10. The multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis according to claim 9, characterized in that: The loss function in step S5 adopts cross entropy loss, which is expressed as follows: , in, represents the probability distribution predicted by the model, represents the true label.

Citation Information

Patent Citations

  • Cerebral stroke risk prediction method and system based on multi-modal data fusion

    CN119943401A

  • Track confidence model

    US20230192145A1