A deep learning framework-based DWI-Flair mismatch state and stroke onset time synchronous prediction system

By fusing multimodal medical images and text information through an AFCA network based on a deep learning framework, the simultaneous prediction of DWI-FLAIR mismatch status and stroke onset time was achieved. This solves the problems of poor consistency in assessing mismatch status and insufficient information mining in existing technologies, and provides more accurate clinical decision support.

CN122117387APending Publication Date: 2026-05-29SHANGHAI FOURTH PEOPLES HOSPITAL (SHANGHAI FOURTH PEOPLES HOSPITAL AFFILIATED TO TONGJI UNIV)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI FOURTH PEOPLES HOSPITAL (SHANGHAI FOURTH PEOPLES HOSPITAL AFFILIATED TO TONGJI UNIV)
Filing Date
2026-02-24
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies suffer from high subjectivity and poor consistency in assessing DWI-FLAIR mismatch, and traditional binary classification fails to fully uncover the intrinsic correlation information of patients with acute ischemic stroke, making it difficult to meet the needs of refined clinical decision support.

Method used

Employing an AFCA network based on a deep learning framework, this system achieves simultaneous prediction of DWI-FLAIR mismatch and stroke onset time through the fusion of multimodal medical images and text information. The system comprises a data filtering and preprocessing subsystem, an AFCA deep learning subsystem, and a model training and optimization subsystem. It utilizes 3D convolutional neural networks, a large visual-language model, and a cross-modal attention mechanism to generate multimodal representations and perform dual-task prediction.

Benefits of technology

It provides detailed information on DWI-FLAIR mismatch status and onset time, supports more comprehensive treatment decisions, improves prediction accuracy, identifies potential treatment beneficiaries that may be overlooked by traditional binary criteria, reduces subjective differences, and is suitable for rapid auxiliary diagnosis in emergency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117387A_ABST
    Figure CN122117387A_ABST
Patent Text Reader

Abstract

The application discloses a DWI-Flair mismatch state and stroke onset time synchronous prediction system based on a deep learning framework, and specifically comprises a data screening and preprocessing subsystem, an AFCA deep learning subsystem and a model training and optimization subsystem.The data screening and preprocessing subsystem performs data screening, classification marking and image preprocessing based on a system architecture, thereby providing high-quality and standardized input data for subsequent deep learning models.The AFCA deep learning subsystem realizes deep fusion of multi-modal information and double-task synchronous prediction through cooperative work of multiple units.The model training and optimization subsystem adopts multiple strategies and parameters for training.The application realizes significant improvement in prediction accuracy by introducing a large language model and a multi-modal fusion mechanism, and is suitable for rapid auxiliary diagnosis in an emergency scene due to full-process automation and reduction of subjective differences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical artificial intelligence and computer vision technology, specifically to a synchronous prediction system for DWI-Flair mismatch state and stroke onset time based on a deep learning framework. Background Technology

[0002] Acute ischemic stroke (AIS) is a common critical illness in clinical practice. Intravenous thrombolytic therapy (IVT) is currently the most effective treatment, but its efficacy is subject to strict time window limitations (usually within 4.5 hours of onset). However, approximately 20% of stroke patients experience "awake stroke," with an unknown onset time, making traditional time window criteria inapplicable. To address this issue, the medical community has proposed the concept of DWI-FLAIR mismatch as an imaging biomarker. This concept is based on pathophysiological principles: after an acute ischemic stroke, the ischemic lesion appears immediately on DWI sequences, while high signal intensity on FLAIR sequences typically appears 4-6 hours after onset. Therefore, a "mismatch" pattern—where the infarct core is visible on DWI but no corresponding high signal intensity is observed on FLAIR—suggests that the patient may be within the treatment time window.

[0003] Currently, clinical assessment of DWI-FLAIR mismatch largely relies on the visual judgment of radiologists, which suffers from high subjectivity and poor consistency. Furthermore, the traditional binary classification (DWI-FLAIR match / mismatch) oversimplifies the continuous pathophysiological process, potentially leading to some patients who could still benefit from IVT treatment, especially "partially matched" patients (those showing a small area of ​​high signal intensity in the infarct core region on the FLAIR sequence). Previous studies have confirmed that approximately 50% of "partially matched" patients have an onset time within 4.5 hours. In recent years, although some studies have attempted to use machine learning or deep learning to predict onset time or DWI-FLAIR mismatch status from multimodal MRI images, most are limited to a single binary classification task, failing to fully explore the intrinsic correlation between the two and providing a more refined mismatch classification, thus failing to meet the clinical need for precise and comprehensive decision support systems. Summary of the Invention

[0004] The purpose of this invention is to provide a synchronous prediction system for DWI-Flair mismatch state and stroke onset time based on a deep learning framework, in order to solve the technical problems existing in the prior art.

[0005] Specifically, it includes: a data filtering and preprocessing subsystem, an AFCA deep learning subsystem, and a model training and optimization subsystem. (See below)Figure 1 ; The data filtering and preprocessing subsystem performs data filtering, classification and labeling, and image preprocessing based on the system architecture, providing high-quality, standardized input data for subsequent deep learning models. The AFCA deep learning subsystem achieves deep fusion of multimodal information and simultaneous prediction of dual tasks through the collaborative work of multiple units. The specific network framework of AFCA can be found in [link to AFCA]. Figure 2 ; The multiple units include an image feature extraction module, a text feature generation and extraction module, a two-stage visual fusion module, a cross-modal interaction module, and a dual-task prediction module; The model training and optimization subsystem employs multiple strategies and parameters for training.

[0006] Preferably, the image feature extraction module uses a shared-weight 3D convolutional neural network as the backbone network; The text feature generation and extraction module inputs the preprocessed DWI and FLAIR images into the visual language big model to automatically generate structured image description text. The two-stage visual fusion module enables deep fusion of multimodal image information; The dual-task prediction module contains two parallel, structurally independent classifiers that take multimodal fusion representation as a common input. The two classifiers perform Task 1 and Task 2 respectively.

[0007] Preferably, the data filtering and preprocessing subsystem specifically handles the following three aspects: The data screening and preprocessing subsystem establishes patient inclusion and exclusion criteria and sets up a standardized data preprocessing workflow. To ensure the accuracy of supervised learning labels, the data filtering and preprocessing subsystem adopts a systematic labeling system involving multiple experts, including labels for onset time groups and labels for DWI-FLAIR mismatch states groups. The data filtering and preprocessing subsystem is used to standardize the original images.

[0008] Preferably, the model training and optimization subsystem uses 5-fold cross-validation for model development and hyperparameter tuning; The model training and optimization subsystem employs a combined loss function for dual-task outputs.

[0009] Preferably, the onset time grouping label is based on the accurately calculated time from onset to scan, generating a binary label for each patient. This classification standard is aligned with the key time points of clinical thrombolytic therapy. The DWI-FLAIR mismatch status grouping label is a standardized three-class classification system for labeling.

[0010] Preferably, the two-stage visual fusion module consists of two sub-modules: an average fusion sub-module and an adaptive fusion sub-module. Average fusion submodule receives image features and Preliminary fusion features are generated through element-level averaging. ; The adaptive fusion submodule uses FLAIR features To guide the dynamic calibration of the contributions of each mode.

[0011] Preferably, the fusion feature The calculation formula is: .

[0012] Preferably, the adaptive fusion submodule first... respectively with and Perform convolutional interactions; then apply the Softmax function to the interaction results to generate two sets of modality-specific attention weights; finally, apply the weights to the original features respectively. and The weighted features are then summed to output the final adaptive visual feature vector. .

[0013] Preferably, the feature vector The calculation formula is: .

[0014] This invention also provides a method for simultaneous prediction of DWI-FLAIR mismatch state and stroke onset time based on a deep learning framework. This method is executed by the simultaneous prediction system of this invention and specifically includes: Step 1: Based on the system architecture, perform data filtering, classification and labeling, and image preprocessing to provide high-quality, standardized input data for subsequent deep learning models; Step 2: Through the collaborative work of multiple units, achieve deep fusion of multimodal information and simultaneous prediction of dual tasks; Step 3: Use multiple strategies and parameters for training to ensure its performance and generalization ability.

[0015] The method of the present invention has the following advantages: (1) Provide clinicians with detailed information on DWI-FLAIR mismatch status and onset time, supporting more comprehensive treatment decisions.

[0016] (2) By introducing a large language model (Qwen-3) and a multimodal fusion mechanism, the present invention has achieved a significant improvement in prediction accuracy.

[0017] (3) The three-classification model can identify potential patients with "partial mismatch" and onset time <4.5 hours that may be overlooked by the traditional binary standard, which is expected to enable more stroke patients who have awakened or whose onset time is unknown to have the opportunity for thrombolysis.

[0018] (4) The entire process is automated, reducing subjective differences and making it suitable for rapid auxiliary diagnosis in emergency scenarios. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the synchronous prediction system of the present invention; Figure 2 This is a diagram of the AFCA framework involved in this invention; Figure 3 This is a diagram of the 3D ResNet network framework involved in this invention; Figure 4 This is a schematic diagram illustrating the image feature extraction involved in this invention; Figure 5 This is a schematic diagram illustrating the text feature generation and extraction involved in this invention; Figure 6 This is a flowchart of the synchronous prediction method involved in the present invention; Figure 7 A schematic diagram illustrating the classification of mismatches in DWI-FLAIR images and the classification of onset time in the test set; Figure 8 This is a schematic diagram of the training and testing sets for the model stability experiment in the DWI-FLAIR mismatched three-class classification task involved in this invention; Figure 9 This is a schematic diagram of the training and testing sets for the model stability experiment in the binary classification task of onset time involved in this invention. Detailed Implementation

[0020] The technical solution of the present invention will be clearly and completely described below with reference to specific embodiments. However, those skilled in the art should understand that the embodiments described below are only for illustrating the present invention and should not be regarded as limiting the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] This invention provides a deep learning-based system for the simultaneous prediction of DWI-FLAIR mismatch and stroke onset time. The system involves an Adaptive Fusion via Cross-Attention Network (AFCA) that fuses multimodal medical images and text information to simultaneously predict the DWI-FLAIR mismatch and time since Stroke Onset (TSS) of patients with acute ischemic stroke (AIS). The system first acquires diffusion-weighted imaging (DWI) and fluid-attenuated inversion recovery (FLAIR) image data from the patient's magnetic resonance imaging (MRI); it then extracts depth features from the bimodal images using a pre-trained 3D ResNet deep learning framework; and finally, it employs a system that includes an Average Fusion Module and an Adaptive Fusion Module. The invention employs a two-stage fusion module to dynamically integrate bimodal image features and generate a fused visual feature vector. Simultaneously, the image data is input into the Qwen-3 large language model to generate text descriptions, and the text encoding model BERT is used to extract text feature vectors. Subsequently, a cross-modal attention module allows the text features to interact with the fused visual features as a query, generating the final multimodal representation. Finally, this multimodal representation is input into two independent classifiers to simultaneously complete a three-classification task (complete match / partial match / no match) for DWI-FLAIR matching and a two-classification task (<4.5h / ≥4.5h) for stroke onset time. This invention overcomes the limitations of existing clinical standards that use a single DWI-FLAIR mismatch to select intravenous thrombolysis (IVT) candidates for post-aware stroke and stroke with unknown onset time (AIS), providing a more accurate and comprehensive artificial intelligence (AI)-assisted tool for thrombolytic therapy decisions in stroke patients with unknown onset time.

[0022] The system integrates computer vision, natural language processing and deep learning technologies. Its core lies in the novel hybrid AFCA fusion architecture, which can realize the deep fusion of multimodal information of images and text, and thus complete the synchronous prediction of two key indicators: DWI-FLAIR image mismatch and stroke onset time.

[0023] like Figure 1 As shown, the present invention provides a synchronous prediction system for DWI-Flair mismatch state and stroke onset time based on a deep learning framework, specifically including: a data screening and preprocessing subsystem, an AFCA deep learning subsystem, and a model training and optimization subsystem.

[0024] The data filtering and preprocessing subsystem performs data filtering, classification and labeling, and image preprocessing based on the system architecture, providing high-quality, standardized input data for subsequent deep learning models.

[0025] The data filtering and preprocessing subsystem specifically handles the following three aspects: (1) Data filtering The data screening and preprocessing subsystem establishes patient inclusion and exclusion criteria and a standardized data preprocessing workflow. Data meeting the criteria are retained, while those not meeting the criteria are excluded.

[0026] The specific criteria are as follows: Patient inclusion criteria: A. Patients with clinically diagnosed acute ischemic stroke to ensure the clinical relevance of the data; B. Patients aged ≥18 years to exclude the interference of physical function differences in minors on imaging characteristics and disease patterns; C. Head MRI data containing DWI and FLAIR sequences collected within 24 hours of onset and before treatment to ensure that the imaging data can truly reflect the pathological state of the acute phase of stroke; D. A clear time from stroke onset to MRI scan to provide accurate basis for the annotation of onset time labels.

[0027] Exclusion criteria: A. Acute infarct lesions are located within a region of extensive white matter lesions to avoid assessment interference and ensure that lesion features can be clearly identified and analyzed; B. Image quality is insufficient to support accurate imaging assessment, excluding the impact of blurry, distorted, or other substandard images on model training and prediction results.

[0028] (2) Classification markers To ensure the accuracy of supervised learning labels, the data filtering and preprocessing subsystem employs a systematic labeling system involving multiple experts, including labels for onset time groups and labels for DWI-FLAIR mismatch states groups.

[0029] Among them, the onset time group label is based on the accurately calculated time from onset to scan, generating a binary label for each patient (TSS<4.5 hours or TSS≥4.5 hours). This classification standard is in line with the key time nodes of clinical thrombolytic therapy.

[0030] The DWI-FLAIR mismatch grouping labeling uses a standardized three-class classification system, executed by a team of neuroradiology experts, to ensure the professionalism and accuracy of the labeling results. Specific classification criteria include: Complete match: The high signal area of ​​the lesion on the FLAIR sequence is basically consistent with the high signal area on the DWI, reflecting that the lesion is in a relatively stable stage; Partial match: The area of ​​high signal intensity of the lesion on the FLAIR sequence is significantly smaller than that on the DWI sequence, suggesting that the lesion is still in the process of progression; Mismatch: A clear lesion is visible on the DWI sequence, but no obvious high signal is seen on the corresponding FLAIR sequence, which often indicates that the stroke has a short onset time and is in the hyperacute phase.

[0031] In addition, the above labeling process includes two stages of quality control: preliminary assessment (completed independently by multiple junior physicians) and consensus review (arbitrated by a team of senior experts on disputed cases). Through this dual verification mechanism, the objectivity and authority of the labels are ensured.

[0032] (3) Image preprocessing The data filtering and preprocessing subsystem is used to standardize the original images.

[0033] Specifically, the BET tool from the FSL toolkit was used for skull dissection and grayscale intensity normalization to eliminate scanner differences, and all image centers were cropped to a uniform 224×224×20 voxel size to ensure the consistency of input data.

[0034] The AFCA deep learning subsystem achieves deep fusion of multimodal information and simultaneous prediction of dual tasks through the collaborative work of multiple units.

[0035] The system comprises multiple modules, including an image feature extraction module, a text feature generation and extraction module, a two-stage visual fusion module, a cross-modal interaction module, and a dual-task prediction module.

[0036] The image feature extraction module (Medical Net) employs a weighted 3D convolutional neural network as its backbone. This network is based on an existing 3D ResNet-50 pre-trained on a large 3D medical image dataset, specifically designed for volumetric medical image analysis. (See [link to relevant documentation]). Figure 3 This module processes DWI and FLAIR images separately, outputting the corresponding depth feature tensors. The DWI features are denoted as... The FLAIR feature is denoted as .like Figure 4 As shown.

[0037] The text feature generation and extraction module inputs the preprocessed DWI and FLAIR images into the visual language large model (Qwen-3) to automatically generate structured image description text. Subsequently, the BERT text encoding model is used to extract these text features. Specifically, the embedding vector of the [CLS] token output from the final layer of BERT is taken as the aggregate representation of the entire sentence, and the text features from DWI and FLAIR are concatenated into a single text feature vector. This can enhance the model's semantic understanding capabilities. For example... Figure 5 As shown.

[0038] The two-stage visual fusion module is key to achieving deep fusion of multimodal image information, and consists of two sub-modules: an average fusion sub-module and an adaptive fusion sub-module. Average Fusion Module: Receives image features and Preliminary fusion features are generated through element-level averaging. The calculation formula is: Adaptive Fusion Module: Based on FLAIR features To guide the dynamic calibration of the contributions of each mode, firstly, the contributions of each mode are... respectively with and Convolutional interactions are performed; then, the Softmax function is applied to the interaction results to generate two sets of modality-specific attention weights; finally, the weights are applied to the original features respectively. and The weighted features are then summed to output the final adaptive visual feature vector. The calculation formula is: .

[0039] This design enables the model to dynamically adjust the weights of DWI and fused features based on their contextual relevance to the FLAIR modality.

[0040] The cross-modal interaction module employs a multi-head attention mechanism. Based on the aforementioned text feature vectors... As a query, it adapts to visual features. Simultaneously serving as both key and value, cross-modal attention computation is performed to achieve deep alignment of visual and textual information. This mechanism has an embedding dimension of 768 and uses eight attention heads to facilitate richer feature interactions, outputting the final multimodal fusion representation. The calculation formula is: The dual-task prediction module comprises two parallel, structurally independent classifiers (fully connected layers). This is based on the aforementioned multimodal fusion representation. With a common input, the two classifiers perform Task 1 and Task 2 respectively, that is, one classifier performs Task 1 and the other performs Task 2.

[0041] Task 1: Classify DWI-FLAIR mismatch states into three categories, and output the results as "complete match", "partial match" or "no match".

[0042] Task 2: Binary classification of stroke onset time, outputting results as "<4.5 hours" or "≥4.5 hours". Example of specific classification results is shown below. Figure 7 As shown.

[0043] Figures A and B show examples of DWI-FLAIR image mismatch classification and onset time classification in the test set. Heatmaps generated using Grad-CAM interpretability technology visualize the key regions the model focuses on in the input images when making classification decisions, validating the consistency between its decision logic and clinicopathological features.

[0044] Figure A, first row (“Mismatch” mode cases): Shows the model's analysis of a patient with a “mismatch” between DWI and FLAIR (high signal on DWI, no clearly abnormal signal on FLAIR). The Grad-CAM heatmap shows that the model's attention distribution on the DWI sequence is weak and highly limited, focusing only minimally on the infarct core area. Second row (“Partial Match” mode cases): Shows the model's analysis of a patient with a “partial match” between DWI and FLAIR (both DWI and FLAIR show abnormal signals, but with differences in degree or extent). The generated attention heatmap highlights and covers the areas within the infarct core corresponding to early signal changes. Third row (“Perfect Match” mode cases): Shows the model's analysis of a patient with a “perfect match” between DWI and FLAIR (both DWI and FLAIR show clearly high signal in the ischemic area). The model's attention is highly focused on the ischemic area showing high signal on both sequences.

[0045] The Grad-CAM heatmap in Figure B shows that, regardless of the duration of the illness (<4.5 hours and ≥4.5 hours), the model's core attention is always focused on the infarct core area.

[0046] The model training and optimization subsystem employs multiple strategies and parameters for training to ensure its performance and generalization ability.

[0047] The model training and optimization subsystem employs a 5-fold cross-validation method for model development and hyperparameter tuning. In each fold of training, an additional 20% of the data is allocated from the training set as a validation set to monitor the training process and implement an early stopping strategy to prevent overfitting.

[0048] The model training and optimization subsystem uses the Adam optimizer with the following parameters: a fixed learning rate of 0.001, momentum parameters β1=0.9, β2=0.999, and weight decay of 0. Simultaneously, an exponential learning rate scheduler is employed, multiplying the learning rate by a decay factor γ=0.99 each training cycle to achieve a dynamic and smooth decrease in the learning rate.

[0049] The model training and optimization subsystem employs a combined loss function for dual-task outputs. Specifically, the TSS binary classification task uses binary cross-entropy loss, while the DWI-FLAIR three-class classification task uses classification cross-entropy loss.

[0050] The model training and optimization subsystem consisted of 500 epochs, with a batch size of 16. All experiments were performed on an NVIDIA GeForce RTX 4090 GPU.

[0051] This invention also proposes a method for simultaneous prediction of DWI-FLAIR mismatch state and stroke onset time based on a deep learning framework, specifically including: Step 1: Based on the system architecture, perform data filtering, classification and labeling, and image preprocessing to provide high-quality, standardized input data for subsequent deep learning models.

[0052] Step 1.1: Data Filtering Establish patient inclusion and exclusion criteria, and create a standardized data preprocessing workflow. Data meeting the criteria are retained, while those not meeting the criteria are excluded.

[0053] The specific criteria are as follows: Patient inclusion criteria: A. Patients with clinically diagnosed acute ischemic stroke to ensure the clinical relevance of the data; B. Patients aged ≥18 years to exclude the interference of physical function differences in minors on imaging characteristics and disease patterns; C. Head MRI data containing DWI and FLAIR sequences collected within 24 hours of onset and before treatment to ensure that the imaging data can truly reflect the pathological state of the acute phase of stroke; D. A clear time from stroke onset to MRI scan to provide accurate basis for the annotation of onset time labels.

[0054] Exclusion criteria: A. Acute infarct lesions are located within a region of extensive white matter lesions to avoid assessment interference and ensure that lesion features can be clearly identified and analyzed; B. Image quality is insufficient to support accurate imaging assessment, excluding the impact of blurry, distorted, or other substandard images on model training and prediction results.

[0055] Step 1.2: The following systematic, multi-expert annotation system is adopted, including onset time grouping labels and DWI-FLAIR mismatch status grouping labels.

[0056] Among them, the onset time group label is based on the accurately calculated time from onset to scan, generating a binary label for each patient (TSS<4.5 hours or TSS≥4.5 hours). This classification standard is in line with the key time nodes of clinical thrombolytic therapy.

[0057] The DWI-FLAIR mismatch grouping labeling uses a standardized three-class classification system, executed by a team of neuroradiology experts, to ensure the professionalism and accuracy of the labeling results. Specific classification criteria include: Complete match: The high signal area of ​​the lesion on the FLAIR sequence is basically consistent with the high signal area on the DWI, reflecting that the lesion is in a relatively stable stage; Partial match: The area of ​​high signal intensity of the lesion on the FLAIR sequence is significantly smaller than that on the DWI sequence, suggesting that the lesion is still in the process of progression; Mismatch: A clear lesion is visible on the DWI sequence, but no obvious high signal is seen on the corresponding FLAIR sequence, which often indicates that the stroke has a short onset time and is in the hyperacute phase.

[0058] In addition, the above labeling process includes two stages of quality control: preliminary assessment (completed independently by multiple junior physicians) and consensus review (arbitrated by a team of senior experts on disputed cases). Through this dual verification mechanism, the objectivity and authority of the labels are ensured.

[0059] Step 1.3: Standardize the original image.

[0060] Specifically, the BET tool from the FSL toolkit was used for skull dissection and grayscale intensity normalization to eliminate scanner differences, and all image centers were cropped to a uniform 224×224×20 voxel size to ensure the consistency of input data.

[0061] Step 2: Through the collaborative work of multiple units, achieve deep fusion of multimodal information and simultaneous prediction of dual tasks.

[0062] Step 2.1: A shared-weight 3D convolutional neural network is used as the backbone network. This network is based on the existing 3D ResNet-50 pre-trained on a large 3D medical image dataset, specifically designed for volumetric medical image analysis. See [link to relevant documentation]. Figure 3 This module processes DWI and FLAIR images respectively, outputting the corresponding depth feature tensors. The two image features, the DWI feature and the FLAIR feature, are denoted as follows: The FLAIR feature is denoted as .

[0063] Step 2.2: Input the preprocessed DWI and FLAIR images into the visual language model (Qwen-3) to automatically generate structured image description text. Then, use the BERT text encoding model to extract these text features. Specifically, take the embedding vector of the [CLS] token output from the final layer of BERT as the aggregate representation of the entire sentence, and concatenate the text features from DWI and FLAIR into a single text feature vector. This can enhance the model's semantic understanding capabilities.

[0064] Step 2.3: Receive image features and Preliminary fusion features are generated through element-level averaging. The calculation formula is: Step 2.4, using FLAIR features To guide the dynamic calibration of the contributions of each mode, firstly, the contributions of each mode are... respectively with and Convolutional interactions are performed; then, the Softmax function is applied to the interaction results to generate two sets of modality-specific attention weights; finally, the weights are applied to the original features respectively. and The weighted features are then summed to output the final adaptive visual feature vector. The calculation formula is: .

[0065] This design enables the model to dynamically adjust the weights of DWI and fused features based on their contextual relevance to the FLAIR modality.

[0066] Step 2.5: Employ a multi-head attention mechanism. Using the aforementioned text feature vectors... As a query, it adapts to visual features. Simultaneously serving as both key and value, cross-modal attention computation is performed to achieve deep alignment of visual and textual information. This mechanism has an embedding dimension of 768 and uses eight attention heads to facilitate richer feature interactions, outputting the final multimodal fusion representation. The calculation formula is: Step 2.5: Characterize using the aforementioned multimodal fusion method. To obtain a common input, two tasks are executed separately.

[0067] Task 1: Three-class classification of DWI-FLAIR mismatch states, with output results of "complete match", "partial match" or "no match".

[0068] Task 2: Classify stroke onset time into two categories, outputting results as "<4.5 hours" or "≥4.5 hours".

[0069] Step 3: Use multiple strategies and parameters for training to ensure its performance and generalization ability.

[0070] Specifically, a 5-fold cross-validation method was used for model development and hyperparameter tuning. In each fold of training, an additional 20% of the data was allocated from the training set as a validation set to monitor the training process and implement an early stopping strategy to prevent overfitting.

[0071] The Adam optimizer is used with the following parameters: a fixed learning rate of 0.001, momentum parameters β1=0.9, β2=0.999, and weight decay of 0. An exponential learning rate scheduler is employed, multiplying the learning rate by a decay factor γ=0.99 each training cycle to achieve a dynamic and smooth decrease in the learning rate.

[0072] For dual-task outputs, a combined loss function is employed. Specifically, the TSS binary classification task uses binary cross-entropy loss, while the DWI-FLAIR tri-class classification task uses classification cross-entropy loss.

[0073] The model was trained for a total of 500 epochs with a batch size of 16. All experiments were performed on an NVIDIA GeForce RTX 4090 GPU.

[0074] The technical effects of this invention will be demonstrated below through model training, evaluation, and comparison of experimental data.

[0075] 1. Performance Evaluation Metrics: The model performance is comprehensively evaluated on the training set and independent test set using the following metrics: accuracy, sensitivity, specificity, F1 score, and area under the receiver operating characteristic curve.

[0076] 2. Comparative Experimental Design: To verify the effectiveness of each module of this invention, a rigorous ablation experiment was conducted: Baseline model: It directly processes stitched DWI and FLAIR images using only 3D ResNet, without any special fusion or attention mechanisms.

[0077] AFCA-Image (Image Fusion Only): Includes a two-stage visual fusion module, but does not introduce text modalities.

[0078] AFCA-Text (Text Only): Classifies using only text features extracted by BERT.

[0079] Model-Multimodal (This invention, image fusion + text): A complete adaptive fusion and cross-modal attention multimodal model.

[0080] 3. Experimental Results: (1) Baseline characteristics of the patient cohort This study ultimately included 333 patients, with 267 in the training set and 66 in the external test set. As shown in Table 1, there were no statistically significant differences between the training and test sets in all baseline clinical variables, including age, sex, medical history (hypertension, diabetes, atrial fibrillation, etc.), neurological deficit score at admission, systolic / diastolic blood pressure, and blood glucose levels (p>0.05). Regarding the core predictive objective, there were also no statistically significant differences between the training and test sets in the distribution of time to onset (TSS) (p=0.062) and the proportion of DWI-FLAIR mismatched tri-class classifications (p=0.337), ensuring that the test set could fairly assess the model's generalization ability.

[0081] Table 1. Experimental data on baseline characteristics of the patient cohort IQR (Interquartile Range); NIHSS (National Institutes of Health Stroke Scale); for continuous variables, the Mann-Whitney U test was used to analyze differences between two independent sample groups; for categorical variables, the chi-square test was used. A p-value less than 0.05 was considered statistically significant. (2) Quantitative evaluation of model performance The performance of the model on the independent test set (n=66) was comprehensively evaluated by accuracy, sensitivity, specificity, F1 score and area under the receiver operating characteristic curve.

[0082] 1) DWI-FLAIR does not match the three-class classification task, as shown in Table 2: Table 2 Performance metrics of each comparison model on the mismatch classification task The receiver operating characteristic (ROC) curves of the comparative model on the DWI-Flair matching classification task are as follows: Figure 8As shown, the ROC curve of the complete model of this invention is closest to the top left corner of the coordinate graph, with AUCs of 0.887 and 0.888, respectively, which is the highest among the four comparison models. The curves of the AFCA-Image model (AUC=0.876 and 0.877) and the AFCA-Text model (AUC=0.867 and 0.861) are located below it, while the curve of the Baseline model (AUC=0.857 and 0.854) is the lowest. This result intuitively shows that the introduction of adaptive fusion architecture and cross-modal attention mechanism systematically improves the model's ability to distinguish different mismatch states.

[0083] Training set performance: The AFCA-Multimodal model of this invention achieved the highest accuracy (0.821) and AUC (0.887) on the training set, indicating that the model has strong learning ability.

[0084] Generalization performance on external test sets: On the independent test set, the model of this invention also performs best, with a classification accuracy of 0.810 and an AUC of 0.888. Compared with the baseline model, the model of this invention shows an absolute improvement of 0.105 (10.5 percentage points) in classification accuracy and an absolute improvement of 0.034 in AUC. Compared with the AFCA-Image model that only uses images, the model of this invention achieves an additional improvement of 0.038 (3.8 percentage points) in accuracy by introducing the text modality.

[0085] Ablation experiments show that the AFCA-Image model significantly outperforms the Baseline model (improving test set accuracy by 0.067), demonstrating the effectiveness of the adaptive fusion architecture. The AFCA-Multimodal model further surpasses the AFCA-Image model, validating the value of cross-modal attention mechanisms in integrating textual information.

[0086] like Figure 8 As shown, the model stability is as follows: the accuracy difference between the model of this invention (AFCA-Multimodal) on the training set and the test set is only 0.011 (1.1 percentage points), and the AUC difference is less than 0.001. This small difference indicates that the model maintains a high degree of consistency with the classification ability on unseen test data as it does on the training set, without overfitting, and possesses excellent generalization stability and reliability for clinical deployment.

[0087] 2) Performance on the binary classification task based on onset time, as shown in Table 3: Table 3. Performance of each model on the task of predicting whether the onset time is <4.5 hours. The ROC curves of the comparison model on the Time of Onset (TSS) classification task are as follows: Figure 9As shown: In the training set and external test set, the AUC squares of the complete model of this invention are 0.889 and 0.879, respectively, and its ROC curve also occupies the most dominant position. The curves of the AFCA-Image model (AUC=0.885 and 0.862), AFCA-Text model (AUC=0.871 and 0.862), and Baseline model (AUC=0.855 and 0.839) are arranged in order. The separation between the curves of each model is clearly visible. In particular, the model of this invention can maintain higher specificity in the higher sensitivity range (>0.7), which directly confirms its superior performance in balancing the identification of true early stroke patients and avoiding false positives.

[0088] Training set performance: The model of this invention achieved an accuracy of 0.852, an AUC of 0.889, and an F1 score of 0.888 on the training set, all of which are the best among all models.

[0089] Generalization performance on external test sets: On independent test sets, the model of this invention continues to lead, with a prediction accuracy of 0.855, an AUC of 0.879, and an F1 score of 0.880. Compared to the baseline model, the model of this invention shows an absolute improvement of 0.073 (7.3 percentage points) in prediction accuracy, an absolute improvement of 0.04 in AUC, and an absolute improvement of 0.091 in F1 score. Compared to the AFCA-Image model, the model of this invention has a higher F1 score of 0.072 (7.20 percentage points), further demonstrating the advantages of multimodal fusion.

[0090] Ablation study results: Compared to the baseline model, the AFCA-Image model improved the F1 score on the test set by 0.019 (1.9 percentage points), verifying the effect of introducing an adaptive fusion mechanism on improving the accuracy of temporal prediction. Compared to the AFCA-Image model, the AFCA-Multimodal complete model further improved the F1 score on the test set by 0.072 (7.2 percentage points).

[0091] like Figure 9 As shown, the model stability is as follows: the performance indicators (such as accuracy and F1 score) of the model on the training set and the test set are less than 0.04, which shows good generalization ability and stability, and no obvious overfitting phenomenon is observed.

Claims

1. A synchronous prediction system for DWI-FLAIR mismatch state and stroke onset time based on a deep learning framework, specifically comprising: Data filtering and preprocessing subsystem, AFCA deep learning subsystem, model training and optimization subsystem; The data filtering and preprocessing subsystem performs data filtering, classification and labeling, and image preprocessing based on the system architecture, providing high-quality, standardized input data for subsequent deep learning models. The AFCA deep learning subsystem achieves deep fusion of multimodal information and simultaneous prediction of dual tasks through the collaborative work of multiple units. The multiple units include an image feature extraction module, a text feature generation and extraction module, a two-stage visual fusion module, a cross-modal interaction module, and a dual-task prediction module; The model training and optimization subsystem employs multiple strategies and parameters for training.

2. The synchronous prediction system according to claim 1, characterized in that, The image feature extraction module uses a shared-weight 3D convolutional neural network as its backbone network. The text feature generation and extraction module inputs the preprocessed DWI and FLAIR images into the visual language big model to automatically generate structured image description text. The two-stage visual fusion module enables deep fusion of multimodal image information; The dual-task prediction module contains two parallel, structurally independent classifiers that take multimodal fusion representation as a common input. The two classifiers perform Task 1 and Task 2 respectively.

3. The synchronous prediction system according to claim 1, characterized in that, The data filtering and preprocessing subsystem specifically handles the following three aspects: The data screening and preprocessing subsystem establishes patient inclusion and exclusion criteria and sets up a standardized data preprocessing workflow. To ensure the accuracy of supervised learning labels, the data filtering and preprocessing subsystem adopts a systematic labeling system involving multiple experts, including labels for onset time groups and labels for DWI-FLAIR mismatch states groups. The data filtering and preprocessing subsystem is used to standardize the original images.

4. The synchronous prediction system according to claim 1, characterized in that, The model training and optimization subsystem uses a 5-fold cross-validation method for model development and hyperparameter tuning. The model training and optimization subsystem employs a combined loss function for dual-task outputs.

5. The synchronous prediction system according to claim 3, characterized in that, The onset time grouping label is based on the accurately calculated time from onset to scan, generating a binary label for each patient. This classification standard is aligned with the key time points of clinical thrombolytic therapy. The DWI-FLAIR mismatch status grouping label is a standardized three-class classification system for labeling.

6. The synchronous prediction system according to claim 1, characterized in that, The two-stage visual fusion module consists of two sub-modules: an average fusion sub-module and an adaptive fusion sub-module. Average fusion submodule receives image features and Preliminary fusion features are generated through element-level averaging. ; The adaptive fusion submodule uses FLAIR features To guide the dynamic calibration of the contributions of each mode.

7. The synchronous prediction system according to claim 6, characterized in that, The fusion feature The calculation formula is: 。 8. The synchronous prediction system according to claim 6, characterized in that, The adaptive fusion submodule first will respectively with and Perform convolutional interactions; then apply the Softmax function to the interaction results to generate two sets of modality-specific attention weights; Finally, the weights are applied to the original features respectively. and The weighted features are then summed to output the final adaptive visual feature vector. .

9. The synchronous prediction system according to claim 8, characterized in that, The feature vector The calculation formula is: 。 10. A method for simultaneous prediction of DWI-FLAIR mismatch state and stroke onset time based on a deep learning framework, wherein the method is executed by the simultaneous prediction system of any one of claims 1-9, specifically comprising: Step 1: Based on the system architecture, perform data filtering, classification and labeling, and image preprocessing to provide high-quality, standardized input data for subsequent deep learning models; Step 2: Through the collaborative work of multiple units, achieve deep fusion of multimodal information and simultaneous prediction of dual tasks; Step 3: Use multiple strategies and parameters for training to ensure its performance and generalization ability.