A head magnetic resonance additional examination scheme recommendation method and electronic equipment

CN122573971BActive Publication Date: 2026-09-29SHANGHAI SIXTH PEOPLES HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611055782.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-29
Estimated Expiration
2046-07-16

AI Technical Summary

Technical Problem

例如,在诊断脑肿瘤时,平扫可能无法准确识别肿瘤的边界和性质,需要进一步的增强扫描来提供更多信息

Benefits of technology

[0018]相较于现有技术,本发明针对现有磁共振头部平扫诊断中存在的单一模态局限、多模态比对繁琐、补扫频繁及依赖医生经验等痛点,具备以下优势:(1)本发明通过多模态特征融合,可充分发挥T1W、T2W、FLAIR、DWI各自的成像优势,尤其对微小病灶或表现不典型的病变(如早期脑肿瘤、微小缺血灶),能强化特征辨识度,快速完成病灶层定位、疾病分类及病灶精确定位,大幅提升临床诊断效率;(2)基于平扫多模态数据的自动化异常检测结果,快速判断是否需要补扫,结合预设规则库精准匹配附加检查方案,将补扫方案提示信息即时提供给医生或技师,可以马上安排当场补扫,减少患者因补扫多次往返医院的身体与心理负担,同时降低医院因重复扫描产生的资源消耗。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122573971B_ABST
    Figure CN122573971B_ABST
Patent Text Reader

Abstract

The application provides a head magnetic resonance additional examination scheme recommendation method and electronic equipment, the method comprises the following steps: step 1: obtaining the multi-modal image obtained by head magnetic resonance plain scan of the examinee, the multi-modal image comprises a T1W image, a T2W image, a DWI image and a FLAIR image; step 2: according to the multi-modal image, abnormality detection is carried out based on a deep learning model, to obtain disease type information and lesion location information of the examinee; step 3: according to the disease type information and / or lesion location information, it is determined whether additional examination is needed, if additional examination is needed, the corresponding additional examination scheme is matched according to the disease type information and / or lesion location information; step 4: generating prompt information containing the disease type information, the lesion location information and the additional examination scheme. The application takes advantage of T1W, T2W, FLAIR and DWI to quickly identify lesions and give additional examination schemes, greatly improving the diagnostic accuracy and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, specifically to a method for processing head magnetic resonance images and an electronic device for implementing the method. Background Technology

[0002] Plain MRI scan is a basic magnetic resonance imaging (MRI) examination without contrast agent injection. It relies on the natural signal differences between different tissues in the body to generate images and is the most commonly used type of MRI examination in clinical practice. In clinical practice, plain MRI scans use one or more different scanning sequences to obtain images of different modalities, depending on the area being examined. Different modalities produce different imaging effects and play different roles in the diagnosis of different disease types. For example, T1-weighted imaging (T1W) can clearly show anatomical structures such as bones and organ outlines; T2-weighted imaging (T2W) highlights fluid and edematous tissues; fluid attenuation inversion recovery imaging (FLAIR) can suppress cerebrospinal fluid signals to more clearly show brain lesions; and diffusion-weighted imaging (DWI) is often used to detect cerebral ischemia.

[0003] Single-modal image analysis methods can limit the ability to identify lesions, especially in cases of small lesions or atypical presentations, easily leading to misdiagnosis or missed diagnosis. For example, in the diagnosis of brain tumors, T1-weighted imaging (T1W) may not clearly show the tumor boundaries, while T2-weighted imaging (T2W) can better reflect the edema around the tumor. If doctors rely solely on T1 images, they may underestimate the invasiveness of the tumor, affecting the formulation of subsequent treatment plans. Similarly, in the diagnosis of stroke, in the case of acute ischemic stroke, DWI (diffusion-weighted imaging) can identify ischemic areas of brain tissue in the early stages, while FLAIR (fluid attenuated inversion recovery) is advantageous in assessing cerebral edema and hemorrhage.

[0004] When dealing with complex brain diseases, doctors often need to perform tedious comparisons between images of different modalities and make comprehensive judgments based on these images, which increases the complexity and time cost of diagnosis.

[0005] Furthermore, for some diseases, accurate diagnosis cannot be achieved solely through plain CT scans. Additional examinations are necessary to obtain images in other modalities for a more precise diagnosis. For example, in diagnosing brain tumors, plain CT scans may not accurately identify the tumor's boundaries and nature, requiring contrast-enhanced scans to provide more information. Contrast-enhanced scans use contrast agents to improve image contrast, helping doctors see lesions more clearly. However, this also means patients need to undergo additional examinations and make multiple trips to the hospital, increasing their physical and psychological burden, as well as the hospital's resource consumption.

[0006] A routine MRI head examination usually recommends that patients complete the images of the four modalities of T1W, T2W, FLAIR, and DWI in one scan. Doctors may also arrange additional scanning sequences based on the consultation and experience. However, relying on the doctor's experience, it is not uncommon for the obtained images to be insufficient for the diagnosis, and additional examinations may need to be arranged. Summary of the Invention

[0007] This invention provides a method for processing head magnetic resonance imaging (MRI) images, comprising the following steps: Step 1: Obtain multimodal images from the plain magnetic resonance imaging (MRI) scan of the subject's head, including T1W images, T2W images, DWI images, and FLAIR images; Step 2: Based on the multimodal images, perform anomaly detection using a deep learning model to obtain the subject's disease type information and lesion location information; Step 3: Determine whether additional examinations are needed based on the disease type information and / or lesion location information. If additional examinations are needed, match the corresponding additional examination plan based on the disease type information and / or lesion location information. Step 4: Generate a prompt message containing the disease type information, the lesion location information, and the additional examination plan; display the prompt message to the user.

[0008] Optionally, the multimodal image is preprocessed before being input into the deep learning model; the preprocessing includes denoising, normalization, and image enhancement.

[0009] The anomaly detection based on the deep learning model specifically includes the following steps: Step 2.1: Feature extraction and fusion: Extract features for each modality in the multimodal image to obtain shallow and deep features corresponding to each modality. Perform cross-modal fusion processing on the shallow and deep features to obtain multimodal fused features. Step 2.2: Lesion layer localization: Based on the multimodal fusion features, a classification network is used to determine whether there are lesions in each layer of the scan image, and the target scan layer where the lesion is located is selected. Step 2.3: Disease Classification: Extract the multimodal fusion features corresponding to the target scanning layer, input them into the disease classification network, output the probability distributions corresponding to various disease types, and obtain the disease type information; Step 2.4: Precise localization of lesion location: Based on the multimodal fusion features of the target scanning layer, the boundary coordinate information of the lesion is output through the target detection network to obtain the lesion location information.

[0010] In step 2.1, a CNN network is used to extract shallow local features, and a ViT network is used to extract deep global features. The shallow local features include the texture, edges, and local intensity information of the image, and the deep global features include the semantic association, spatial relationship of anatomical structures, and distribution pattern of multiple lesions in the image.

[0011] The cross-modal fusion processing in step 2.1 adopts a two-branch attention fusion mechanism: the first branch performs intra-modal attention weighting on the shallow local features extracted by the CNN network to highlight the unique local details of each modality; the second branch performs inter-modal attention weighting on the deep global features extracted by the ViT network to highlight the global correlation between different modalities; the output features of the two branches are concatenated and compressed through a fully connected layer to obtain the final multimodal fusion features.

[0012] In the bi-branch attention fusion mechanism, the weighting coefficients of the intra-modal attention weighting are calculated in the following way: Let the shallow local features of the i-th modality extracted by the CNN network be... Then the attention weights of this modality satisfy: in, This is a similarity measure for shallow local features within the i-th mode, calculated using cosine similarity. The local features after intramodal attention weighting are: .

[0013] In the aforementioned dual-branch attention fusion mechanism, the weight matrix for inter-modal attention weighting is calculated in the following manner: Let the deep global feature of the i-th modality extracted by the ViT network be... Correlation weights between different modalities The degree of correlation between the i-th and j-th modes is represented as follows: in, , It is a matrix used to adjust features. It is a scaling factor; softmax is used to ensure that the sum of the weights is 1. The integrated global features after intermodal attention weighting are: .

[0014] Step 2 involves the overall training process of the feature extraction and fusion network, lesion layer localization network, disease classification network, and target detection network, including: Step a: Construct a training dataset, which contains multimodal images of patients with various disease types, including T1W images, T2W images, DWI images, and FLAIR images, as well as corresponding complete annotation information. The complete annotation information includes the target scanning layer where the lesion is located, the disease type label, and the true boundary coordinates of the lesion. Step b: Input the multimodal images from the training dataset into the feature extraction and fusion network to obtain multimodal fusion features; Step c: Input the multimodal fusion features into the lesion layer localization network and output the predicted target scanning layer; extract the fusion features corresponding to the predicted target scanning layer from the multimodal fusion features, and input them into the disease classification network and the target detection network respectively to obtain the predicted lesion type and the predicted lesion boundary coordinates; Step d: Calculate the overall loss value, which is a weighted sum of the lesion layer localization loss, disease classification loss, and lesion boundary coordinate regression loss; Step e: Based on the overall loss value, synchronously update all parameters of the feature extraction and fusion network, the lesion layer localization network, the disease classification network, and the target detection network; Step f: Repeat steps b to e until the overall loss value drops below the preset threshold to complete the training.

[0015] The matching of the corresponding additional examination plan includes: determining the additional examination plan based on the disease type information and a preset rule base; The construction of the rule base includes: establishing a mapping relationship between various disease types and corresponding additional examination plans; the rule base allows administrators to configure, modify, or add mapping relationships.

[0016] The rule base contains preset standard scanning parameters for various additional inspection schemes; The additional examination plan corresponding to the matching also includes: adapting and adjusting the conventional scanning parameters according to the location and size of the lesion.

[0017] The present invention also provides an electronic device for processing head magnetic resonance images, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned method.

[0018] Compared with existing technologies, this invention addresses the pain points of existing magnetic resonance head plain scan diagnosis, such as the limitation of single modality, cumbersome multimodal comparison, frequent rescanning, and reliance on doctor experience, and has the following advantages: (1) This invention can give full play to the imaging advantages of T1W, T2W, FLAIR, and DWI through multimodal feature fusion. Especially for small lesions or atypical lesions (such as early brain tumors and small ischemic lesions), it can enhance feature recognition, quickly complete lesion layer localization, disease classification, and lesion precise localization, and greatly improve the efficiency of clinical diagnosis; (2) Based on the automated abnormal detection results of plain scan multimodal data, it can quickly determine whether rescanning is needed. Combined with the preset rule library, it can accurately match the additional examination plan and provide the rescanning plan prompt information to the doctor or technician in real time. Rescanning can be arranged on the spot immediately, reducing the physical and psychological burden of patients having to go to the hospital multiple times for rescanning, and reducing the resource consumption of the hospital due to repeated scanning. Attached Figure Description

[0019] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A schematic diagram of the overall workflow of this invention; Figure 2 A schematic diagram of the image processing flow in step 2 of this invention; Figure 3 Schematic diagram of model training in this invention; Figure 4 Schematic diagram of image preprocessing in this invention; Figure 5 Schematic diagram of the implementation effect study process of this invention; Figure 6 A schematic diagram illustrating the effects of implementing this invention. Detailed Implementation

[0020] The present disclosure will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present disclosure, but do not limit the present disclosure in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present disclosure. These all fall within the protection scope of the present disclosure.

[0021] This embodiment aims to implement an automated, high-precision system for detecting abnormalities in head magnetic resonance imaging (MRI) images and recommending supplementary examination plans. The core objective is to automatically identify lesions and match the optimal supplementary examination plan using a deep learning model, thereby improving the efficiency and accuracy of clinical diagnosis. Specific implementation details are as follows (see...). Figure 1 ): Step 1: Multimodal Image Acquisition The patient's head was scanned using an MRI scanner. The scanning sequence and parameters strictly followed the standardized clinical procedure to obtain the patient's T1W, T2W, DWI, and FLAIR images.

[0022] Step 2: Anomaly detection based on deep learning models (see...) Figure 2 , Figure 3 , Figure 4 ) This step adopts a progressive four-stage process, from multimodal feature fusion to lesion localization and disease classification, to achieve end-to-end anomaly detection, as detailed below: Step 2.1: Feature Extraction and Fusion A dual-network architecture is adopted, consisting of CNN for shallow local feature extraction and ViT for deep global feature extraction. A dual-branch attention fusion mechanism is used to achieve efficient fusion of cross-modal features, fully utilizing the complementary information of each modality. Specific implementation details are as follows: (1) Shallow local feature extraction (CNN network) ResNet-50 was selected as the backbone model of the CNN network. The last fully connected layer and the classification layer of the original model were removed, and the first four convolutional blocks (conv1-conv4) were retained to extract shallow local features (texture, edges, local intensity) of images of each modality. The specific network configuration is as follows: The conv1 block contains a 7×7 convolutional kernel (stride 2, padding=3) with 64 output channels. It is followed by a batch normalization layer and a ReLU activation function. The purpose of this layer is to initially compress the image size (512×512→256×256) and extract low-frequency edge features. The conv2-conv4 blocks each consist of three bottleneck modules. Each bottleneck module contains a structure of 1×1 convolution (dimensionality reduction) - 3×3 convolution (feature extraction) - 1×1 convolution (dimensionality increase). The number of output channels for conv2, conv3, and conv4 are 64, 128, and 256, respectively. Residual connections are used to avoid the gradient vanishing problem in deep networks and ensure the effective transfer of local detail features. Output features: After processing by the CNN network, the images of each modality output a feature map (number of channels × height × width) with a size of 256×64×64. Among them, the feature map of T1W modality highlights the edges of anatomical structures, the feature map of DWI modality highlights the high signal local area of ​​ischemic lesions, and the feature map of FLAIR modality highlights the texture details of edema areas.

[0023] (2) Deep global feature extraction (ViT network) ViT-B / 16 (Vision Transformer-Base) was selected as the backbone model of the ViT network. The input image size was adjusted to 224×224 pixels (achieved through bilinear interpolation) to extract deep global features (semantic associations, spatial relationships of anatomical structures, and distribution patterns of multiple lesions) from each modality of the image. The specific network configuration is as follows: Patch Embedding layer: The 224×224 image is divided into 14×14 16×16 patches. Each patch is transformed into a 768-dimensional vector through linear projection. At the same time, a learnable class embedding vector (ClassToken) is added for global feature aggregation, and a position embedding vector (Position Embedding) is added to preserve spatial location information. The output dimension is 197×768 (14×14+1=197 vectors, each vector is 768-dimensional). Transformer Encoder: Contains a 12-layer Transformer encoder, each layer consisting of a Multi-Head Self-Attention (MHSA) module and a Multi-Layer Perceptron (MLP) module. The MHSA module contains 12 attention heads, each with a dimension of 64 (768 / 12), capturing global semantic relationships in the image (such as the locational relationship between ischemic lesions and the middle cerebral artery supply area) through a self-attention mechanism. The MLP module contains two fully connected layers with a hidden layer dimension of 3072, using GELU as the activation function. Output features: The category embedding vector output by the last encoder layer of the ViT network is taken as the deep global feature of this modality, with a dimension of 768. For example, the global features of the T1W modality of brain tumor patients include the spatial distribution relationship between tumor enhancement foci and surrounding edema, and the global features of the DWI modality include the association between tumor necrosis foci and diffusion-restricted areas.

[0024] (3) Two-branch attention fusion mechanism A dual-branch attention fusion mechanism is adopted to perform weighted fusion of shallow local features and deep global features separately, solving the problem of complementarity between "local details and global correlation" in multimodal features. The specific implementation is as follows: ① First branch: Intramodal attention weighting (highlighting local details unique to each modality) For the four modalities of shallow local features extracted by CNN, F1, F2, F3, and F4 correspond to T1W, T2W, DWI, and FLAIR, respectively. The weights are calculated through intramodal attention to enhance key local features for disease diagnosis (such as the local signal of ischemic lesions on DWI). The steps are as follows: A. Feature Pooling: Global Average Pooling is performed on the shallow local feature maps of each modality to convert the 256×64×64 feature maps into 256-dimensional vectors; B. Similarity Calculation: Cosine similarity is used to calculate the autocorrelation of features within each modality. This measures the cohesion of the local features of the modality (the higher the cohesion, the more consistent the local details of the modality, and the more critical it is for diagnosis). C. Weight normalization: This is achieved by applying the softmax function to... Normalization is performed to obtain the attention weights for each modality. To ensure the total weights are equal to 1, the formula is: ; D. Feature Weighting: This involves weighting the shallow local feature map of each modality with its corresponding weight. Multiply to obtain the weighted local features. This enables the enhancement of local details in key modes.

[0025] ② Second branch: Intermodal attention weighting (highlighting the global correlation between different modalities) Deep global features extracted from four modalities by ViT Each feature Both belong to the real number space The association weights are calculated through intermodal attention to capture global semantic associations across modalities (such as the spatial overlap between FLAIR edema areas and DWI ischemic lesions). The steps are as follows: A. Feature Projection: Define two learnable projection matrices and The initial values ​​are obtained using a Xavier uniform distribution (to ensure numerical stability during forward and backward propagation); the global features of each mode are... respectively with , Multiply; B. Correlation Weight Calculation: Calculate the correlation weight matrix between modes. ,in The global correlation between the i-th and j-th modes is expressed by the formula: ,in It is a scaling factor; softmax is used to ensure that the sum of the weights is 1. C. Global Feature Integration: The integrated global features after inter-modal attention weighting are: .

[0026] ③ Feature splicing and compression A. Local feature concatenation: The shallow local features weighted by the four modalities are concatenated according to the channel dimension to obtain local fusion features; B. Global Feature Adaptation: The weighted global fusion features are adapted to the dimensions of the local fusion features by passing them through a 1×1 convolutional layer (1024 output channels, 1024 convolutional kernels). C. Final fusion features: The local fusion features and the global fusion features are concatenated according to the channel dimension to obtain a feature map with a size of 2048×64×64. Then, the features are compressed through a fully connected layer (input dimension 2048×64×64, output dimension 1024×64×64) to finally obtain the multimodal fusion features.

[0027] Step 2.2: Localization of the lesion layer A binary classification network (outputting lesion layer or non-lesion layer) is used to screen the scanning layer where the lesion is located. The core is to determine whether there is a lesion in each layer from the multimodal fusion features, thereby reducing the computational load of subsequent tasks and focusing on key layers. The specific process is as follows: (1) Input the multimodal fusion features into the binary classification network and output the lesion layer probability P of each scanning image; (2) Set a probability threshold. If P is greater than the probability threshold, the layer is determined to be a lesion layer; otherwise, it is a non-lesion layer; (3) Output the target scanning layer index and use the fusion features of these target scanning layers for subsequent steps. Step 2.3: Disease Classification A multi-classification network is used to classify various brain diseases. The core principle is to output the probability distribution of each disease based on the fusion features of the target scan layer, thus determining the most likely disease type. In this embodiment, an improved MobileNetV3-Large model is used to adapt to the feature distribution of medical images while maintaining lightweight design. The inference process is as follows: extract the multimodal fusion features of the target scan layer, and take the average value of the features from each layer as the input features (reducing inter-layer noise interference); input the input features into the multi-classification network to output the probability distribution of various diseases; take the disease type corresponding to the maximum probability as the final disease type information, and simultaneously output the confidence level (i.e., the maximum probability value) of that disease.

[0028] As another embodiment, multiple binary classification networks can be used to classify various brain diseases. Each binary classification network corresponds to a disease type. The multimodal fusion features of the target scanning layer are input into multiple binary classification networks to obtain the probability of each disease type.

[0029] Step 2.4: Precise localization of the lesion The YOLOv8-nano target detection network is used to accurately locate the lesion boundary coordinates. The core is to identify the lesion region in the fusion features of the target scanning layer and output the pixel coordinates of the upper left and lower right corners.

[0030] The specific process is as follows: input the fused features of the target scanning layer into the YOLOv8-nano network, the network matches the anchor box that is closest to the size of the lesion; predict the offset of the bounding box relative to the anchor box; calculate the actual boundary coordinates of the lesion based on the initial coordinates and offset of the anchor box; and output the final lesion boundary coordinate information.

[0031] Step 2 involves the overall training process of the feature extraction and fusion network, lesion layer localization network, disease classification network, and target detection network, including: ① Construct a training dataset, which contains multimodal images of patients with various disease types, including T1W images, T2W images, DWI images, and FLAIR images, as well as corresponding complete annotation information. The complete annotation information includes the target scanning layer where the lesion is located, the disease type label, and the true boundary coordinates of the lesion.

[0032] Hundreds of real head MRI scans from hospitals were collected, covering common clinical brain diseases such as ischemic stroke, brain tumors, cerebral hemorrhage, and cerebral atrophy. These images were annotated by physicians, with annotation information including: ① the target scan layer where the lesion is located; ② the disease type label; and ③ the actual boundary coordinates of the lesion. All DICOM format images were converted to 16-bit PNG format, preserving the original image grayscale information. Bilinear interpolation was used to adjust the resolution of all images to 512×512 pixels to avoid size inconsistencies caused by differences in scanning parameters. Image pixels were normalized, mapping pixel values ​​to the [0, 255] interval. The dataset was divided into training, validation, and test sets in a 7:2:1 ratio to ensure that the proportion of each disease type after the division was consistent with the original dataset, avoiding data bias.

[0033] ② Input the multimodal images from the training dataset into the feature extraction and fusion network to obtain multimodal fusion features.

[0034] ③ Input the multimodal fusion features into the lesion layer localization network and output the predicted target scanning layer; extract the fusion features corresponding to the predicted target scanning layer from the multimodal fusion features, and input them into the disease classification network and the target detection network respectively to obtain the predicted lesion type and the predicted lesion boundary coordinates.

[0035] ④ Calculate the overall loss value, which is a weighted sum of the lesion layer localization loss, disease classification loss, and lesion boundary coordinate regression loss.

[0036] ⑤ Based on the overall loss value, simultaneously update all parameters of the feature extraction and fusion network, lesion layer localization network, disease classification network, and target detection network; ⑥ Repeat steps ② to ⑤ until the overall loss value drops below the preset threshold to complete the training.

[0037] Step 3: Additional inspection plan matching and parameter adjustment Based on the disease type, lesion location, and size information obtained in step 2, the system automatically matches additional examination plans using a preset rule base and adjusts routine parameters to ensure the plan's relevance and rationality.

[0038] As an example only, some of the rule base content is shown in the table below: Table 1: This rule base is stored locally, allowing administrators to configure, modify, or add mapping relationships. Furthermore, the standard scanning parameters can be adjusted based on the location and size of the lesion. As an example, the parameter adjustment rules are as follows: Table 2: Step 4: Generation and Display of Visual Prompt Messages A visual interface is constructed to integrate disease type information, lesion location information, and additional examination plans into intuitive prompts, allowing doctors or scanning technicians to quickly view and make decisions. The specific implementation is as follows.

[0039] Study on the effects of this invention: like Figure 5As shown, this study collected two sets of brain magnetic resonance imaging (MRI) data. All data came from adult patients who underwent brain MRI examinations at the Sixth People's Hospital affiliated with Shanghai Jiao Tong University. All MRI scans were performed using a 3.0T clinical MRI scanner with a standard 48-channel head and neck coil. Core sequences included T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), liquid attenuation inversion recovery sequence (FLAIR), and diffusion-weighted imaging (DWI, b=0, 1000 s / mm²). The scanning procedure strictly followed standard clinical brain plain scan protocols. The model was constructed based on brain plain scan data collected from January 2012 to June 2025, and the modeling process strictly followed the above-mentioned scan sequence parameters. The exclusion criteria for this cohort included: 1) positive results irrelevant to the training objective; 2) scan images with severe artifacts; 3) registration failure between multiple sequences; 4) incomplete acquisition of all scan sequences required by the algorithm; and 5) repeated scan data from the same patient. The prospective clinical validation cohort consisted of continuously acquired clinical brain CT scans performed with system assistance between May and August 2025. All scans were performed using a 3.0T clinical MRI scanner with a 48-channel head and neck coil. Exclusion criteria included: 1) scans with severe artifacts; 2) registration errors between multiple sequences; and 3) incomplete acquisition of all scan sequences required by the algorithm. To establish a clinical control group, we retrospectively collected imaging report data from patients who underwent brain magnetic resonance imaging (MRI) between January 2020 and December 2025. Raw MRI images were not included in the analysis. The control group was defined as screened patients who underwent additional MRI or combined MRA+SWI scans within one month of their initial brain CT scan, and whose additional scan decisions conformed to the AGRAI system algorithm logic.

[0040] The study initially screened 33,601 adult patients admitted between December 2012 and December 2025, of whom 14,155 were excluded due to duplicate scans and other quality issues. The remaining data were allocated to a pre-training set (n=18,257), a proprietary model training set (n=1,028), and a retrospective test set (n=161) for model development and validation. For clinical efficacy validation, two independent cohorts were analyzed: a retrospective control cohort comprising 4,479 adult patients identified through radiology report metadata who underwent scans between January 2020 and December 2025; and a prospective clinical validation cohort consisting of 1,738 consecutively enrolled adult patients between May and August 2025, during which the AGRAI system was deployed to assist in real-time scan decisions. After screening according to exclusion criteria (severe artifacts, multiple sequence registration errors, and incomplete acquisition of required sequences), 260 patients in the prospective cohort were removed, resulting in 1,478 patients for efficacy analysis.

[0041] Diagnostic performance of the model on a retrospective test set: The trained model demonstrated high recommendation accuracy on the retrospective test set. For recommending MRA scans, the accuracy and sensitivity were 96.4% (53 / 55) and 93.0% (53 / 56), respectively, with a false negative rate of 5.4% and a false recommendation rate of 3.6% (2 / 55). For MRA + SWI examinations, the accuracy and sensitivity were 95.4% (41 / 43) and 91.1% (41 / 45), respectively, with a false negative rate of 8.9% and a false recommendation rate of 4.7% (2 / 43). The specificity of the test set was 96.7% (58 / 60).

[0042] Diagnostic performance of the model on the prospective test set: In a prospective cohort study, the model using pre-trained weights (pre-trained model) significantly outperformed the model without pre-trained weights (non-pre-trained model) on all diagnostic metrics. For MRA scan recommendations, the pre-trained model achieved an accuracy of 94.2% (193 / 205) and a sensitivity of 91.1% (193 / 212), while the non-pre-trained model achieved 88.2% (186 / 211) and 87.7% (186 / 212), respectively. The false negative rate and false recommendation rate decreased from 12.3% and 11.8% (25 / 211) to 8.9% and 5.8% (12 / 205), respectively. For MRA+SWI sequences, the pre-trained model achieved an accuracy of 94.6% (88 / 93) and a sensitivity of 89.8% (88 / 98), while the non-pre-trained model achieved 92.1% (82 / 89) and 83.7% (82 / 98), respectively. The false negative rate and false recommendation rate decreased from 16.3% and 7.1% (7 / 98) to 10.2% and 5.4% (5 / 93), respectively. The specificity of the pre-trained model was 98.8% (1,154 / 1,168), slightly higher than that of the non-pre-trained model (97.7%, 1,141 / 1,168).

[0043] Inter-observer consistency and diagnostic efficacy among radiologists: This study evaluated inter-observer consistency and diagnostic efficacy for all prospective clinical imaging examinations in middle-aged and junior radiologists. Within-class correlation coefficient (ICC) and diagnostic indicators (accuracy, sensitivity, and missed diagnosis rate) were used for analysis. Results showed that inter-observer consistency was moderately better than average in all assessments. The ICC values ​​for middle-aged radiologists were significantly higher than those for junior radiologists (ICC value for middle-aged group was 0.756, mean of 0.714 for 5 participants; ICC value for junior group was 0.632; LYT was not included in the ICC calculation in the 4-participant analysis). Middle-aged radiologists demonstrated higher accuracy than junior radiologists in both tasks, with the highest accuracy observed in MRA+SWI sequences. The overall accuracy rate was 73.35%±4.18% (intermediate group) compared to 66.96%±9.61% (primary group); the suggested accuracy rates for MRA + SWI were 77.80%±8.99% (intermediate group) and 69.82%±14.92% (primary group), respectively. Intermediate radiologists had higher sensitivity than the primary group, and the missed diagnosis rates for all sequences were correspondingly lower. The sensitivity of MRA sequences was higher than that of MRA + SWI sequences: MRA sensitivity was 68.37%±6.96% (intermediate group) compared to 61.48%±8.84% (primary group), while the sensitivity of MRA + SWI was only 38.01%±11.59% (intermediate group) and 26.79%±16.61% (primary group). The false negative rate of MRA+SWI sequences was the highest among all sequences, reaching 61.99%±11.59% (intermediate group) and 73.21%±16.61% (primary group).

[0044] Clinical efficacy analysis of the supplementary testing protocol recommendation system: In a 5-year retrospective clinical efficacy analysis, a total of 4,479 patients were included. Among them, 95.08% of patients underwent only a single scan (4,259 / 4,479); 4.40% (197 / 4,479) and 0.52% (23 / 4,479) of patients underwent 2 and ≥3 scans, respectively.

[0045] like Figure 6 As shown, the supplemental testing recommendation system significantly reduced the time from plain scan to final diagnosis. The median time from plain scan to a single plain scan report (Group A, equivalent to the system's diagnostic time) was 14.22 hours (Q1=3.75 hours, Q3=19.82 hours), while the median time from plain scan to the final report requiring supplemental scans was 91.35 hours (Group B, Q1=61.95 hours, Q3=165.40 hours). The difference between the two groups was statistically significant (p<0.001). After applying this system, the average time for patients to obtain a final diagnosis was reduced by 82%.

[0046] The specific embodiments of this disclosure have been described above. It should be understood that this disclosure is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the substantive content of this disclosure. The above-described preferred features can be used in any combination without conflict.

Claims

1. A recommended method for supplementary head magnetic resonance imaging (MRI) examination, characterized in that, Includes the following steps: Step 1: Obtain multimodal images from the plain magnetic resonance imaging (MRI) scan of the subject's head, including T1W images, T2W images, DWI images, and FLAIR images; Step 2: Based on the multimodal image, perform anomaly detection using a deep learning model to obtain the subject's disease type information and lesion location information; the anomaly detection using a deep learning model specifically includes the following steps: Step 2.1: Feature Extraction and Fusion: Features are extracted for each modality of the multimodal image to obtain shallow and deep features corresponding to each modality. Cross-modal fusion processing is then performed on these shallow and deep features to obtain multimodal fused features. Specifically, a CNN network is used to extract shallow local features, and a ViT network is used to extract deep global features. The shallow local features include image texture, edges, and local intensity information, while the deep global features include semantic associations, spatial relationships of anatomical structures, and multi-lesion distribution patterns. The cross-modal fusion processing employs a two-branch attention fusion mechanism: the first branch performs intra-modal attention weighting on the shallow local features extracted by the CNN network, highlighting the unique local details of each modality; the second branch performs inter-modal attention weighting on the deep global features extracted by the ViT network, highlighting the global associations between different modalities. The output features of the two branches are concatenated and compressed through a fully connected layer to obtain the final multimodal fused features. In the bi-branch attention fusion mechanism, the weighting coefficients of the intra-modal attention weighting are calculated in the following way: Let the shallow local features of the i-th modality extracted by the CNN network be... Then the attention weights of this modality satisfy: in, This is a similarity measure for shallow local features within the i-th modality, calculated using cosine similarity; the local features after in-modality attention weighting are: ; Step 2.2: Lesion layer localization: Based on the multimodal fusion features, a classification network is used to determine whether there are lesions in each layer of the scan image, and the target scan layer where the lesion is located is selected. Step 2.3: Disease Classification: Extract the multimodal fusion features corresponding to the target scanning layer, input them into the disease classification network, output the probability distributions corresponding to various disease types, and obtain the disease type information; Step 2.4: Precise localization of lesion location: Based on the multimodal fusion features of the target scanning layer, the boundary coordinate information of the lesion is output through the target detection network to obtain the lesion location information; Step 3: Determine whether additional examinations are needed based on the disease type information and / or lesion location information. If additional examinations are needed, match the corresponding additional examination plan based on the disease type information and / or lesion location information. Step 4: Generate prompt information including the disease type information, the lesion location information, and the additional examination plan.

2. The method according to claim 1, characterized in that, In the aforementioned dual-branch attention fusion mechanism, the weight matrix for inter-modal attention weighting is calculated in the following manner: Let the deep global feature of the i-th modality extracted by the ViT network be... Correlation weights between different modalities The degree of correlation between the i-th and j-th modes is represented as follows: in, , It is a matrix used to adjust features. It is a scaling factor; softmax is used to ensure that the sum of the weights is 1. The integrated global features after intermodal attention weighting are: .

3. The method according to claim 1, characterized in that, Step 2 involves the overall training process of the feature extraction and fusion network, lesion layer localization network, disease classification network, and target detection network, including: Step a: Construct a training dataset, which contains multimodal images of patients with various disease types, including T1W images, T2W images, DWI images, and FLAIR images, as well as corresponding complete annotation information. The complete annotation information includes the target scanning layer where the lesion is located, the disease type label, and the true boundary coordinates of the lesion. Step b: Input the multimodal images from the training dataset into the feature extraction and fusion network to obtain multimodal fusion features; Step c: Input the multimodal fusion features into the lesion layer localization network and output the predicted target scanning layer; extract the fusion features corresponding to the predicted target scanning layer from the multimodal fusion features, and input them into the disease classification network and the target detection network respectively to obtain the predicted lesion type and the predicted lesion boundary coordinates; Step d: Calculate the overall loss value, which is a weighted sum of the lesion layer localization loss, disease classification loss, and lesion boundary coordinate regression loss; Step e: Based on the overall loss value, synchronously update all parameters of the feature extraction and fusion network, the lesion layer localization network, the disease classification network, and the target detection network; Step f: Repeat steps b to e until the overall loss value drops below the preset threshold to complete the training.

4. The method according to claim 1, characterized in that, The matching of the corresponding additional examination plan includes: determining the additional examination plan based on the disease type information and a preset rule base; The construction of the rule base includes: establishing a mapping relationship between various disease types and corresponding additional examination plans; the rule base allows administrators to configure, modify, or add mapping relationships.

5. The method according to claim 4, characterized in that, The rule base contains preset standard scanning parameters for various additional inspection schemes; The additional examination plan corresponding to the matching also includes: adapting and adjusting the conventional scanning parameters according to the location and size of the lesion.

6. An electronic device for processing head magnetic resonance imaging, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Medical diagnosis intelligent decision-making system based on multi-modal data fusion

    CN119495423A

  • Intelligent detection method for brain tumor focus area

    CN120689354A