Large language model coronary artery CTA image analysis method based on multi-modal fusion

By constructing image vision encoder and bridging module in large language models, the problem of difficulty in aligning visual feature information and text feature information in the prior art is solved, efficient analysis and report generation of coronary CTA images are realized, and the performance and efficiency of the model in the field of cardiac medicine are improved.

CN119941698AActive Publication Date: 2025-05-06QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510081910.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

When analyzing coronary CTA images, existing large language models are difficult to effectively align visual feature information and text feature information, resulting in insufficient performance in the field of cardiac medicine, and are difficult and slow in training.

Method used

A large language model coronary CTA image analysis method based on multimodal fusion is adopted. By constructing an image visual encoder and a bridge module, the image feature information is extracted and aligned with the dimensions embedded in the space of the large language model to generate an information report of the coronary CTA image.

Benefits of technology

It improves the accuracy and efficiency of coronary CTA image information reporting, reduces the false positive rate, enhances the model's understanding and analysis ability of medical images, improves its response ability when facing complex medical problems, and slows down doctor pressure and doctor-patient communication costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941698A_ABST
    Figure CN119941698A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model coronary artery CTA image analysis method based on multi-modal fusion, and relates to the technical field of large language models. A visual image feature vector is mapped into a visual feature vector with the same dimension as a large language model embedding space through a simple linear projection layer, so that visual feature information and text feature information can be aligned, the accuracy and efficiency of coronary artery CTA image information report generation can be improved, and the false positive rate is reduced. The understanding and analysis ability of the model on the medical image is enhanced, the response ability of the model in the face of complex medical problems is improved, and the pressure of doctors and the doctor-patient communication cost are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and in particular to a large language model coronary artery CTA image analysis method based on multimodal fusion. Background Art

[0002] Coronary CTA, or Coronary Computed Tomography Angiography, is a non-invasive coronary artery imaging technique. Coronary CTA can clearly show vascular lesions, such as atherosclerosis, aneurysms, thrombosis, etc., and can accurately measure the size, morphology and location of diseased blood vessels, providing important basis for clinical diagnosis and treatment. Traditional machine learning methods often rely on manual features to analyze images, which cannot adapt to the high spatial correlation features of coronary CTA images. The use of multimodal large language models can not only identify and process similar images but also generate meaningful reports. An important research area in natural language processing (NLP) and computer vision (CV) is to explore the learning technology of large language vision models (LLVM), which aims to bridge the gap between visual information and text information, and enable machines to understand and generate content that combines the two modes by aligning the features of the two modes. For example, a discrete tagger is used to directly encode an image or video into a series of tags, and these visual (and text) tags are processed by a multimodal Transformer. Finally, with the help of appropriate decoders or diffusion designs, multimodal outputs can be directly generated. However, this mode usually trains models for the generality of image feature recognition and lacks the ability to understand the complexity of images in the medical field, resulting in poor performance in specific fields such as cardiology. In addition, they usually need to be trained from scratch, which is difficult and slow. Summary of the invention

[0003] In order to overcome the deficiencies of the above technologies, the present invention provides a method for effectively solving the problem of difficulty in aligning visual feature information and text feature information, and can analyze and answer open questions about coronary artery CTA images.

[0004] The technical solution adopted by the present invention to overcome the technical problems is:

[0005] A method for analyzing coronary artery CTA images based on a large language model with multimodal fusion, comprising:

[0006] a) obtaining a coronary artery CTA image X;

[0007] b) Preprocessing the coronary artery CTA image X to obtain a preprocessed coronary artery CTA image X 2 ;

[0008] c) Establish an image visual encoder consisting of a feature extraction module, a residual module, and a multi-head attention mechanism module;

[0009] d) The preprocessed coronary artery CTA image X 2 Input to the feature extraction module of the image visual encoder, and output the feature map Coronary_L 1 ;

[0010] e) Coronary_L 1 Input into the residual module of the image visual encoder, and output the feature map Coronary_L 2 ;

[0011] f) Coronary_L 2 Input into the multi-head attention mechanism module of the image visual encoder, and output the high-dimensional image feature vector Coronary_L 3 ;

[0012] g) Establish a visual and language feature bridging module consisting of a residual module and a dilated convolution module;

[0013] h) Transform the high-dimensional image feature vector Coronary_L 3 Input into the residual module of the visual and language feature bridge module, and output feature maps bridge_res_1, bridge_res_2, and bridge_res_3;

[0014] i) Coronary_L 1 、Characteristic graph Coronary_L 2 , high-dimensional image feature vector Coronary_L 3 , feature maps bridge_res_1, bridge_res_2, and bridge_res_3 are input into the dilated convolution module of the visual and language feature bridge module, and the visual feature vector Linear_C is output;

[0015] j) Input the visual feature vector Linear_C into the trained large language model to obtain the information report of the coronary artery CTA image X.

[0016] Further, step b) comprises the following steps:

[0017] b-1) Use the torchvision.transforms module in PyTorch to convert the coronary artery CTA image X into a coronary artery CTA image X in Tensor format 1 ;

[0018] b-2) Use the transforms.Normalize() function in PyTorch to transform the coronary artery CTA image X 1 Standardization is performed to obtain the preprocessed coronary artery CTA image X 2 .

[0019] Further, step d) comprises the following steps:

[0020] d-1) The feature extraction module of the image visual encoder is composed of a first basic convolution module, a second basic convolution module, and a third basic convolution module;

[0021] d-2) The first basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. The preprocessed coronary artery CTA image X 2 Input into the first basic convolution module and output the feature map

[0022] d-3) The second basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. Input into the second basic convolution module and output the feature map

[0023] d-4) The third basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. Input into the third basic convolution module, and output the feature map Coronary_L 1 .

[0024] Further, step e) comprises the following steps:

[0025] e-1) The residual module of the image visual encoder is composed of a first basic convolution module and a second basic convolution module;

[0026] e-2) The first basic convolution module of the residual module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. 1 Input into the first basic convolution module of the residual module, and output the feature map res_f 1 ;

[0027] e-3) The second basic convolution module of the residual module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. 1Input into the second basic convolution module of the residual module, and output the feature map res_f 2 ;

[0028] e-4) The feature map res_f 1 、Feature map res_f 2 、Characteristic graph Coronary_L 1 Perform the addition operation to obtain the feature map Coronary_L 2 .

[0029] Further, step f) comprises the following steps:

[0030] f-1) The multi-head attention mechanism module of the image visual encoder consists of a Dropout layer, a BN layer, an attention mechanism, and a Relu activation function;

[0031] f-2) Coronary_L 2 Input into the multi-head attention mechanism module in turn through the Dropout layer and the BN layer, and output the feature map attention_C 1 ;

[0032] f-3) The feature map attention_C 1 Input into the attention mechanism of the multi-head attention mechanism module, and output the feature map attention_C 2 ;

[0033] f-4) The feature map attention_C 2 Input into the Relu activation function of the multi-head attention mechanism module, and output the high-dimensional image feature vector Coronary_L 3 .

[0034] Further, step h) comprises the following steps:

[0035] h-1) The residual module of the visual and language feature bridge module is composed of the first branch and the second branch;

[0036] h-2) The first branch of the residual module is composed of a three-dimensional convolutional layer, a BN layer, and a Relu activation function, which converts the high-dimensional image feature vector Coronary_L 3 Input into the first branch of the residual module, and output the feature map res_f 3 ;

[0037] h-3) The second branch of the residual module consists of a three-dimensional convolutional layer, which transforms the high-dimensional image feature vector Coronary_L 3 Input into the second branch of the residual module, and output the feature map res_f 4;

[0038] h-4) The feature map res_f 3 With the feature map res_f 4 Perform addition operation to obtain the feature map RES_Cor;

[0039] h-5) Through the formula bridge_res_1 = α 1 Coronary_L 1 +β 1 RES_Cor calculation feature map Coronary_L 1 The weighted sum of the feature map RES_Cor and the feature map bridge_res_1, where α 1 and β 1 are weights, α 1 =0.3,β 1 =0.7;

[0040] h-6) Coronary_L 2 Multiply it with the feature map RES_Cor to obtain the feature map bridge_res_2;

[0041] h-7) Through the formula bridge_res_3 = α 2 Coronary_L 3 +β 2 RES_Cor calculates the high-dimensional image feature vector Coronary_L 3 The weighted sum of the feature map RES_Cor and the feature map bridge_res_3, where α 2 and β 2 are weights, α 2 =0.7,β 2 =0.3.

[0042] Further, step i) comprises the following steps:

[0043] i-1) The dilated convolution module of the visual and language feature bridging module consists of a first three-dimensional convolution layer, a first dilated convolution layer, a second three-dimensional convolution layer, a second dilated convolution layer, a Relu activation function, a maximum pooling layer, and a linear transformation layer;

[0044] i-2) The high-dimensional image feature vector Coronary_L 3 The first three-dimensional convolutional layer, the first dilated convolutional layer, the second three-dimensional convolutional layer, the second dilated convolutional layer, and the Relu activation function are sequentially input into the dilated convolutional module, and the feature map Conv_cor is output. The feature map Conv_cor is added to the feature map RES_Cor to obtain the feature map RES_CONV_Cor.

[0045] i-3) Input the feature map RES_CONV_Cor into the maximum pooling layer of the dilated convolution module to obtain the compressed feature map RES_CONV_Cor_1;

[0046] i-4) Input the compressed feature map RES_CONV_Cor_1 into the first branch of the residual module of the visual and language feature bridge module, and output the feature map RES_f 1 , the compressed feature map RES_CONV_Cor_1 is input into the second branch of the residual module of the visual and language feature bridge module, and the output is the feature map RES_f 2 , the feature map RES_f 1 With the characteristic graph RES_f 2 Perform addition operation to obtain feature map RES_CONV_Cor_1.1;

[0047] i-5) Input the compressed feature map RES_CONV_Cor_1 into the hole convolution module, which consists of the first three-dimensional convolution layer, the first hole convolution layer, the second three-dimensional convolution layer, the second hole convolution layer, and the Relu activation function, and output the feature map Conv_cor′. Add the feature map Conv_cor′ to the feature map RES_Cor to obtain the feature map RES_CONV_Cor_1.2;

[0048] i-6) Add the feature map RES_CONV_Cor_1.1, the feature map RES_CONV_Cor_1.2, and the compressed feature map RES_CONV_Cor_1 to obtain the feature map R_C_C_P;

[0049] i-7) Through the formula bridge_dilated_1 = α 3 Coronary_L 1 +β 3 R_C_C_P Calculation Characteristic Map Coronary_L 1 The feature map bridge_dilated_1 is the weighted sum of the feature map R_C_C_P, where α 3 and β 3 are weights, α 3 =0.3,β 3 =0.7;

[0050] i-8) Combine the feature graph R_C_C_P with the feature graph Coronary_L 2 Perform multiplication operation to obtain feature map bridge_dilated_2;

[0051] i-9) Through the formula bridge_dilated_3 = α 4 Coronary_L 3 +β 4 R_C_C_P calculates the high-dimensional image feature vector Coronary_L 3 The feature map bridge_dilated_3 is the weighted sum of the feature map R_C_C_P, where α 4 and β 4 are weights, α 4 =0.3,β 4 =0.7;

[0052] i-10) multiplying the feature map bridge_res_1 and the feature map bridge_res_2 to obtain the feature map Bridge_Res_1, and adding the feature map Bridge_Res_1 and the feature map bridge_res_3 to obtain the feature map Bridge_Res;

[0053] i-11) by formula

[0054] Calculate the feature map Bridge_Dilated which is the weighted sum of the feature map bridge_dilated_1, the feature map bridge_dilated_2 and the feature map bridge_dilated_3, where α 5 , β 5 , γ 1 are weights, α 5 =0.2,β 5 =0.3,γ 1 =0.5;

[0055] i-12) Add the feature map Bridge_Res and the feature map Bridge_Dilated to obtain the feature map Bridge_output;

[0056] i-13) Input the feature map Bridge_output into the linear transformation layer of the hole convolution module, and output the visual feature vector Linear_C.

[0057] Preferably, the large language model in step j) is a Vicuna1.5-7B model.

[0058] Furthermore, the Vicuna1.5-7B model is trained using the ROCOv2 dataset to obtain a trained large language model.

[0059] The beneficial effects of the present invention are: by constructing a visual encoder and a bridge module to extract image feature information, and then mapping the visual image feature vector into a visual feature vector with the same dimension as the large language model embedding space through a simple linear projection layer, this can not only align the visual feature information with the text feature information, but also improve the accuracy and efficiency of the generation of coronary CTA image information reports and reduce the false positive rate. It enhances the model's ability to understand and analyze medical images, improves the model's ability to respond to complex medical problems, and effectively reduces the pressure on doctors and the cost of communication between doctors and patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flow chart of the method of the present invention;

[0061] Figure 2 is a structural diagram of an image visual encoder of the present invention;

[0062] Figure 3 It is a structural diagram of the bridge module of the present invention. DETAILED DESCRIPTION

[0063] The following is combined with Figure 1 , Attachment Figure 2 , Attachment Figure 3 The present invention is further described.

[0064] A method for analyzing coronary artery CTA images based on a large language model with multimodal fusion, comprising:

[0065] a) Obtain a coronary artery CTA image X.

[0066] b) Preprocessing the coronary artery CTA image X to obtain a preprocessed coronary artery CTA image X 2 c) Establish an image visual encoder consisting of a feature extraction module, a residual module, and a multi-head attention mechanism module. d) Transform the preprocessed coronary artery CTA image X 2 Input to the feature extraction module of the image visual encoder, and output the feature map Coronary_L 1 .

[0067] e) Coronary_L 1 Input into the residual module of the image visual encoder, and output the feature map Coronary_L 2 .

[0068] f) Coronary_L 2 Input into the multi-head attention mechanism module of the image visual encoder, and output the high-dimensional image feature vector Coronary_L 3 .

[0069] g) Establish a visual and language feature bridging module consisting of a residual module and a dilated convolution module.

[0070] h) Transform the high-dimensional image feature vector Coronary_L 3 Input into the residual module of the visual and language feature bridge module, and output the feature maps bridge_res_1, bridge_res_2, and bridge_res_3.

[0071] i) Coronary_L 1 、Characteristic graph Coronary_L 2 , high-dimensional image feature vector Coronary_L 3 , feature map bridge_res_1, feature map bridge_res_2, and feature map bridge_res_3 are input into the hole convolution module of the visual and language feature bridging module, and the visual feature vector Linear_C is output.

[0072] j) Input the visual feature vector Linear_C into the trained large language model to obtain the information report of the coronary artery CTA image X.

[0073] By constructing a medical image encoder network and a bridge module, and using the self-attention mechanism to extract features from coronary CT angiography images, the spatial correlation of coronary CTA images can be effectively adapted, the false positive rate can be effectively reduced, and an effective connection between the visual encoder and the large language model can be achieved. This method can analyze and answer open questions about coronary CTA images, achieve accurate analysis of coronary CTA, enhance the model's ability to understand and analyze medical images, improve the model's ability to respond to complex medical problems, and reduce the pressure on doctors and the cost of communication between doctors and patients.

[0074] In one embodiment of the present invention, step b) comprises the following steps:

[0075] b-1) Use the torchvision.transforms module in PyTorch to convert the coronary artery CTA image X into a coronary artery CTA image X in Tensor format 1 .

[0076] b-2) Use the transforms.Normalize() function in PyTorch to transform the coronary artery CTA image X 1 Standardization is performed to obtain the preprocessed coronary artery CTA image X 2 .

[0077] In one embodiment of the present invention, step d) comprises the following steps:

[0078] d-1) The feature extraction module of the image visual encoder is composed of a first basic convolution module, a second basic convolution module, and a third basic convolution module.

[0079] d-2) The first basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. The preprocessed coronary artery CTA image X 2 Input into the first basic convolution module and output the feature map The convolution kernel size of the three-dimensional convolution layer of the first basic convolution module is 3×3×3, the step size is 1, and the padding is 2. The drop probability of the Dropout layer of the first basic convolution module is 0.5.

[0080] d-3) The second basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. Input into the second basic convolution module and output the feature map The convolution kernel size of the three-dimensional convolution layer of the second basic convolution module is 3×3×3, the step size is 1, and the padding is 2. The drop probability of the Dropout layer of the second basic convolution module is 0.5.

[0081] d-4) The third basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. Input into the third basic convolution module, and output the feature map Coronary_L 1 The convolution kernel size of the three-dimensional convolution layer of the third basic convolution module is 3×3×3, the stride is 1, and the padding is 2. The drop probability of the Dropout layer of the third basic convolution module is 0.5.

[0082] In one embodiment of the present invention, step e) comprises the following steps:

[0083] e-1) The residual module of the image visual encoder is composed of a first basic convolution module and a second basic convolution module.

[0084] e-2) The first basic convolution module of the residual module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. 1 Input into the first basic convolution module of the residual module, and output the feature map res_f 1The convolution kernel size of the three-dimensional convolution layer of the first basic convolution module of the residual module is 3×3×3, the stride is 1, the padding is 2, and the dropout probability of the Dropout layer is 0.3.

[0085] e-3) The second basic convolution module of the residual module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. 1 Input into the second basic convolution module of the residual module, and output the feature map res_f 2 The convolution kernel size of the three-dimensional convolution layer of the second basic convolution module of the residual module is 3×3×3, the stride is 1, the padding is 2, and the dropout probability of the Dropout layer is 0.3.

[0086] e-4) The feature map res_f 1 、Feature map res_f 2 、Characteristic graph Coronary_L 1 Perform the addition operation to obtain the feature map Coronary_L 2 .

[0087] In one embodiment of the present invention, step f) comprises the following steps:

[0088] f-1) The multi-head attention mechanism module of the image visual encoder consists of a Dropout layer, a BN layer, an attention mechanism, and a Relu activation function.

[0089] f-2) Coronary_L 2 Input into the multi-head attention mechanism module in turn through the Dropout layer and the BN layer, and output the feature map attention_C 1 .

[0090] f-3) The feature map attention_C 1 Input into the attention mechanism of the multi-head attention mechanism module, and output the feature map attention_C 2 .

[0091] f-4) The feature map attention_C 2 Input into the Relu activation function of the multi-head attention mechanism module, and output the high-dimensional image feature vector Coronary_L 3 .

[0092] In one embodiment of the present invention, step h) comprises the following steps:

[0093] h-1) The residual module of the visual and language feature bridge module consists of a first branch and a second branch.

[0094] h-2) The first branch of the residual module is composed of a three-dimensional convolutional layer, a BN layer, and a Relu activation function, which converts the high-dimensional image feature vector Coronary_L 3 Input into the first branch of the residual module, and output the feature map res_f 3 The convolution kernel size of the 3D convolutional layer of the first branch is 5×5×5, the stride is 2, and the padding is 1.

[0095] h-3) The second branch of the residual module consists of a three-dimensional convolutional layer, which transforms the high-dimensional image feature vector Coronary_L 3 Input into the second branch of the residual module, and output the feature map res_f 4 The convolution kernel size of the 3D convolution layer of the second branch is 3×3×3, the stride is 2, and the padding is 1.

[0096] h-4) The feature map res_f 3 With the feature map res_f 4 Perform the addition operation to obtain the characteristic graph RES_Cor. h-5) Through the formula bridge_res_1 = α 1 Coronary_L 1 +β 1 RES_Cor calculation feature map Coronary_L 1 The weighted sum of the feature map RES_Cor and the feature map bridge_res_1, where α 1 and β 1 are weights, α 1 =0.3,β 1 =0.7.

[0097] h-6) Coronary_L 2 Multiply it with the feature map RES_Cor to obtain the feature map bridge_res_2.

[0098] h-7) Through the formula bridge_res_3 = α 2 Coronary_L 3 +β 2 RES_Cor calculates the high-dimensional image feature vector Coronary_L 3 The weighted sum of the feature map RES_Cor and the feature map bridge_res_3, where α 2 and β 2 are weights, α 2 =0.7,β 2 =0.3.

[0099] In one embodiment of the present invention, step i) comprises the following steps:

[0100] i-1) The dilated convolution module of the visual and language feature bridging module consists of the first three-dimensional convolution layer, the first dilated convolution layer, the second three-dimensional convolution layer, the second dilated convolution layer, the Relu activation function, the maximum pooling layer, and the linear transformation layer. The convolution kernel size of the first three-dimensional convolution layer is 3×3×3, the step size is 2, and the padding is 1. The convolution kernel size of the first dilated convolution layer is 5, the expansion rate is 3, the step size is 1, and the padding is 3. The convolution kernel size of the second three-dimensional convolution layer is 3×3×3, the step size is 2, and the padding is 1. The convolution kernel size of the second dilated convolution layer is 5, the expansion rate is 3, the step size is 1, and the padding is 3.

[0101] i-2) The high-dimensional image feature vector Coronary_L 3 The first three-dimensional convolutional layer, the first dilated convolutional layer, the second three-dimensional convolutional layer, the second dilated convolutional layer, and the Relu activation function are sequentially input into the dilated convolution module, and the feature map Conv_cor is output. The feature map Conv_cor is added to the feature map RES_Cor to obtain the feature map RES_CONV_Cor.

[0102] i-3) Input the feature map RES_CONV_Cor into the maximum pooling layer of the dilated convolution module to obtain the compressed feature map RES_CONV_Cor_1.

[0103] i-4) Input the compressed feature map RES_CONV_Cor_1 into the first branch of the residual module of the visual and language feature bridge module, and output the feature map RES_f 1 , the compressed feature map RES_CONV_Cor_1 is input into the second branch of the residual module of the visual and language feature bridge module, and the output is the feature map RES_f 2 , the feature map RES_f 1 With the characteristic graph RES_f 2 The addition operation is performed to obtain the feature map RES_CONV_Cor_1.1.

[0104] i-5) The compressed feature map RES_CONV_Cor_1 is input into the hole convolution module in sequence, which consists of the first three-dimensional convolution layer, the first hole convolution layer, the second three-dimensional convolution layer, the second hole convolution layer, and the Relu activation function, and the feature map Conv_cor′ is output. The feature map Conv_cor′ is added to the feature map RES_Cor to obtain the feature map RES_CONV_Cor_1.2.

[0105] i-6) Add the feature map RES_CONV_Cor_1.1, the feature map RES_CONV_Cor_1.2, and the compressed feature map RES_CONV_Cor_1 to obtain the feature map R_C_C_P.

[0106] i-7) Through the formula bridge_dilated_1 = α 3 Coronary_L 1 +β 3 R_C_C_P Calculation Characteristic Map Coronary_L 1 The feature map bridge_dilated_1 is the weighted sum of the feature map R_C_C_P, where α 3 and β 3 are weights, α 3 =0.3,β 3 =0.7.

[0107] i-8) Combine the feature graph R_C_C_P with the feature graph Coronary_L 2 Perform a multiplication operation to obtain the feature map bridge_dilated_2.

[0108] i-9) Through the formula bridge_dilated_3 = α 4 Coronary_L 3 +β 4 R_C_C_P calculates the high-dimensional image feature vector Coronary_L 3 The feature map bridge_dilated_3 is the weighted sum of the feature map R_C_C_P, where α 4 and β 4 are weights, α 4 =0.3,β 4 =0.7.

[0109] i-10) Multiply the feature map bridge_res_1 and the feature map bridge_res_2 to obtain the feature map Bridge_Res_1, and add the feature map Bridge_Res_1 and the feature map bridge_res_3 to obtain the feature map Bridge_Res.

[0110] i-11) by formula

[0111] Calculate the feature map Bridge_Dilated which is the weighted sum of the feature map bridge_dilated_1, the feature map bridge_dilated_2 and the feature map bridge_dilated_3, where α 5 , β 5 , γ 1 are weights, α 5 =0.2,β 5 =0.3,γ 1 =0.5.

[0112] i-12) Add the feature map Bridge_Res and the feature map Bridge_Dilated to obtain the feature map Bridge_output.

[0113] i-13) Input the feature map Bridge_output into the linear transformation layer of the hole convolution module, and output the visual feature vector Linear_C. The dimension of the visual feature vector Linear_C is 512. In one embodiment of the present invention, the large language model in step j) is the Vicuna1.5-7B model. In this embodiment, the Vicuna1.5-7B model is trained using the ROCOv2 data set to obtain a trained large language model. The ROCOv2 public data set contains a large number of X-ray images of the human body and text descriptions of the images. By training the Vicuna1.5-7B model therewith, the trained Vicuna1.5-7B model can output accurate information reports for the visual feature vector Linear_C.

[0114] In order to verify the reliability of this method, the ROUGE score is used to compare the model proposed by this method with the existing baseline model MINIGPT-4 as shown in Table 1:

[0115] Comparing our model with the baseline model using the Rouge score on the ARCADE dataset

[0116] Model\Rating Criteria R-1 R-2 RL MINIGPT-4 Baseline Model 0.3008 0.0721 0.1496 Method of the present invention 0.3209 0.0808 0.1775

[0117] In our experimental method, we compared the most advanced visual language model MINIGPT-4. The ROUGE score calculates precision, recall and F1 score, taking into account the presence and order of n-grams (continuous word sequences). The higher the ROUGE score, the better the match between the generated text and the reference text, the higher the accuracy of the generated text, and the higher the quality and coherence. In terms of R-1 score, our method outperforms MINIGPT-4 with an absolute gain of 2%. This shows that our method has certain advantages, which are significant improvements in accuracy.

[0118] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for coronary artery CTA image analysis based on a large language model of multimodal fusion, characterized in that: include: a) obtaining a coronary artery CTA image X; b) preprocessing the coronary artery CTA image X to obtain a preprocessed coronary artery CTA image X2; c) Establish an image visual encoder consisting of a feature extraction module, a residual module, and a multi-head attention mechanism module; d) inputting the preprocessed coronary artery CTA image X2 into the feature extraction module of the image visual encoder, and outputting the feature map Coronary_L1; e) Input the feature map Coronary_L1 into the residual module of the image visual encoder, and output the feature map Coronary_L2; f) Input the feature map Coronary_L2 into the multi-head attention mechanism module of the image visual encoder, and output the high-dimensional image feature vector Coronary_L3; g) Establish a visual and language feature bridging module consisting of a residual module and a dilated convolution module; h) Input the high-dimensional image feature vector Coronary_L3 into the residual module of the visual and language feature bridge module, and output feature maps bridge_res_1, bridge_res_2, and bridge_res_3; i) Input feature map Coronary_L1, feature map Coronary_L2, high-dimensional image feature vector Coronary_L3, feature map bridge_res_1, feature map bridge_res_2, feature map bridge_res_3 into the hole convolution module of the visual and language feature bridge module, and output the visual feature vector Linear_C; j) Input the visual feature vector Linear_C into the trained large language model to obtain the information report of the coronary artery CTA image X.

2. The method for coronary artery CTA image analysis based on multimodal fusion and large language model according to claim 1, characterized in that: Step b) comprises the following steps: b-1) Use the torchvision.transforms module in PyTorch to convert the coronary artery CTA image X into a coronary artery CTA image X1 in Tensor format; b-2) Use the transforms.Normalize() function in PyTorch to normalize the coronary artery CTA image X1 to obtain the preprocessed coronary artery CTA image X2.

3. The method for coronary artery CTA image analysis based on multimodal fusion and large language model according to claim 1, characterized in that: Step d) comprises the following steps: d-1) The feature extraction module of the image visual encoder is composed of a first basic convolution module, a second basic convolution module, and a third basic convolution module; d-2) The first basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer in sequence. The preprocessed coronary artery CTA image X2 is input into the first basic convolution module, and the feature map is output. d-3) The second basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. Input into the second basic convolution module and output the feature map d-4) The third basic convolution module of the feature extraction module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer. Input into the third basic convolution module and output the feature map Coronary_L1.

4. The method for coronary artery CTA image analysis based on multimodal fusion and large language model according to claim 1, characterized in that: Step e) comprises the following steps: e-1) The residual module of the image visual encoder is composed of a first basic convolution module and a second basic convolution module; e-2) The first basic convolution module of the residual module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer in sequence. The feature map Coronary_L1 is input into the first basic convolution module of the residual module, and the feature map res_f1 is output; e-3) The second basic convolution module of the residual module is composed of a three-dimensional convolution layer, a BN layer, a Relu activation function, and a Dropout layer in sequence. The feature map Coronary_L1 is input into the second basic convolution module of the residual module, and the feature map res_f2 is output; e-4) Add the feature map res_f1, the feature map res_f2, and the feature map Coronary_L1 to obtain the feature map Coronary_L2.

5. The method for coronary artery CTA image analysis based on multimodal fusion and large language model according to claim 1, characterized in that: Step f) comprises the following steps: f-1) The multi-head attention mechanism module of the image visual encoder consists of a Dropout layer, a BN layer, an attention mechanism, and a Relu activation function; f-2) Input the feature map Coronary_L2 into the multi-head attention mechanism module in turn through the Dropout layer and the BN layer, and output the feature map attention_C1; f-3) Input the feature map attention_C1 into the attention mechanism of the multi-head attention mechanism module, and output the feature map attention_C2; f-4) Input the feature map attention_C2 into the Relu activation function of the multi-head attention mechanism module, and output the high-dimensional image feature vector Coronary_L3.

6. The method for coronary artery CTA image analysis based on multimodal fusion and large language model according to claim 5, characterized in that: Step h) comprises the following steps: h-1) The residual module of the visual and language feature bridge module is composed of the first branch and the second branch; h-2) The first branch of the residual module is composed of a three-dimensional convolutional layer, a BN layer, and a Relu activation function in sequence. The high-dimensional image feature vector Coronary_L3 is input into the first branch of the residual module, and the feature map res_f3 is output; h-3) The second branch of the residual module is composed of a three-dimensional convolutional layer. The high-dimensional image feature vector Coronary_L3 is input into the second branch of the residual module, and the feature map res_f4 is output; h-4) Add the feature map res_f3 and the feature map res_f4 to obtain the feature map RES_Cor; h-5) Calculate the feature map bridge_res_1 which is the weighted sum of the feature map Coronary_L1 and the feature map RES_Cor by the formula bridge_res_1=α1Coronary_L1+β1RES_Cor, where α1 and β1 are weights, α1=0.3, β1=0.7; h-6) Multiply the feature map Coronary_L2 and the feature map RES_Cor to obtain the feature map bridge_res_2; h-7) The feature map bridge_res_3 which is the weighted sum of the high-dimensional image feature vector Coronary_L3 and the feature map RES_Cor is calculated by the formula bridge_res_3=α2Coronary_L3+β2RES_Cor, where α2 and β2 are both weights, α2=0.7, β2=0.

3.

7. The method for coronary artery CTA image analysis based on multimodal fusion and large language model according to claim 6, characterized in that: Step i) comprises the following steps: i-1) The dilated convolution module of the visual and language feature bridging module consists of a first three-dimensional convolution layer, a first dilated convolution layer, a second three-dimensional convolution layer, a second dilated convolution layer, a Relu activation function, a maximum pooling layer, and a linear transformation layer; i-2) The high-dimensional image feature vector Coronary_L3 is sequentially input into the dilated convolution module, which consists of the first three-dimensional convolution layer, the first dilated convolution layer, the second three-dimensional convolution layer, the second dilated convolution layer, and the Relu activation function, and the feature map Conv_cor is output. The feature map Conv_cor is added to the feature map RES_Cor to obtain the feature map RES_CONV_Cor; i-3) Input the feature map RES_CONV_Cor into the maximum pooling layer of the dilated convolution module to obtain the compressed feature map RES_CONV_Cor_1; i-4) inputting the compressed feature map RES_CONV_Cor_1 into the first branch of the residual module of the visual and language feature bridge module, outputting the feature map RES_f1, inputting the compressed feature map RES_CONV_Cor_1 into the second branch of the residual module of the visual and language feature bridge module, outputting the feature map RES_f2, adding the feature map RES_f1 and the feature map RES_f2 to obtain the feature map RES_CONV_Cor_1.1; i-5) Input the compressed feature map RES_CONV_Cor_1 into the hole convolution module, which consists of the first three-dimensional convolution layer, the first hole convolution layer, the second three-dimensional convolution layer, the second hole convolution layer, and the Relu activation function, and output the feature map Conv_cor′. Add the feature map Conv_cor′ to the feature map RES_Cor to obtain the feature map RES_CONV_Cor_1.2; i-6) Add the feature map RES_CONV_Cor_1.1, the feature map RES_CONV_Cor_1.2, and the compressed feature map RES_CONV_Cor_1 to obtain the feature map R_C_C_P; i-7) Calculate the feature map bridge_dilated_1 which is the weighted sum of the feature map Coronary_L1 and the feature map R_C_C_P by the formula bridge_dilated_1=α3Coronary_L1+β3R_C_C_P, where α3 and β3 are weights, α3=0.3, β3=0.7; i-8) Multiply the feature map R_C_C_P with the feature map Coronary_L2 to obtain the feature map bridge_dilated_2; i-9) Calculate the feature map bridge_dilated_3 which is the weighted sum of the high-dimensional image feature vector Coronary_L3 and the feature map R_C_C_P by the formula bridge_dilated_3=α4Coronary_L3+β4R_C_C_P, where α4 and β4 are weights, α4=0.3, β4=0.7; i-10) Multiply the feature map bridge_res_1 and the feature map bridge_res_2 to obtain the feature map Bridge_Res_1, and add the feature map Bridge_Res_1 and the feature map bridge_res_3 to obtain the feature map Bridge_Res; i-11) by formula Calculate the feature map Bridge_Dilated which is the weighted sum of the feature maps bridge_dilated_1, bridge_dilated_2 and bridge_dilated_3, where α5, β5 and γ1 are weights, α5 = 0.2, β5 = 0.3 and γ1 = 0.5; i-12) Add the feature map Bridge_Res and the feature map Bridge_Dilated to obtain the feature map Bridge_output; i-13) Input the feature map Bridge_output into the linear transformation layer of the hole convolution module, and output the visual feature vector Linear_C.

8. The method for coronary artery CTA image analysis based on multimodal fusion and large language model according to claim 1, characterized in that: In step j), the large language model is the Vicuna1.5-7B model.

9. The method for coronary artery CTA image analysis based on multimodal fusion and large language model according to claim 8, characterized in that: Use the ROCOv2 dataset to train the Vicuna1.5-7B model and obtain the trained large language model.

Citation Information

Patent Citations

  • Nodule calcification medical image processing method

    CN113920082A

  • 3D coronary artery CTA plaque identification method based on deep learning

    CN115713626A

  • Multi-modal dialogue generation method fusing ASPP module and cross-modal interaction

    CN117290461A

  • Multi-modal multi-language medical image text report generation method and device

    CN118098480A