Lung injury intelligent detection system based on lung image and pathological data

Through the medical image segmentation method based on bidirectional gating cross-attention multimodal feature fusion and multi-layer fusion optimization, the problem of accurate lung injury detection of lung images and pathological data in the prior art is solved, and the accuracy of lung injury detection is improved.

CN120260893AActive Publication Date: 2025-07-04QINGDAO UNIV

Patent Information

Application Number
CN202510750578.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Existing artificial intelligence technology is difficult to effectively carry out accurate lung injury detection of lung imaging and pulmonary pathological data.

Method used

The multimodal feature fusion method of cross-attention based on bidirectional gating is adopted, combined with 3D Swin Transformer and BioClinicalBERT, the local and global features of lung CT images are extracted, and modal interaction is achieved through the bidirectional gating attention mechanism. Combined with the medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement, precise segmentation and feature fusion are performed.

Benefits of technology

The accuracy of lung injury detection has been improved, and the accuracy of lung injury detection has been improved through cross-modal feature fusion and precise segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260893A_ABST
    Figure CN120260893A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent medicine, in particular to a lung injury intelligent detection system based on lung images and pathological data, which comprises a multi-modal data acquisition module, an intelligent segmentation module, a cross-modal feature fusion module and a lung injury classification detection module, a trans-attention multi-modal feature fusion method based on bidirectional gating is adopted, a lung CT image and a pathological text are coded, modal interaction is achieved through a bidirectional gating attention mechanism, a lesion area is highlighted, a channel-space attention gate is introduced, feature fusion is optimized, and accurate lung injury detection is achieved. A medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement is adopted, through multi-layer feature optimization based on depth separable convolution and three-layer parallel feature enhancement based on a residual mechanism, key features, namely a focus area, are reserved, local-global features are effectively integrated while nonlinear information is captured, and the image segmentation efficiency is improved. The lung lesion area is accurately segmented, and the accuracy of lung injury detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent medicine, and specifically relates to an intelligent lung injury detection system based on lung images and pathological data. Background Art

[0002] Lung injury detection is an important application in the medical field, mainly analyzing and diagnosing lung image data and clinical pathological data, aiming to detect and evaluate lung injuries, lesions or diseases at an early stage. Existing artificial intelligence technologies are difficult to effectively and accurately detect lung injuries when faced with lung images and clinical pathological data of lung patients. Summary of the Invention

[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides an intelligent lung injury detection system based on lung images and pathological data. Aiming at the problem that existing artificial intelligence technologies are difficult to effectively and accurately detect lung injuries when faced with lung images and clinical pathological data of lung patients, the present invention creatively adopts a cross-attention multimodal feature fusion method based on bidirectional gating, uses 3D Swin Transformer to extract local and global features of lung CT images, retains the spatial hierarchical structure, uses BioClinicalBERT to perform dynamic semantic encoding on pathological texts, captures medical entities and time series features in symptom descriptions, and realizes modal interaction through a bidirectional gating attention mechanism, locates the text semantic focus corresponding to the image features, enhances the visual saliency of the lesion area, introduces a channel-spatial attention gate, dynamically adjusts the fusion weights of multimodal features, and finally realizes optimal feature fusion and activation classification to achieve accurate lung injury detection; the present invention creatively adopts a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement to segment lung CT images. Through multi-layer feature optimization based on depthwise separable convolution, the detailed local features and global context information of lung CT are effectively extracted. Through three-layer parallel feature enhancement based on the residual mechanism, key features, namely the lesion area, are retained, and local-global features are effectively integrated while capturing non-linear information. Finally, based on this, the lung lesion area is accurately segmented, thereby optimizing cross-modal feature fusion and further increasing the accuracy of lung injury detection.

[0004] The intelligent lung injury detection system based on lung images and pathological data provided by the present invention includes a multimodal data acquisition module, an intelligent segmentation module, a cross-modal feature fusion module, and a lung injury classification and detection module;

[0005] The multimodal data acquisition module collects lung CT images from a lung medical image database and collects pathological data of patients visiting the pulmonary department.

[0006] The intelligent segmentation module uses a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement to segment the lung CT image and obtain the lesion image;

[0007] The cross-modal feature fusion module uses a cross-attention multi-modal feature fusion method based on bidirectional gating to fuse and extract features from the pathological data and the lesion image, and obtains the cross-modal fusion feature;

[0008] The lung injury classification and detection module activates and classifies the fusion feature to obtain the classification and detection result of lung injury.

[0009] Further, in the intelligent segmentation module, a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement specifically includes the following steps:

[0010] Step S1: Multi-layer downsampling. Input the lung CT image into a pre-trained three-layer encoder network for downsampling to obtain three-level encoded features , , ;

[0011] Step S2: Optimization of the attention module based on position embedding. Optimize the encoded feature using the attention module based on position embedding to obtain the refined feature, which specifically includes the following steps:

[0012] Step S21: Feature position encoding embedding. Perform position encoding on the encoded feature to obtain the position-encoded feature;

[0013] Step S22: Matrix mapping. Map and embed the position-encoded feature to obtain the query matrix Q, the key matrix K, and the value matrix V;

[0014] Step S23: Self-attention mechanism processing. Perform self-attention processing on the query matrix Q, the key matrix K, and the value matrix V to obtain the self-attention feature;

[0015] Step S24: Feature fusion. Connect the encoded feature with the self-attention feature to obtain the refined feature ;

[0016] Step S3: Multi-layer feature optimization based on depthwise separable convolution. Optimize the refined feature using multi-layer feature optimization based on depthwise separable convolution to obtain the optimized feature : ;

[0017] where represents batch normalization, represents the ReLU activation function, represents the depthwise separable convolution with a convolution kernel size of k, where k = 2i - 1, represents the standard convolution with a convolution kernel size of k×k, represents the optimized feature ;

[0018] Step S4: Three - layer parallel feature enhancement based on the residual mechanism, perform three - layer parallel feature enhancement on the encoded feature to obtain the enhanced feature , which specifically includes the following steps:

[0019] Step S41: Standard convolution and activation layer processing, perform standard convolution, batch normalization, and ReLU activation on the encoded feature to obtain the standard feature : ;

[0020] In the formula, represents the standard feature ;

[0021] Step S42: Local - global feature integration layer processing based on dual - pooling, perform max - pooling and average - pooling on the standard feature respectively, and then connect the obtained results, perform standard convolution, batch normalization, and ReLU activation to obtain the enhanced feature : i 2 =ReLU(BN(Con v 3 [ p 3 avg ( i 1 )⨂ p 3 max ( i 1 )])) ;

[0022] In the formula, represents the average - pooling processing with a convolution kernel size of 3×3, represents the max - pooling processing with a convolution kernel size of 3×3, represents the connection operation, represents the enhanced feature ;

[0023] Step S43: Global pooling layer processing, perform global average - pooling on the encoded feature and then perform standard convolution, batch normalization, and Sigmoid activation to obtain the enhanced feature ;

[0024] Step S44: Residual mechanism processing. For the enhanced feature and the enhanced feature perform dot product multiplication fusion, and add the result of the fusion to the encoded feature to obtain the enhanced feature ;

[0025] Step S5: Multi-layer upsampling. For the enhanced feature , the optimized feature and the encoded feature , perform three-layer upsampling decoding of multi-scale feature fusion based on depthwise separable convolution and three-layer parallel spatial enhancement based on the residual mechanism to obtain the segmentation mask: ; ; ; ;

[0026] wherein, represents the upsampling operation, represents the three-layer parallel spatial enhancement based on the residual mechanism, represents the multi-scale feature fusion based on depthwise separable convolution, represents the standard convolution with a kernel size of 1×1 for dimension alignment before feature connection, , , respectively represent the results of the three-layer upsampling decoding, represents the Sigmoid activation function, represents the segmentation mask;

[0027] Step S6: Generation of the segmented image. Generate the lung segmentation map, i.e., the lesion image, according to the segmentation mask. Specifically, it includes the following steps:

[0028] Step S61: Load the segmentation mask and the corresponding original image;

[0029] Step S62: Convert the segmentation mask into a grayscale image or a binary image to ensure that the value of each pixel represents different segmentation categories or objects;

[0030] Step S63: Generate the segmentation map according to the segmentation mask. Traverse each pixel of the segmentation mask and determine the pixel value in the corresponding position of the segmentation image according to the pixel value;

[0031] Step S64: Visualize the generated segmentation image to display different segmentation regions or objects in the image.

[0032] Furthermore, in the cross-modal feature fusion module, a cross-attention multi-modal feature fusion method based on bidirectional gating specifically includes the following steps:

[0033] Step Q1: Dual-modal feature encoder. Use 3D Swin Transformer to extract local and global features of the lesion image to obtain lesion features, and use BioClinicalBERT to perform dynamic semantic encoding on the pathological data text to obtain text features;

[0034] Step Q2: Cross-modal cross-attention. Implement modal interaction through a bidirectional gating attention mechanism:

[0035] Image → text guidance: Use the lesion features to generate a query vector Q, and use the text features as the key vector K and value vector V. Through the attention mechanism, locate the text semantic focus corresponding to the imaging features to generate interactive text features;

[0036] Text → image guidance: Use the text features to generate a query vector Q, and use the lesion features as the key vector K and value vector V. Through the attention mechanism, enhance the visual saliency of the lesion area to generate interactive visual features;

[0037] Step Q3: Cross-modal feature fusion optimization based on channel-spatial attention gating. Optimize the interactive text features based on channel attention to obtain optimized text features, optimize the interactive visual features based on spatial attention to obtain optimized visual features, and perform weighted fusion on the optimized text features to obtain cross-modal fusion features.

[0038] The beneficial effects achieved by the present invention using the above solution are as follows:

[0039] (1) The present invention creatively adopts a cross-attention multi-modal feature fusion method based on bidirectional gating. It uses 3D Swin Transformer to extract local and global features of the lung CT image, retains the spatial hierarchical structure, uses BioClinicalBERT to perform dynamic semantic encoding on the pathological text, captures medical entities and time series features in the symptom description, and realizes modal interaction through the bidirectional gating attention mechanism, locates the text semantic focus corresponding to the imaging features, enhances the visual saliency of the lesion area, introduces a channel-spatial attention gate, dynamically adjusts the fusion weight of multi-modal features, and finally realizes optimal feature fusion and activation classification to achieve accurate lung injury detection;

[0040] (2) The present invention creatively adopts a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement to segment lung CT images. Through multi-layer feature optimization based on depthwise separable convolution, it effectively extracts the detailed local features and global context information of lung CT. Through three-layer parallel feature enhancement based on the residual mechanism, it retains the key features, namely the lesion areas, and effectively integrates local-global features while capturing non-linear information. Finally, based on this, it accurately segments the lung lesion areas, further optimizes cross-modal feature fusion, and further improves the accuracy of lung injury detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a module diagram of the lung injury intelligent detection system based on lung imaging and pathological data provided by the present invention;

[0042] Figure 2 It is a schematic flowchart of a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement;

[0043] Figure 3 It is a schematic flowchart of a cross-attention multi-modal feature fusion method based on bidirectional gating.

[0044] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0046] Embodiment 1, refer to Figure 1 The lung injury intelligent detection system based on lung imaging and pathological data includes a multi-modal data acquisition module, an intelligent segmentation module, a cross-modal feature fusion module, and a lung injury classification and detection module;

[0047] The multi-modal data acquisition module collects lung CT images from the lung medical image database and collects the pathological data of patients visiting the pulmonary department;

[0048] The intelligent segmentation module adopts a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement to segment the lung CT images and obtain lesion images;

[0049] The cross-modal feature fusion module adopts a cross-attention multi-modal feature fusion method based on bidirectional gating to fuse and extract features from pathological data and lesion images, and obtains cross-modal fusion features;

[0050] The lung injury classification and detection module activates and classifies the fusion features to obtain the classification and detection results of lung injury.

[0051] Example 2, refer to Figure 2 , based on the above example, in the intelligent segmentation module, a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement specifically includes the following steps:

[0052] Step S1: Multi-layer downsampling. Input the lung CT image into a pre-trained three-layer encoder network for downsampling to obtain three-level encoded features , , ;

[0053] Step S2: Optimization of the attention module based on position embedding. Optimize the encoded features using the attention module based on position embedding to obtain refined features;

[0054] Step S3: Multi-layer feature optimization based on depthwise separable convolution. Optimize the refined features using multi-layer feature optimization based on depthwise separable convolution to obtain optimized features : ;

[0055] In the formula, represents batch normalization, represents the ReLU activation function, represents the depthwise separable convolution with a convolution kernel size of k, where k = 2i - 1, represents the standard convolution with a convolution kernel size of k×k, represents the optimized feature ;

[0056] Step S4: Three-layer parallel feature enhancement based on the residual mechanism. Perform three-layer parallel feature enhancement on the encoded features using the residual mechanism to obtain enhanced features ;

[0057] Step S5: Multi-layer upsampling. For the enhanced features , the optimized features and the encoded features , Perform three - layer upsampling decoding with multi - scale feature fusion based on depth - separable convolution and three - layer parallel spatial enhancement based on the residual mechanism to obtain a segmentation mask: ; ; ; ;

[0058] wherein, represents the upsampling operation, represents three - layer parallel spatial enhancement based on the residual mechanism, represents multi - scale feature fusion based on depth - separable convolution, represents a standard convolution with a kernel size of 1×1, used for dimension alignment before feature connection, , , respectively represent the results of the three - layer upsampling decoding, represents the Sigmoid activation function, represents the segmentation mask;

[0059] Step S6: Generate a segmented image. Generate a lung segmentation map, that is, a lesion image, according to the segmentation mask.

[0060] Example 3. This example is based on the above example. Step S2 specifically includes the following steps:

[0061] Step S21: Feature position encoding embedding. Perform position encoding on the encoded feature to obtain a position - encoded feature;

[0062] Step S22: Matrix mapping. Map and embed the position - encoded feature to obtain a query matrix Q, a key matrix K, and a value matrix V;

[0063] Step S23: Self - attention mechanism processing. Perform self - attention processing on the query matrix Q, the key matrix K, and the value matrix V to obtain a self - attention feature;

[0064] Step S24: Feature fusion. Connect the encoded feature with the self - attention feature to obtain a refined feature .

[0065] Example 4. This example is based on the above example. Step S4 specifically includes the following steps:

[0066] Step S41: Standard convolution and activation layer processing. Perform standard convolution, batch normalization, and ReLU activation on the encoded feature to obtain a standard feature : ;

[0067] In the formula, represents a standard feature ;

[0068] Step S42: Based on the processing of the local-global feature integration layer with dual pooling, perform max pooling and average pooling on the standard feature respectively, and connect, standard-convolve, batch-normalize, and ReLU-activate the obtained results to obtain an enhanced feature : i 2 =ReLU(BN(Con v 3 [ p 3 avg ( i 1 )⨂ p 3 max ( i 1 )])) ;

[0069] In the formula, represents average pooling with a convolution kernel size of 3×3, represents max pooling with a convolution kernel size of 3×3, represents a concatenation operation, represents the enhanced feature ;

[0070] Step S43: Global pooling layer processing. After performing global average pooling on the encoded feature , perform standard convolution, batch normalization, and Sigmoid activation to obtain an enhanced feature ;

[0071] Step S44: Residual mechanism processing. Perform dot-product multiplication fusion on the enhanced feature and the enhanced feature , add the fused result to the encoded feature to obtain an enhanced feature .

[0072] Example 5. Refer to Figure 3 . Based on the above example, in the cross-modal feature fusion module, a cross-attention multi-modal feature fusion method based on bidirectional gating specifically includes the following steps:

[0073] Step Q1: The dual-modal feature encoder uses 3D Swin Transformer to extract local and global features of the lesion images, obtaining lesion features, and uses BioClinicalBERT to perform dynamic semantic encoding on the pathological data text, obtaining text features;

[0074] Step Q2: Cross-modal cross-attention realizes modal interaction through a bidirectional gated attention mechanism:

[0075] Image → Text guidance: Use the lesion features to generate a query vector Q, and the text features as the key vector K and value vector V. Through the attention mechanism, locate the text semantic focus corresponding to the imaging features and generate interactive text features;

[0076] Text → Image guidance: Generate a query vector Q with the text features, and the lesion features as the key vector K and value vector V. Through the attention mechanism, enhance the visual saliency of the lesion area and generate interactive visual features;

[0077] Step Q3: Cross-modal feature fusion optimization based on channel-spatial attention gating. Optimize the interactive text features based on channel attention to obtain optimized text features, optimize the interactive visual features based on spatial attention to obtain optimized visual features, and perform weighted fusion on the optimized text features to obtain cross-modal fusion features.

[0078] Example Six, based on the above example, the training design of the system:

[0079] Loss function

[0080] Classification loss: Cross-entropy loss is used for disease classification;

[0081] Optimizer and learning rate

[0082] Use the AdamW optimizer, and the learning rate strategy adopts learning rate decay;

[0083] Training and inference

[0084] Data augmentation:

[0085] Image: Use flipping and cropping for enhancement;

[0086] Text: Use synonym replacement and masking;

[0087] Inference:

[0088] Input the patient's CT image and pathological text, and output the lung disease classification result.

[0089] Example Seven, based on the above example, the classification and detection rules for lung injury are as follows:

[0090] Pneumonia:

[0091] including different types such as bacterial pneumonia and viral pneumonia;

[0092] Detection manifestation: showing features such as infiltrative shadows in CT scans;

[0093] Tuberculosis:

[0094] Pulmonary tuberculosis lesions;

[0095] Detection manifestation: having features such as nodules and cavities;

[0096] Lung cancer:

[0097] including malignant tumors and benign tumors;

[0098] Detection manifestation: having abnormal structures such as lung masses and tumors;

[0099] Emphysema:

[0100] Detection manifestation: destruction and overinflation of lung tissue;

[0101] Pulmonary edema:

[0102] Detection manifestation: imaging changes caused by the accumulation of fluid in lung tissue;

[0103] Pulmonary embolism:

[0104] Detection manifestation: pulmonary embolism caused by thrombosis in the pulmonary artery or its branches;

[0105] Lung ischemia-reperfusion injury:

[0106] Detection manifestation: having uneven infiltrative shadows, thickening of the alveolar wall, blurred lung markings, blurred lung contour, and blurred pleural contour.

[0107] Example 8. Based on the above example, an optimized implementation of a cross-attention multimodal feature fusion method based on bidirectional gating:

[0108] Step Q1: Bimodal feature encoding

[0109] Image encoding: Use 3D Swin Transformer to extract local and global features of lung CT images, retaining the spatial hierarchical structure;

[0110] Text encoding: Use BioClinicalBERT to perform dynamic semantic encoding on pathological texts, capturing medical entities and time series features in symptom descriptions;

[0111] Step Q2: Optimization of cross-modal cross-attention mechanism

[0112] Achieve modal interaction through a bidirectional gating attention mechanism:

[0113] Image→Text Guidance: Use CT features to generate a query vector (Q), with text features as keys / values (K / V) to locate the text semantic focus corresponding to imaging features;

[0114] Text→Image Guidance: Generate Q with text features and use CT features as K / V to enhance the visual saliency of the lesion area;

[0115] Step Q3: Multi-scale Feature Fusion

[0116] Introduce a channel-spatial attention gate to dynamically adjust the fusion weights of multi-modal features, design a pyramid fusion structure, and splice and convolve cross-modal features at 4 scales (original resolution → 1 / 8 downsampling) to capture fine-grained lesion features;

[0117] Dynamic Sparse Attention Mechanism

[0118] Introduce Top-K sparsification during cross-modal interaction, only retaining the top 20% of attention connections;

[0119] Step Q4: Lesion Localization-Classification Joint Network

[0120] Adopt a dual-branch structure:

[0121] Localization Branch: Regress the coordinates of the lesion area through a heatmap (such as the bounding box of ground-glass opacity)

[0122] Classification Branch: Combine multi-modal features to output the probability of disease types (pneumonia / lung cancer / fibrosis, etc.).

[0123] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0124] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

[0125] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. An intelligent lung injury detection system based on lung images and pathological data, characterized in that: It includes a multi-modal data acquisition module, an intelligent segmentation module, a cross-modal feature fusion module, and a lung injury classification and detection module; The multi-modal data acquisition module collects lung CT images from a lung medical image database and collects pathological data of patients visiting the pulmonary department; The intelligent segmentation module uses a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement to segment the lung CT images and obtain lesion images; The cross-modal feature fusion module uses a cross-attention multi-modal feature fusion method based on bidirectional gating to fuse and extract features from the pathological data and the lesion images to obtain cross-modal fusion features; The lung injury classification and detection module activates and classifies the fusion features to obtain the classification and detection results of lung injury.

2. The intelligent lung injury detection system based on lung images and pathological data according to claim 1, characterized in that: In the intelligent segmentation module, a medical image segmentation method based on multi-layer fusion optimization and multi-layer parallel enhancement specifically includes the following steps: Step S1: Perform three-layer downsampling on the lung CT images to obtain three-level encoded features , , ; Step S2: Optimize the encoded features using the attention module based on positional embedding to obtain refined features; Step S3: Perform multi-layer feature optimization on the refined features using depthwise separable convolution to obtain optimized features : ; Wherein, represents batch normalization, represents the ReLU activation function, represents depthwise separable convolution with a convolution kernel size of k, where k = 2i - 1, represents standard convolution with a convolution kernel size of k×k, represents optimized features ; Step S4: Perform three-layer parallel feature enhancement on the encoded features to obtain enhanced features ; Step S5: For the enhanced features , the optimized features , and the encoded features , perform three-layer upsampling decoding of multi-scale feature fusion based on depthwise separable convolution and three-layer parallel spatial enhancement based on the residual mechanism to obtain the segmentation mask; Step S6: Generate a lung segmentation map, that is, a lesion image, according to the segmentation mask.

3. The intelligent lung injury detection system based on lung images and pathological data according to claim 2, characterized in that: Step S4 specifically includes the following steps: Step S41: Perform standard convolution, batch normalization, and ReLU activation on the encoded feature to obtain a standard feature ; Step S42: For the standard features perform max pooling and average pooling respectively, and concatenate, standard convolution, batch normalization, and ReLU activation on the obtained results to obtain enhanced features ; Step S43: Perform global average pooling, standard convolution, batch normalization, and Sigmoid activation on the encoded feature to obtain an enhanced feature ; Step S44: After fusing the enhanced feature with the enhanced feature , add it to the encoded feature to obtain the enhanced feature .

4. The intelligent lung injury detection system based on lung images and pathological data according to claim 1, characterized in that: In the cross-modal feature fusion module, a cross-attention multi-modal feature fusion method based on bidirectional gating specifically includes the following steps: Step Q1: Use a visual feature encoder to extract local and global features of the lesion image to obtain lesion features, and use a text feature encoder to perform dynamic semantic encoding on the pathological data text to obtain text features; Step Q2: Implement modal interaction through a bidirectional gating attention mechanism to obtain interactive text features and interactive visual features; Step Q3: Optimize the interactive text features based on channel attention to obtain optimized text features, optimize the interactive visual features based on spatial attention to obtain optimized visual features, and perform weighted fusion on the optimized text features to obtain cross-modal fusion features.

Citation Information

Patent Citations

  • Ultrasonic breast tumor automatic segmentation method based on attention enhancement improved U-shaped network

    CN112785598A

  • CT image lung lobe image segmentation system based on attention mechanism

    CN113936011A

  • Pulmonary nodule detection and identification system based on YOLOv5

    CN115661029A

  • Microblog user depression tendency detection method and device based on multi-feature fusion

    CN117094319A

  • Multi-modal fusion recognition system and device for mycobacterium pulmonary granuloma infection

    CN119360167A

Cited By

  • Intelligent lung cancer detection system based on PET / MR (positron emission tomography / magnetic resonance) multi-mode image

    CN120635098A

  • A smart lung cancer detection system based on PET / MR multimodal imaging

    CN120635098B

  • Lung ventilation-perfusion development area image segmentation and quantitative analysis system

    CN120953301A

  • Video saliency prediction method based on multi-scale bidirectional gating and perception guidance fusion

    CN121999414A