Pneumonia image segmentation and prediction recognition system based on sam-clip multimodal multitask

CN122821113APending Publication Date: 2026-09-25GUILIN MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610793810.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0008]本发明的目的在于提供基于SAM-CLIP 多模态多任务的肺炎影像分割与预测识别系统,解决现有肺炎影像分析方法中存在单模态特征局限性、任务协同缺失和大模型适配不足的技术问题

Benefits of technology

[0019](1)本发明通过 CLIP 模型对医学文本进行编码,将临床语义信息转化为分割任务的引导信号,解决了传统模型对病灶边界模糊、对比度低的肺炎影像分割易出现过分割的问题。采用图文跨模态注意力机制,动态加权融合影像视觉特征与文本语义特征,生成更具临床针对性的粗粒度掩模,为 SAM 模型提供高质量的初始提示,引导其实现病灶边界的精细化分割,显著提升分割精度与鲁棒性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821113A_ABST
    Figure CN122821113A_ABST
Patent Text Reader

Abstract

The application provides a pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitask, belongs to the technical field of pneumonia image processing, and comprises a medical text information acquisition module, a pneumonia image acquisition module, a SAM-CLIP cross-modal interaction module, a task feature fusion module and a model result output module. The output ends of the medical text information acquisition module and the pneumonia image acquisition module are connected with the SAM-CLIP cross-modal interaction module, the acquisition end of the medical text information acquisition module is connected with an external hospital information system, the acquisition end of the pneumonia image acquisition module is connected with an external hospital image system, the SAM-CLIP cross-modal interaction module is connected with the task feature fusion module, and the task feature fusion module is connected with the model result output module. The medical text is encoded through the CLIP model, the clinical semantic information is converted into a guide signal of the segmentation task, and the problem that the traditional model is prone to over-segmentation when segmenting the pneumonia image with a fuzzy lesion boundary and low contrast is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pneumonia image processing technology, and in particular to a pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking. Background Technology

[0002] Pneumonia, as a highly prevalent clinical disease, relies on precise lesion localization, quantitative assessment, and prognostic risk stratification for accurate diagnosis and treatment. Currently, pneumonia image analysis in clinical workflows generally faces the following technical bottlenecks:

[0003] (1) Insufficient image segmentation accuracy and poor robustness of lesion identification Traditional pneumonia segmentation methods rely heavily on single-modal image features, which are poorly adapted to problems such as blurred lesion boundaries, uneven textures, and low contrast with surrounding tissues in CT images, and are prone to over-segmentation or under-segmentation. At the same time, single-modal models lack semantic guidance from clinical text information, making it difficult to combine clinical descriptions to enhance lesion features in a targeted manner, resulting in a deviation between the segmentation results and clinical needs, and failing to provide a reliable basis for quantitative assessment.

[0004] (2) The segmentation and prognostic tasks are disconnected and lack a multi-task collaborative mechanism. In the existing technology, the segmentation of pneumonia lesions and the prognostic risk assessment of patients are usually carried out as two independent tasks: the segmentation task only outputs the lesion area mask and cannot be directly related to the clinical outcome; the prognostic classification task relies on global image features and lacks accurate support from local lesion features. Moreover, the two types of tasks have not established a two-way feedback mechanism, which makes it impossible for the segmentation results to provide effective guidance for prognostic assessment. The prognostic labels are also difficult to optimize the segmentation model in reverse. The performance of both is limited by the information island effect.

[0005] (3) The large model is not adaptable enough and it is difficult to implement in medical scenarios. General segmentation large models perform well in natural images, but when directly applied to medical images, they have problems such as insufficient understanding of medical lesion features, lack of clinical semantic guidance, and weak generalization ability in small sample scenarios. At the same time, although multimodal models such as CLIP have the ability to understand text and image semantics, they have not designed cross-modal interaction mechanisms for medical image segmentation scenarios, and cannot achieve accurate mapping from text semantics to lesion segmentation. The advantages of the two have not been effectively integrated.

[0006] (4) Insufficient utilization of multi-scale features and serious loss of hierarchical information Traditional multi-task models are insufficient in mining multi-scale features of images. Low-level detail features and high-level semantic features are not effectively integrated, and there is a lack of adaptive aggregation mechanism for multi-scale features. As a result, the model has poor adaptability to pneumonia lesions of different sizes and shapes, and it is difficult to simultaneously achieve both segmentation accuracy and classification stability.

[0007] In summary, current pneumonia image analysis technology suffers from three major shortcomings: First, the limitations of single-modal features: There is a lack of deep integration between medical text semantics and image features, resulting in segmentation results lacking clinical semantic guidance. Second, the lack of task collaboration: Segmentation and prognostic tasks are independent, and a bidirectional optimization mechanism has not been established, making it impossible to achieve synergistic improvement in lesion localization and risk assessment. Third, insufficient adaptation to large models: General-purpose models such as SAM and CLIP have not been customized for medical imaging scenarios, making it difficult to directly meet the clinical needs of pneumonia diagnosis and treatment. Therefore, it is necessary to design a pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal and multi-task approaches. Summary of the Invention

[0008] The purpose of this invention is to provide a pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking, which solves the technical problems of single-modal feature limitations, lack of task collaboration and insufficient large model adaptation in existing pneumonia image analysis methods.

[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0010] The SAM-CLIP multimodal and multi-task-based pneumonia image segmentation and prediction recognition system includes a medical text information acquisition module, a pneumonia image acquisition module, a SAM-CLIP cross-modal interaction module, a task feature fusion module, and a model result output module. The output terminals of the medical text information acquisition module and the pneumonia image acquisition module are connected to the SAM-CLIP cross-modal interaction module. The acquisition terminal of the medical text information acquisition module is connected to an external hospital information system, the acquisition terminal of the pneumonia image acquisition module is connected to an external hospital imaging system, the SAM-CLIP cross-modal interaction module is connected to the task feature fusion module, and the task feature fusion module is connected to the model result output module.

[0011] The medical text information acquisition module is used to collect pneumonia-related text information data from the hospital information system. The pneumonia image acquisition module is used to collect pneumonia image data from the hospital imaging system. The SAM-CLIP cross-modal interaction module is used to generate coarse-grained masks by utilizing CLIP's semantic understanding capabilities, fuse image and text features through an adaptive attention mechanism, and guide the SAM model to finely segment pneumonia lesions. The task feature fusion module uses a hierarchical group aggregation bridge structure to fuse multi-scale image features and mask information. By jointly optimizing segmentation loss, stroke classification / prognostic classification loss, and several task perception losses, it achieves collaborative analysis of lesion localization and clinical risk assessment. The model result output module is used to output the segmentation mask of the lesion region and the prognostic risk classification of the patient.

[0012] Furthermore, the SAM-CLIP cross-modal interaction module includes a CLIP image encoder, a CLIP text encoder, a cross-modal interaction submodule, a segmentation model submodule, and a mask image input submodule. The input end of the cross-modal interaction submodule is connected to the CLIP image encoder and the CLIP text encoder, respectively. The output end of the cross-modal interaction submodule is connected to the segmentation model submodule. The mask image input submodule is connected to the segmentation model submodule. The segmentation model submodule outputs the segmentation result after fusing image and text information.

[0013] 3. The pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking as described in claim 2, characterized in that: the mask image input submodule is used to input a mask image and generate prompt points; the mask image input submodule is equipped with randomly selected points, which are sampled from the mask image and used as point prompts for the segmentation model submodule.

[0014] Furthermore, the cross-modal interaction submodule includes max pooling, three convolutional layers, an attention mechanism, a normalization layer, and a matrix adder. The input of max pooling is connected to the output of the CLIP image encoder for compressing image features. The output of max pooling is connected to the first convolutional layer. The query branch of the attention mechanism at the output of the first convolutional layer is connected to the attention mechanism. The second and third convolutional layers are both connected to the CLIP text encoder. The key branch of the attention mechanism in the second convolutional layer is connected to the attention mechanism. The value branch of the attention mechanism in the third convolutional layer is connected to the attention mechanism. The attention mechanism is connected to the matrix adder via the normalization layer.

[0015] Furthermore, the segmentation model submodule includes a segmentation convolutional layer, a segmentation matrix adder, a cue encoder, and a mask decoder. The segmentation convolutional layer is connected to the CLIP mask branch of the matrix adder, the output of the segmentation convolutional layer is connected to the segmentation matrix adder, the input of the cue encoder is connected to the randomly selected points in the code image input submodule, and the cue encoder is connected to the mask decoder.

[0016] Furthermore, the task feature fusion module includes four feature fusion cross-modal interaction sub-modules, four global attention modules, and a feature transformation convolutional layer. The four feature fusion cross-modal interaction sub-modules are connected to four feature map outputs of different scales obtained by downsampling the original image four times. The four feature fusion cross-modal interaction sub-modules are also connected to a mask image and medical text, respectively. The mask image is used to provide prompts and supervision labels for the segmentation task, and the medical text is text data of the patient's clinical information and diagnosis description, used for cross-modal guidance. The four feature fusion cross-modal interaction sub-modules are connected to four global attention modules, and each global attention module is connected to a feature transformation convolutional layer.

[0017] Furthermore, the model output module includes a medical image analysis output submodule and a clinical decision support output submodule. Both the medical image analysis output submodule and the clinical decision support output submodule are connected to the feature transformation convolutional layer. The medical image analysis output submodule is used for segmentation masking of lesion regions to realize medical image analysis, and the clinical decision support output submodule is used for patient prognostic risk classification to realize clinical decision support.

[0018] The present invention, by adopting the above-described technical solution, has the following beneficial effects:

[0019] (1) This invention encodes medical text using the CLIP model, transforming clinical semantic information into guiding signals for segmentation tasks. This solves the problem of oversegmentation that traditional models often encounter when segmenting pneumonia images with blurred lesion boundaries and low contrast. By employing a cross-modal attention mechanism, the image visual features and text semantic features are dynamically and weightedly fused to generate more clinically targeted coarse-grained masks. This provides high-quality initial prompts for the SAM model, guiding it to achieve refined segmentation of lesion boundaries and significantly improving segmentation accuracy and robustness.

[0020] (2) A collaborative learning framework for lesion segmentation and prognostic classification was constructed through a multi-task feature fusion module. A hierarchical group aggregation bridge structure was adopted to effectively fuse image features and mask information at different scales, while capturing the local detail texture and global distribution features of lesions, providing more comprehensive feature support for segmentation and classification tasks. By jointly optimizing the segmentation loss, prognostic classification loss, and multi-task perception loss, bidirectional supervision and mutual promotion between the two tasks were achieved. The lesion mask output by the segmentation task provides accurate local lesion features for prognostic classification, improving the accuracy of risk assessment. The global semantic information of the prognostic classification task guides the segmentation model in reverse, making it pay more attention to key lesion areas related to prognosis, further optimizing the segmentation effect. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the system structure of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and preferred embodiments. However, it should be noted that many details listed in the specification are merely to provide the reader with a thorough understanding of one or more aspects of the present invention, and these aspects of the invention can be implemented even without these specific details.

[0023] like Figure 1As shown, the pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking includes a medical text information acquisition module, a pneumonia image acquisition module, a SAM-CLIP cross-modal interaction module, a task feature fusion module, and a model result output module. The output ends of the medical text information acquisition module and the pneumonia image acquisition module are connected to the SAM-CLIP cross-modal interaction module. The acquisition end of the medical text information acquisition module is connected to an external hospital information system, the acquisition end of the pneumonia image acquisition module is connected to an external hospital imaging system, the SAM-CLIP cross-modal interaction module is connected to the task feature fusion module, and the task feature fusion module is connected to the model result output module.

[0024] The medical text information acquisition module is used to collect pneumonia-related text information data from the hospital information system. The pneumonia image acquisition module is used to collect pneumonia image data from the hospital imaging system. The SAM-CLIP cross-modal interaction module is used to generate coarse-grained masks by utilizing CLIP's semantic understanding capabilities, fuse image and text features through an adaptive attention mechanism, and guide the SAM model to finely segment pneumonia lesions. The task feature fusion module uses a hierarchical group aggregation bridge structure to fuse multi-scale image features and mask information. By jointly optimizing segmentation loss, stroke classification / prognostic classification loss, and several task perception losses, it achieves collaborative analysis of lesion localization and clinical risk assessment. The model result output module is used to output the segmentation mask of the lesion region and the prognostic risk classification of the patient.

[0025] In this embodiment of the invention, the SAM-CLIP cross-modal interaction module includes a CLIP image encoder, a CLIP text encoder, a cross-modal interaction submodule, a segmentation model submodule, and a mask image input submodule. The input end of the cross-modal interaction submodule is connected to the CLIP image encoder and the CLIP text encoder, respectively. The output end of the cross-modal interaction submodule is connected to the segmentation model submodule. The mask image input submodule is connected to the segmentation model submodule. The segmentation model submodule outputs the segmentation result after fusing image and text information.

[0026] The mask image input submodule is used to input the mask image and generate cue points. The mask image input submodule is set to randomly select points, which are sampled from the mask image and used as cue points for the segmentation model submodule.

[0027] The cross-modal interaction submodule includes max pooling, three convolutional layers, an attention mechanism, a normalization layer, and a matrix adder. The input of max pooling is connected to the output of the CLIP image encoder for compressing image features. The output of max pooling is connected to the first convolutional layer. The query branch of the attention mechanism at the output of the first convolutional layer is connected to the attention mechanism. The second and third convolutional layers are both connected to the CLIP text encoder. The key branch of the attention mechanism in the second convolutional layer is connected to the attention mechanism. The value branch of the attention mechanism in the third convolutional layer is connected to the attention mechanism. The attention mechanism is connected to the matrix adder via the normalization layer.

[0028] The segmentation model submodule includes a segmentation convolutional layer, a segmentation matrix adder, a cue encoder, and a mask decoder. The segmentation convolutional layer is connected to the CLIP mask branch of the matrix adder. The output of the segmentation convolutional layer is connected to the segmentation matrix adder. The input of the cue encoder is connected to the randomly selected points in the code image input submodule. The cue encoder is connected to the mask decoder.

[0029] The advanced semantic understanding capabilities of the CLIP model are used to generate an initial coarse mask, which is then used as a guiding constraint for the SAM model to improve segmentation performance.

[0030] Image and text inputs are processed to obtain an initial CLIP mask. During this process, an attention mechanism for adaptive weighting is introduced to facilitate interaction between image and text features. A query (Q) is generated by performing max pooling and convolution (Conv) operations on the image features, while two different sets of convolutional layers process the text features separately, producing keys (K) and values ​​(V). Subsequently, an attention (Att) mechanism is used to combine Q, K, and V, enabling cross-modal feature interaction. This process is further enhanced by a residual connection containing the query output and a normalization layer.

[0031] The formula for the CLIP mask is:

[0032] .

[0033] Using the coarse CLIP mask from the previous step, this study refines it using SAM, guided by the provided bounding box and mask cues, to generate an effective mask. The specific steps involve first integrating the bounding box cues and the coarse mask cues to create a bounding box (p) and a set of randomly selected points (b), respectively, and then feeding them into SAM's cue encoder. Processing is then performed. Finally, the SAM mask encoder... This is responsible for generating the effective mask for SAM-CLIP, which fuses the coarse segmentation results with the original image, while also being influenced by cue constraints. Therefore,

[0034] The effective mask for SAM-CLIP is represented as follows:

[0035] .

[0036] By comprehensively comparing different configurations of this module, this study intends to use pre-trained... , And it was decided to keep the attention component unchanged, unmodified, and frozen. To achieve the best...

[0037] Excellent performance; this study uses a multimodal lung adenocarcinoma dataset to analyze... Make minor adjustments.

[0038] In this embodiment of the invention, the task feature fusion module includes four feature fusion cross-modal interaction sub-modules, four global attention modules, and a feature transformation convolutional layer. The four feature fusion cross-modal interaction sub-modules are respectively connected to four feature map outputs of different scales obtained by downsampling the original image four times. The four feature fusion cross-modal interaction sub-modules are also respectively connected to a mask image and medical text. The mask image is used to provide prompts and supervision labels for segmentation tasks, and the medical text is text data of the patient's clinical information and diagnosis description, used for cross-modal guidance. The four feature fusion cross-modal interaction sub-modules are respectively connected to four global attention modules, and each global attention module is connected to a feature transformation convolutional layer.

[0039] Medical lesions in images can manifest at different levels and scales, posing challenges to algorithm design. This requires...

[0040] The goal is to find an algorithm capable of handling the diversity of images while considering the limitations of computational resources in clinical practice. This study proposes a multi-task feature fusion module, which is capable of handling multiple tasks. The multi-task feature fusion (MTFF) module consists of four consecutive scales or stages, arranged hierarchically from bottom to top. Inspired by the group bridging structure in EGE-UNet, this study uses a group aggregation bridge (GAB) to fuse image features from each stage. The study employs a segmentation prediction method, which takes three inputs: the corresponding valid mask (VM) and features from the previous stage. The fused features from each stage serve as input for the next stage. These fused features are processed element-wise by adding them together and applying convolutional layers, generating a prediction mask at each stage. These predictions are then evaluated against their respective segmentation labels to compute a loss function. In the final stage, the study obtains segmentation predictions and multi-scale fused features, denoted as segmentation output (S) and feature output (F), respectively. The GAB design contains three inputs: image features, the valid mask, and lower-level features. Initially, these inputs are sequentially concatenated within a group, forming four such groups. Each group is then processed using dilated convolutional layers with a kernel size of 3 and dilation rates of {1, 2, 5, 7} corresponding to the group. Finally, the outputs of the four groups are concatenated and fused through an additional convolutional layer.

[0041] Loss Function: The loss function is carefully designed to facilitate the optimization process. It includes separate losses for segmentation and classification tasks, as well as a multi-task-aware loss for establishing correlations between tasks.

[0042] Segmentation Loss: For each scale level in the segmentation task, this study computes DSC and Jaccard to evaluate the consistency between the prediction mask and the ground truth. Overall Segmentation Loss , where is the weighted sum of these indicators across four scales, with scale weights of 1, 0.75, 0.5, and 0.25 for scales ranging from 0 to 3. Classification Loss: For the risk-return assessment classification task, this study uses weighted cross-entropy loss to quantify the deviation between the predicted results and the true labels. Its mathematical formula is as follows:

[0043]

[0044] Class-specific weights are derived from the inverse frequencies of the two classes across the entire dataset to mitigate the effects of potential sample imbalance. Indicates an indicator function, Indicates the predicted probability. It is the predicted true label.

[0045] Multi-task perceptual loss: Segmentation and classification are two closely related tasks, especially in the context of clinical applications, as consistency between them is crucial. This study obtains pixel-level probability values ​​P by passing the feature output F through a softmax layer, and utilizes Jensen-Shannon Divergence to quantify the difference between P and the segmentation output S, expressed as: This serves as an indicator to ensure the required consistency between segmentation and classification tasks.

[0046] The total losses are as follows:

[0047] Class-specific weights derived from the inverse frequencies of the two classes across the entire dataset are used to mitigate the effects of potential sample imbalance.

[0048] Multi-task perceptual loss: Segmentation and classification are two closely related tasks, especially in the context of clinical applications, as consistency between them is crucial. This study obtains pixel-level probability values ​​P by passing the feature output F through a softmax layer, and utilizes Jensen-Shannon Divergence to quantify the difference between P and the segmentation output S, expressed as: This serves as an indicator to ensure the required consistency between segmentation and classification tasks.

[0049] The total losses are as follows:

[0050]

[0051] Where and are weighted hyperparameters, respectively coordinating... and The contribution of these hyperparameters was determined experimentally, with values ​​set to 0.2 and 0.8 to ensure a balance between segmentation and classification objectives.

[0052] In this embodiment of the invention, the model result output module includes a medical image analysis output submodule and a clinical auxiliary decision output submodule. Both the medical image analysis output submodule and the clinical auxiliary decision output submodule are connected to the feature transformation convolutional layer. The medical image analysis output submodule is used for segmentation masking of lesion areas to realize medical image analysis, and the clinical auxiliary decision output submodule is used for patient prognostic risk classification to realize clinical auxiliary decision-making.

[0053] Matters not covered in this invention are common knowledge.

[0054] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking, characterized in that: It includes a medical text information acquisition module, a pneumonia image acquisition module, a SAM-CLIP cross-modal interaction module, a task feature fusion module, and a model result output module. The output ends of the medical text information acquisition module and the pneumonia image acquisition module are connected to the SAM-CLIP cross-modal interaction module. The acquisition end of the medical text information acquisition module is connected to an external hospital information system. The acquisition end of the pneumonia image acquisition module is connected to an external hospital imaging system. The SAM-CLIP cross-modal interaction module is connected to the task feature fusion module. The task feature fusion module is connected to the model result output module. The medical text information acquisition module is used to collect pneumonia-related text information data from the hospital information system. The pneumonia image acquisition module is used to collect pneumonia image data from the hospital imaging system. The SAM-CLIP cross-modal interaction module is used to generate coarse-grained masks using CLIP's semantic understanding capabilities, fuse image and text features through an adaptive attention mechanism, and guide the SAM model to finely segment pneumonia lesions. The task feature fusion module uses a hierarchical group aggregation bridge structure to fuse multi-scale image features and mask information. By jointly optimizing segmentation loss, stroke classification / prognostic classification loss, and several task perception losses, it achieves collaborative analysis of lesion localization and clinical risk assessment. The model result output module is used to output the segmentation mask of the lesion region and the patient's prognostic risk classification.

2. The pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking as described in claim 1, characterized in that: The SAM-CLIP cross-modal interaction module includes a CLIP image encoder, a CLIP text encoder, a cross-modal interaction submodule, a segmentation model submodule, and a mask image input submodule. The input of the cross-modal interaction submodule is connected to the CLIP image encoder and the CLIP text encoder, respectively. The output of the cross-modal interaction submodule is connected to the segmentation model submodule. The mask image input submodule is connected to the segmentation model submodule. The segmentation model submodule outputs the segmentation result after fusing image and text information.

3. The pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking as described in claim 2, characterized in that: The mask image input submodule is used to input the mask image and generate cue points. The mask image input submodule is set to randomly select points, which are sampled from the mask image and used as cue points for the segmentation model submodule.

4. The pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking as described in claim 2, characterized in that: The cross-modal interaction submodule includes max pooling, three convolutional layers, an attention mechanism, a normalization layer, and a matrix adder. The input of max pooling is connected to the output of the CLIP image encoder for compressing image features. The output of max pooling is connected to the first convolutional layer. The query branch of the attention mechanism at the output of the first convolutional layer is connected to the attention mechanism. The second and third convolutional layers are both connected to the CLIP text encoder. The key branch of the attention mechanism in the second convolutional layer is connected to the attention mechanism. The value branch of the attention mechanism in the third convolutional layer is connected to the attention mechanism. The attention mechanism is connected to the matrix adder via the normalization layer.

5. The pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking as described in claim 4, characterized in that: The segmentation model submodule includes a segmentation convolutional layer, a segmentation matrix adder, a cue encoder, and a mask decoder. The segmentation convolutional layer is connected to the CLIP mask branch of the matrix adder. The output of the segmentation convolutional layer is connected to the segmentation matrix adder. The input of the cue encoder is connected to the randomly selected points in the code image input submodule. The cue encoder is connected to the mask decoder.

6. The pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking as described in claim 1, characterized in that: The task feature fusion module includes four feature fusion cross-modal interaction sub-modules, four global attention modules, and a feature transformation convolutional layer. The four feature fusion cross-modal interaction sub-modules are connected to four feature map outputs of different scales obtained by downsampling the original image four times. The four feature fusion cross-modal interaction sub-modules are also connected to a mask image and medical text, respectively. The mask image is used to provide prompts and supervision labels for the segmentation task, and the medical text is text data containing the patient's clinical information and diagnostic descriptions for cross-modal guidance. The four feature fusion cross-modal interaction sub-modules are connected to four global attention modules, and each global attention module is connected to a feature transformation convolutional layer.

7. The pneumonia image segmentation and prediction recognition system based on SAM-CLIP multimodal multitasking as described in claim 6, characterized in that: The model output module includes a medical image analysis output submodule and a clinical decision support output submodule. Both the medical image analysis output submodule and the clinical decision support output submodule are connected to the feature transformation convolutional layer. The medical image analysis output submodule is used for lesion region segmentation masking to realize medical image analysis, and the clinical decision support output submodule is used for patient prognostic risk classification to realize clinical decision support.