Disease segmentation method and system based on large language model

By combining the image encoder that integrates the Transformer module and the Adapter module and the semantic prompt encoder of the large language model, the problems of low accuracy and difficulty in automated segmentation in engineering disease detection are solved, efficient semantic segmentation of disease images is achieved, and the accuracy and automation of disease recognition are improved.

CN120279035APending Publication Date: 2025-07-08GUANGDONG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510297101.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing deep learning-based disease detection methods have problems such as low accuracy, target domain offset and automated segmentation in the engineering field.

Method used

The image encoder based on the fusion Transformer module and the Adapter module is used to encode the diseased image, and the semantic prompt encoder is combined with the large language model semantic prompt encoder to generate target segmentation content prompts, and the mask decoder is used for mask decoding to realize the semantic segmentation of the diseased image.

Benefits of technology

The comprehensive, complete and adaptive segmentation capabilities of disease categories are improved, the ability to automatically identify new disease types under fewer samples training is enhanced, and the success rate of target disease extraction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279035A_ABST
    Figure CN120279035A_ABST
Patent Text Reader

Abstract

The invention discloses a disease segmentation method and system based on a language large model. The method comprises the steps of obtaining a to-be-segmented disease image; an image encoder based on a fusion Transform module and an Adapter module encodes the disease image, and the encoded image is output and embedded; generating a target segmentation content prompt based on a large language model semantic prompt encoder; and performing mask decoding by using a mask decoder according to the encoded image block embedding and the target segmentation content prompt to obtain a disease image after semantic segmentation. According to the invention, comprehensive, complete and adaptive segmentation is carried out on disease types by using a powerful visual large model migration network, and the accuracy of automatically identifying new disease types under the condition of less sample training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a disease segmentation method and system based on a large language model. Background Art

[0002] During the long-term use of construction projects, they may be affected by various factor changes, such as geological changes, natural disasters, traffic loads, construction errors, etc., resulting in surface diseases. Surface diseases may lead to the rapid deterioration of the main structure performance, not only may endanger the stability of the engineering structure, but also may pose a potential threat to the use safety of the project. Therefore, timely and accurate detection and analysis of diseases have become an important link to ensure the safe operation of the project.

[0003] Based on deep learning technology, its powerful generalization ability has been applied in the engineering field, which can solve two major challenges in the detection of engineering surface diseases. The first challenge stems from the complex acquisition environment, such as insufficient lighting or shadow occlusion angles, as well as the interference of complex background textures other than the diseases themselves. This may lead to the lack of local features and incorrect identification of disease positions. The second challenge stems from the variability of diseases. For example, the same type of disease may present different orientations, textures, and colors, and most diseases are linearly distributed. The collected samples cannot fully cover all disease categories, resulting in incorrect identification of disease types or inability to segment potential diseases.

[0004] Based on the DETR (Detection Transformer) end-to-end vision large model, a research has proposed the Segment Anything Model (SAM). By given an image and prompts such as points, boxes, or rough masks, the content of interest in the image can be segmented. The release of the model once again marks an important breakthrough of large models in the field of computer vision. The SAM large model consists of three parts: an image encoder, a prompt encoder, and a mask decoder. Among them, the image encoder is based on ViT and has a large number of parameters, reflecting powerful information abstraction ability. The prompt encoder encodes the input prompts (such as points, boxes, or masks) into vector form and combines them with image features as embedding vectors. The mask decoder synthesizes image features and prompt vectors and generates pixel-level segmentation masks through cross-attention. Using the powerful network of SAM to achieve precise extraction of engineering diseases, applying the rich pre-trained background knowledge and powerful network generalization ability of large models in the engineering field has great development prospects and can well solve the limitation that current deep neural networks are limited to some established fields.

[0005] SAM has been proven to have good generalization capabilities for most natural images. However, while the SAM model is powerful, it also brings difficulties to special downstream tasks. The first is the deviation from the background knowledge domain. SAM pre-training knowledge and experience are concentrated on a large number of natural images. When the downstream tasks focus on special fields (such as engineering diseases), the target field will deviate from the knowledge distribution of natural images. Then the unprompted segmentation ability is likely to become unstable or even invalid. At this time, it is necessary to rely on the information provided by the prompt encoder. Secondly, the image encoder has a large number of parameters and cannot be trained. If the background domain knowledge needs to be fine-tuned, it is necessary to find a transfer learning method suitable for large models. Finally, artificial prompts are given for each segmentation, which is contrary to the concept of automated segmentation, and therefore poses a challenge to the downstream personalized field segmentation tasks. In summary, for downstream task development, it is necessary to solve the problems of target domain deviation and automated fine segmentation.

[0006] Therefore, the prior art needs to be improved. Summary of the invention

[0007] The technical problem to be solved by the present invention is that, in view of the defects of the prior art, the present invention provides a disease segmentation method and system based on a large language model to solve the problem of low accuracy of the existing disease segmentation method based on deep learning.

[0008] The technical solution adopted by the present invention to solve the technical problem is as follows: In a first aspect, the present invention provides a disease segmentation method based on a large language model, comprising: Obtaining a disease image to be segmented; Encoding the diseased image based on an image encoder that fuses a Transformer module and an Adapter module, and outputting an encoded image embedding; Generate object segmentation content hints based on a large language model semantic hint encoder; According to the encoded image block embedding and the target segmentation content prompt, a mask decoder is used to perform mask decoding to obtain a semantically segmented disease image.

[0009] In one implementation, the image encoder based on the fusion of the Transformer module and the Adapter module encodes the disease image and outputs the encoded image embedding, including: Dividing the disease image into a plurality of image blocks of fixed size, and flattening each image block; wherein each image block forms a vector after being flattened; The flattened image block is mapped to the preset feature space through a linear projection layer to obtain the image block embedding; Add position information to each image patch embedding; Embed the image patch after the addition position into the calculation through multiple of the Transformer modules, and supplement the frequency domain information based on the Adapter module to obtain the supplemented image patch embedding; Reduce the dimension of the supplemented image patch embedding through a convolutional module to obtain the encoded image embedding.

[0010] In one implementation, the embedding the image patch after the addition position into the calculation through multiple of the Transformer modules, and supplementing the frequency domain information based on the Adapter module to obtain the supplemented image patch embedding includes: Perform self-attention calculation on the image patch embedding after the addition position to obtain the context global information association; Provide non-linear feature representation ability through a feed-forward network to enhance the feature information at each position; Perform channel-wise multi-layer perception based on the Adapter module after self-attention calculation, and perform multi-layer perception based on the Adapter module parallel to the feed-forward neural network.

[0011] In one implementation, the performing channel-wise multi-layer perception based on the Adapter module after self-attention calculation, and performing multi-layer perception based on the Adapter module parallel to the feed-forward neural network includes: Perform a dimensionality reduction operation on the image patch embedding after self-attention calculation, perform non-linear processing through an activation function, then perform dimensionality increase to restore to the original dimension, and finally add the residual path and input it into the feed-forward neural network; In the feed-forward neural network, perform channel-wise and spatial multi-layer perception based on the parallel Adapter module; wherein, the channel-wise perception in the feed-forward neural network is the same as that of the Adapter module after multi-head attention calculation, and the spatial perception is performed by convolution and deconvolution.

[0012] In one implementation, the performing channel-wise and spatial multi-layer perception based on the parallel Adapter module includes: Use multiple MLP structures to perform non-linear scale alignment of the frequency domain information, and use a shared projection layer to inject the information into the input of the Transformer module to align the dimensions required for the input of the Transformer module.

[0013] In one implementation, the generating the target segmentation content prompt based on the large language model semantic prompt encoder includes: Generate natural language descriptions using multiple prompt templates; According to the natural language description, multiple category text embeddings are generated through a CLIP text encoder. Multiple calibrated text formats are tokenized and converted into corresponding tokens, and the average semantic description is obtained by taking the average of multiple tokens. Based on the average semantic description, positional encoding is added and input into the SAM prompt encoder to generate the target segmentation content prompt.

[0014] In one implementation, the method of using a mask decoder to perform mask decoding based on the encoded image patch embeddings and the target segmentation content prompt to obtain the diseased image after semantic segmentation includes: Copy the encoded image patch embeddings and the target segmentation content prompt multiple times; the number of copies is the same as the number of segmentation categories. Input the copied image patch embeddings and target segmentation content prompts into the same two-layer cross-attention module for information exchange. Perform a cross-attention calculation from token to image on the output embeddings once, then obtain the foreground mask tokens responsible for each category, and perform matrix multiplication with the corresponding deconvolved image embeddings respectively to obtain the predicted masks for the corresponding categories. Output the diseased image after semantic segmentation.

[0015] In a second aspect, the present invention provides a disease segmentation system based on a large language model, including: An image acquisition module for acquiring the diseased image to be segmented. An image encoder module for encoding the diseased image based on an image encoder that combines a Transformer module and an Adapter module, and outputting the encoded image embeddings. A semantic prompt encoder module for generating a target segmentation content prompt based on a large language model semantic prompt encoder. A mask decoder module for using a mask decoder to perform mask decoding based on the encoded image patch embeddings and the target segmentation content prompt to obtain the diseased image after semantic segmentation.

[0016] In a third aspect, the present invention provides a terminal, including: a processor and a memory. The memory stores a disease segmentation program based on a large language model. When the disease segmentation program based on the large language model is executed by the processor, it is used to implement the operations of the disease segmentation method based on the large language model as described in the first aspect.

[0017] Fourthly, the present invention also provides a medium, which is a computer-readable storage medium storing a disease segmentation program based on a language large model. When the disease segmentation program based on the language large model is executed by a processor, it is used to implement the operations of the disease segmentation method based on the language large model as described in the first aspect.

[0018] The present invention adopts the above technical solutions and has the following effects: 1) The present invention utilizes a powerful visual large model migration network to comprehensively, completely, and adaptively segment disease categories, so as to improve the ability to automatically identify new disease types with less sample training.

[0019] 2) The present invention can be fine-tuned according to specific engineering field objectives, thereby improving the success rate of target disease extraction; 3) The present invention combines and utilizes a pre-trained visual image-language model to achieve the segmentation ability of open vocabulary, automatically generates a text description of the information contained in the image, guides the image decoder to perform category segmentation, and facilitates subsequent image inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the structures shown in these drawings.

[0021] Figure 1 is a flowchart of the disease segmentation method based on the language large model in the present invention.

[0022] Figure 2 is a schematic diagram of the disease segmentation framework based on the language large model in the present invention.

[0023] Figure 3 is a functional schematic diagram of a terminal in an implementation manner of the present invention.

[0024] The implementation, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the object, technical solutions, and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0026] Exemplary Method The SAM (Segment Anything Model) has been proven to have good generalization ability for most natural images. However, while the SAM model is powerful, it also brings difficulties to special downstream tasks. Firstly, there is a deviation in the background knowledge domain. The pre-trained knowledge and experience of SAM are concentrated in a large number of natural images. When the downstream task focuses on a special field (such as engineering diseases), the target domain will deviate from the knowledge distribution of natural images. Then, the zero-shot segmentation ability is likely to become unstable or even ineffective. At this time, it is necessary to rely on the information provided by the prompt encoder. Secondly, the number of parameters of the image encoder is huge and cannot be trained. If it is necessary to fine-tune the background domain knowledge, it is necessary to find a transfer learning method suitable for large models. Finally, giving prompts manually for each segmentation is contrary to the concept of automatic segmentation, so it poses a challenge to the downstream personalized domain segmentation task. In summary, for the development of downstream tasks, it is indeed necessary to solve the problems of target domain deviation and automatic fine segmentation.

[0027] To address the above technical problems, in an embodiment of the present invention, a disease segmentation method based on a large language model is provided. The method encodes the disease image to be segmented by using an image encoder that integrates a Transformer module and an Adapter module, outputs the encoded image embedding, and generates a target segmentation content prompt based on the large language model semantic prompt encoder. Then, according to the encoded image patch embedding and the target segmentation content prompt, a mask decoder is used for mask decoding to obtain the disease image after semantic segmentation.

[0028] As Figure 1 shown, an embodiment of the present invention provides a disease segmentation method based on a large language model, including the following steps: Step S100, obtain a disease image to be segmented.

[0029] In this embodiment, in the problem of semantic segmentation of engineering surface diseases, the target domain is significantly deviated from the natural domain, and there are optimization tasks for the deployment of the SAM large model in downstream tasks. It is necessary to focus on solving the challenges of knowledge transfer learning and automatic semantic segmentation. Therefore, in this embodiment, a disease segmentation framework based on a large language model is proposed (as Figure 2 shown).

[0030] In this embodiment, the disease segmentation framework based on the large language model aims at disease images of construction projects (such as disease images caused by geological changes, natural disasters, traffic loads, construction errors, etc.) to perform the task of semantic segmentation of engineering surface diseases; after obtaining the disease image to be segmented, the image encoder that integrates a Transformer module and an Adapter module in this embodiment is used to encode the disease image, and the image embedding is output.

[0031] As Figure 1 shown, an embodiment of the present invention provides a disease segmentation method based on a large language model, including the following steps: Step S200, encode the disease image based on an image encoder that fuses a Transformer module and an Adapter module, and output the encoded image embedding.

[0032] In this embodiment, an improvement point of the proposed disease segmentation framework based on a large language model is that: in each Transformer module in the SAM encoder, a lightweight and parallel Adapter module is inserted. Without changing the original pre-trained parameters, by adding a very small number of trainable parameters, information in the target domain is merged into the segmentation network, thereby increasing the transferability of the image encoder. In addition, the frequency domain information of the input image is also emphasized. This part of the information can be selected as a specific filtering range according to the characteristics of the target domain. This part of the information is fused with the image embedding and aligned to the input dimension of each Transformer module through an MLP module. Injecting them into each layer of the Transformer module can enable the network to learn more information in the target frequency domain, thereby realizing fine-tuning training of the image encoder.

[0033] Specifically, in one implementation manner of this embodiment, step S200 includes the following steps: Step S201, divide the disease image into multiple image blocks of a fixed size, and flatten each image block; wherein, each flattened image block forms a vector; Step S202, map the flattened image blocks to a preset feature space through a linear projection layer to obtain image block embeddings; Step S203, add position information to each image block embedding; Step S204, calculate the image block embeddings added with positions through multiple of the Transformer modules, and supplement frequency domain information based on the Adapter module to obtain supplemented image block embeddings.

[0034] In this embodiment, the image encoder of the SAM model is built based on ViT (Vision Transformer, an image processing model based on the Transformer architecture). It has powerful global modeling capabilities and feature representation effects, and is responsible for extracting high-dimensional features from the input image to provide basic information for subsequent segmentation tasks. The basic process is as follows: 1) Image Blocking. The input image (i.e., the disease image) is sliced into multiple image blocks of a fixed size (e.g., 16*16 image blocks). After each image block is flattened, it forms a vector, which serves as the input to the subsequent Transformer module.

[0035] 2) Image Block Embedding. The flattened image blocks are mapped to a high-dimensional feature space (i.e., a preset feature space) through a linear projection layer.

[0036] 3) Adding Position Embedding. Position information is added to each image block embedding to ensure that the model can capture the spatial structure of the image.

[0037] 4) Transformer Computing Module. The image blocks go through 16 Transformer computing modules in total. Each module captures the global relationships between different regions in the image through the multi-head self-attention mechanism. Subsequently, the attention output is further processed using a feed-forward network to enhance the feature representation ability.

[0038] Specifically, in one implementation of this embodiment, step S204 includes the following steps: Step S204a, perform self-attention calculation on the image block embeddings after adding positions to obtain context global information associations; Step S204b, provide non-linear feature representation ability through a feed-forward network to enhance the feature information at each position; Step S204c, perform channel-wise multi-layer perception based on the Adapter module after self-attention calculation, and perform multi-layer perception based on the Adapter module parallel to the feed-forward neural network.

[0039] In this embodiment, an Adapter module is added to the Transformer module of the image encoder. It can obtain powerful transfer learning ability by only fine-tuning a small number of additional parameters, enabling the efficient transfer of a large pre-trained Transformer model to downstream tasks, as shown in Figure 2 as shown in (a) in the figure. The main body of the Transformer module can be divided into two major parts. One part is to perform self-attention calculation on the image embeddings to obtain context global information associations. The other part is to add a feed-forward network to provide stronger non-linear feature representation ability and enhance the feature richness at each position, making it easier for the model to capture complex patterns in the high-dimensional space.

[0040] In this embodiment, trainable Adapter modules are added to the tails of these two parts of the main body of the Transformer module to contain information in the transfer domain, which is used for fine-tuning specific tasks.

[0041] Specifically, in one implementation of this embodiment, the Adapter module after self-attention calculation performs channel-wise multi-layer perception, and the Adapter module parallel to the feed-forward neural network performs multi-layer perception, including: performing a dimensionality reduction operation on the image patch embedding after self-attention calculation, performing non-linear processing through an activation function, then performing dimensionality increase to restore to the original dimension, and finally adding the sum to the residual path and inputting it into the feed-forward neural network; in the feed-forward neural network, channel-wise and spatial multi-layer perception is performed based on the parallel Adapter module; wherein, the channel-wise perception in the feed-forward neural network is the same as that of the Adapter module after multi-head attention calculation, and the spatial direction is perceived by convolution and deconvolution.

[0042] Specifically, in this embodiment, the Adapter module after multi-head attention calculation mainly realizes channel-wise multi-layer perception, that is, performing a dimensionality reduction operation on the image embedding, then performing non-linear processing through an activation function, then performing dimensionality increase to restore to the original dimension, and finally adding the sum to the residual path and feeding it into the feed-forward network part. In the feed-forward network part, in this embodiment, an Adapter module is parallel to the feed-forward neural network to realize channel-wise and spatial multi-layer perception. Among them, the channel-wise perception of the Adapter module parallel to the feed-forward neural network is the same as that of the Adapter module after multi-head attention calculation. The spatial direction is obtained by convolution and deconvolution perception.

[0043] Specifically, in one implementation of this embodiment, the channel-wise and spatial multi-layer perception based on the parallel Adapter module includes: using multiple MLP structures to perform non-linear scale alignment of frequency domain information, and using a shared projection layer to inject information into the input of the Transformer module to align the dimensions required for input of the Transformer module.

[0044] In this embodiment, a parallel frequency domain Adapter module is constructed in the image encoder to supplement the ignored frequency domain information. The image filter structure is as Figure 2 shown in (b) in the figure.

[0045] Perform a two-dimensional Fourier transform in the image filter structure to obtain the frequency-domain information of the image. Perform an inverse two-dimensional Fourier transform on the filtered frequency-domain information to obtain the high-pass filtered image information, which is fused with the image patch embedding and used as the input of the MLP module. As an example, in this embodiment, two MLP structures are used to achieve non-linear scale alignment of the frequency-domain information, and then a shared projection layer is used to inject the information into the input of the Transformer module to align the dimensions required for the input of the Transformer module. In this simple and efficient way, the frequency-domain information of the original image is added to the image encoder layer by layer through a parallel module, thereby injecting the guiding information of the target task and better fine-tuning the network to downstream tasks.

[0046] Step S205: Reduce the dimension of the supplemented image patch embedding through a convolution module to obtain the encoded image embedding.

[0047] In this embodiment, after passing through the Transformer calculation module, the following processing flow is also required: 5) Image feature post-processing. Finally, reduce the dimension of the image embedding through convolution, and output the image embedding, that is, obtain the encoded image embedding.

[0048] In this embodiment, the Adapter module is fused in the image encoder Transformer module, which can obtain powerful transfer learning capabilities by only fine-tuning a small number of additional parameters, enabling the efficient transfer of large pre-trained Transformer models to downstream tasks; moreover, by constructing a parallel frequency-domain Adapter module in the image encoder to supplement the ignored frequency-domain information, as knowledge on the target task, it can be used as a hint for the target domain to guide the encoder, thereby enhancing the generalization ability of the model to downstream tasks.

[0049] As Figure 1 shown, an embodiment of the present invention provides a disease segmentation method based on a large language model, including the following steps: Step S300: Generate a target segmentation content hint based on a large language model semantic hint encoder.

[0050] In this embodiment, another improvement point of the proposed disease segmentation framework based on a large language model is that in the hint encoder, in order to achieve fully automated semantic segmentation, that is, without giving hints for each segmentation, this embodiment uses a well-pre-trained image-language large model, and uses a preset natural language as a hint during the training phase to generate corresponding image hint tokens to guide the segmentation work of the subsequent image decoder.

[0051] Specifically, in one implementation manner of this embodiment, step S300 includes the following steps: Step S301, generate natural language descriptions using multiple prompt templates; Step S302, according to the natural language description, generate multiple category text embeddings through the CLIP text encoder, tokenize multiple calibrated text formats and convert them into corresponding tokens, and take the average of the multi-tokens to obtain an average semantic description; Step S303, add positional encoding based on the average semantic description and input it into the SAM prompt encoder to generate the target segmentation content prompt.

[0052] In the original SAM prompt encoder, it can accept point prompts, box prompts, language prompts, and coarse mask prompts to guide segmentation for the mask decoder. However, for the engineering field, relevant prompts cannot be given in each segmentation prediction, which contradicts the original intention of the automated segmentation project. In addition, during the training process, it is time-consuming and laborious to perform point annotation, box annotation, or coarse mask annotation on the training images. In comparison, adding target semantics is simple. By determining a certain number of fixed-number text frameworks, the categories of the targets in the image can be added to the sentence pattern.

[0053] Therefore, in this embodiment, the large language pre-training model CLIP is combined in the prompt encoder, as shown in Figure 2 (c) below. Multiple prompt templates (Prompting) are used in the prompt encoder module to generate natural language descriptions, such as: "A photo of {} disease.", "There is {} disease in the scene.", "This is a picture containing {} disease.".

[0054] This embodiment generates multiple category text embeddings through the CLIP text encoder, tokenizes multiple calibrated simple text formats and converts them into token tokens, and takes the average of the multi-tokens to obtain an average semantic description, avoiding overfitting caused by a single prompt and enhancing the model's adaptability to different contexts. After adding positional encoding, it is input into the SAM prompt encoder and subsequently participates as a prompt token in the mask decoder. These semantic target information can better guide the mask decoder with information in the target domain as prompts, thereby achieving more accurate learning.

[0055] This embodiment realizes segment-level semantic alignment by introducing the pre-trained vision-language model CLIP in the prompt encoder, realizes the conversion from text prompts to text prompt embeddings, and fills in the missing target segmentation content prompts.

[0056] As Figure 1 shown, the embodiment of the present invention provides a disease segmentation method based on a language large model, including the following steps: Step S400: According to the encoded image patch embedding and the target segmentation content prompt, use a mask decoder to perform mask decoding to obtain a disease image after semantic segmentation.

[0057] In this embodiment, in addition to adding an additional semantic prompt module to the prompt encoder, an additional semantic prompt module is also added to the image decoder; by converting the instance segmentation ability into semantic segmentation ability, that is, in the module, first copy the image embedding, and the number is the total number of categories to be segmented. Similarly, copy the input embedding of the prompt token the same number of times. Input both into the mask decoder to finally obtain semantic segmentation masks of each category and the IoU confidence. The final predicted target mask is generated by the segmentation network, and a loss function is used for supervised training with the ground truth label to adjust the weight parameters in the adapter module of the image encoder and the mask decoder.

[0058] Specifically, in one implementation manner of this embodiment, step S400 includes the following steps: Step S401: Copy the encoded image patch embedding and the target segmentation content prompt multiple times; where the number of copies is the same as the number of segmentation categories; Step S402: Input the copied image patch embedding and target segmentation content prompt into the same two-layer cross-attention module for information exchange; Step S403: Perform a cross-attention calculation from token to image on the output embedding, then obtain the foreground mask tokens responsible for each category, and perform matrix multiplication with the corresponding image embeddings after deconvolution to obtain the predicted masks for the corresponding categories; Step S404: Output the disease image after semantic segmentation.

[0059] In this embodiment, in SAM, the original image passes through the image encoder to obtain an image embedding, and the user gives a target prompt to obtain a prompt token, which together with the IoU token, background mask token, and foreground mask token form the output embedding. The two are input into the mask decoder to obtain the final predicted mask and confidence score, as shown in (d) in Figure 2 shown. For each prompt input by the user, SAM will establish an independent output embedding and thus obtain a brand-new predicted mask. Each class specified by each prompt is independent of each other, which meets the requirements of instance segmentation. However, in the downstream engineering field, it is difficult to provide fine-grained prompts for each segmentation, and it may be necessary to obtain predicted masks of the same class but with non-adjacent appearance positions, and SAM cannot recognize and merge masks belonging to the same class, which poses a challenge to the transformation of downstream tasks. Therefore, this embodiment uses a simple but effective transformation method to enable SAM to complete the semantic segmentation task.

[0060] During the decoding process of this embodiment, compared with the original SAM mask decoder, the image embedding and the pruned input embedding are copied the same number of times as the number of categories. The purpose is to have corresponding image embeddings and output embeddings for separate reasoning for each predicted category. Subsequently, both are input into the same two-layer cross-attention module (Two-Way Transformer) for information exchange. Furthermore, a cross-attention calculation from tokens to the image is performed on the output embedding, and then the foreground Mask tokens responsible for each category are obtained. These are respectively matrix-multiplied with the corresponding deconvolved image embeddings to obtain the predicted mask for this category, that is, the segmentation result of the diseased image (the semantic segmentation masks and IoU confidence levels of various diseases).

[0061] In the decoder of this embodiment, output embeddings are designed for each category to predict masks. It should be noted that the decoder is equivalent to calculating a mask in parallel for each category, and the mask decoder is lightweight. Therefore, the increase in the number of categories will not significantly increase the computational load of the network.

[0062] This embodiment takes visual images and text prompts during training (not required during prediction) as inputs to achieve precise semantic segmentation of disease categories in the target engineering field, and can be widely applied to fields such as pixel-level localization of engineering appearance diseases, such as specific engineering scenarios like bridge cable disease extraction, tunnel lining disease extraction, and house appearance disease extraction.

[0063] The deformation scheme of this embodiment based on the language-image large model disease segmentation framework can be as follows: 1) Replace the image encoder with other modules based on convolutional or other Transformer frameworks for extracting image features.

[0064] 2) Replace the language-image pre-trained large model module with other modules with similar language-image alignment functions.

[0065] 3) Modify the image encoder for functional variants of modules for image classification, object detection, and instance segmentation tasks.

[0066] 4) Frame variants that modify the data input format types (such as image file format, annotation file format). Such frameworks have a similar structure to the framework of the present invention, and can meet the technical requirements by equivalent replacement or modification according to the technical solutions and concepts of this embodiment.

[0067] This embodiment achieves the following technical effects through the above technical solutions: 1) This embodiment uses a powerful visual large model migration network to comprehensively, completely, and adaptively segment disease categories to improve the ability to automatically identify new disease types with less sample training.

[0068] 2) This embodiment can be fine-tuned according to specific engineering field goals to improve the success rate of target disease extraction; 3) This embodiment combines and utilizes a pre-trained vision-language model to achieve the ability of open-vocabulary segmentation, automatically generates a text description of the information contained in the image, guides the image decoder for category segmentation, and facilitates subsequent image inspection.

[0069] Exemplary device Based on the above embodiments, the present invention further provides a disease segmentation system based on a large language model, including: An image acquisition module for acquiring a disease image to be segmented; An image encoder module for encoding the disease image based on an image encoder that combines a Transformer module and an Adapter module, and outputting an encoded image embedding; A semantic prompt encoder module for generating a target segmentation content prompt based on a large language model semantic prompt encoder; A mask decoder module for performing mask decoding using a mask decoder according to the encoded image patch embedding and the target segmentation content prompt to obtain a disease image after semantic segmentation.

[0070] This embodiment achieves the following technical effects through the above technical solutions: 1) This embodiment uses a powerful vision large model transfer network to comprehensively, completely, and adaptively segment disease categories to improve the ability to automatically identify new disease types with less sample training.

[0071] 2) This embodiment can be fine-tuned according to specific engineering field goals to improve the success rate of target disease extraction; 3) This embodiment combines and utilizes a pre-trained vision-language model to achieve the ability of open-vocabulary segmentation, automatically generates a text description of the information contained in the image, guides the image decoder for category segmentation, and facilitates subsequent image inspection.

[0072] Based on the above embodiments, the present invention further provides a terminal, and its principle block diagram can be as Figure 3 shown.

[0073] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected via a system bus; wherein, the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and the computer program in the storage medium; the interface is used to connect external devices; the display screen is used to display corresponding information; the communication module is used to communicate with a cloud server or other devices.

[0074] When the computer program is executed by the processor, it is used to implement the operations of the disease segmentation method based on a language large model.

[0075] Those skilled in the art can understand that Figure 3 The principle block diagram shown is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0076] In one embodiment, a terminal is provided, which includes: a processor and a memory. The memory stores a disease segmentation program based on a language large model. When the disease segmentation program based on the language large model is executed by the processor, it is used to implement the operations of the disease segmentation method based on the language large model as described above.

[0077] In one embodiment, a storage medium is provided, which stores a disease segmentation program based on a language large model. When the disease segmentation program based on the language large model is executed by the processor, it is used to implement the operations of the disease segmentation method based on the language large model as described above.

[0078] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and volatile memories.

[0079] In summary, the present invention provides a disease segmentation method and system based on a large language model, including: obtaining a disease image to be segmented; encoding the disease image by an image encoder integrating a Transformer module and an Adapter module to output an encoded image embedding; generating a target segmentation content prompt by a large language model semantic prompt encoder; and performing mask decoding by a mask decoder according to the encoded image patch embedding and the target segmentation content prompt to obtain the disease image after semantic segmentation. The present invention utilizes a powerful visual large model migration network to comprehensively, completely and adaptively segment disease categories, improving the accuracy of automatically identifying new disease types with less sample training.

[0080] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A disease segmentation method based on a large language model, characterized in that Including: Obtain the disease image to be segmented; Encode the disease image using an image encoder based on a fused Transformer module and an Adapter module, and output the encoded image embedding; Generate a target segmentation content prompt based on a large language model semantic prompt encoder; According to the encoded image patch embedding and the target segmentation content prompt, use a mask decoder to perform mask decoding to obtain the disease image after semantic segmentation.

2. The disease segmentation method based on a large language model according to claim 1, wherein, The encoding of the disease image by the image encoder based on the fused Transformer module and the Adapter module to output the encoded image embedding includes: Slice the disease image into multiple image patches of a fixed size and flatten each image patch; wherein, each flattened image patch forms a vector; Map the flattened image patches to a preset feature space through a linear projection layer to obtain image patch embeddings; Add position information to each image patch embedding; Perform calculations on the image patch embeddings with added positions through multiple of the Transformer modules, and supplement frequency domain information based on the Adapter module to obtain the supplemented image patch embeddings; Reduce the dimension of the supplemented image patch embeddings through a convolutional module to obtain the encoded image embedding.

3. The disease segmentation method based on a large language model according to claim 2, wherein The performing calculations on the image patch embeddings with added positions through multiple of the Transformer modules and supplementing frequency domain information based on the Adapter module to obtain the supplemented image patch embeddings includes: Perform self-attention calculation on the image patch embeddings with added positions to obtain context global information association; Provide non-linear feature representation ability through a feed-forward network to enhance the feature information at each position; Perform channel-wise multi-layer perception based on the Adapter module after self-attention calculation, and perform multi-layer perception based on the Adapter module parallel to the feed-forward neural network.

4. The disease segmentation method based on a large language model according to claim 3, wherein The performing channel-wise multi-layer perception based on the Adapter module after self-attention calculation and performing multi-layer perception based on the Adapter module parallel to the feed-forward neural network includes: Perform a dimensionality reduction operation on the image patch embeddings after self-attention calculation, perform non-linear processing through an activation function, then perform dimensionality restoration to the original dimension, and finally add the sum of the residual paths to input the feed-forward neural network; In the feed-forward neural network, perform channel-wise and spatial multi-layer perception based on the parallel Adapter module; wherein, the channel-wise perception in the feed-forward neural network is the same as that of the Adapter module after multi-head attention calculation, and the spatial perception is performed by convolution and deconvolution.

5. The disease segmentation method based on a large language model according to claim 4, wherein The performing channel-wise and spatial multi-layer perception based on the parallel Adapter module includes: Use multiple MLP structures to perform non-linear scale alignment of frequency domain information, and use a shared projection layer to inject the information into the input of the Transformer module to align the dimensions required for input to the Transformer module.

6. The disease segmentation method based on a large language model according to claim 1, wherein, The generating of the target segmentation content prompt based on the large language model semantic prompt encoder includes: Generate natural language descriptions using multiple prompt templates; According to the natural language description, generate multiple category text embeddings through the CLIP text encoder, tokenize multiple calibrated text formats and convert them into corresponding tokens, and take the average of multiple tokens to obtain an average semantic description; Based on the average semantic description, add positional encoding and input it into the SAM prompt encoder to generate the target segmentation content prompt.

7. The disease segmentation method based on a large language model according to claim 1, characterized in that According to the encoded image patch embeddings and the target segmentation content prompt, use the mask decoder to perform mask decoding to obtain the diseased image after semantic segmentation, including: Copy the encoded image patch embeddings and the target segmentation content prompt multiple times; where the number of copies is the same as the number of segmentation categories; Input the copied image patch embeddings and target segmentation content prompts into the same two-layer cross-attention module for information exchange; Perform a cross-attention calculation from tokens to images on the output embeddings, then obtain the foreground mask tokens responsible for each category, and perform matrix multiplication with the corresponding image embeddings after deconvolution to obtain the predicted masks for the corresponding categories; Output the diseased image after semantic segmentation.

8. A disease segmentation system based on a large language model, characterized in that, Including: An image acquisition module for acquiring the diseased image to be segmented; An image encoder module for encoding the diseased image based on an image encoder that combines a Transformer module and an Adapter module, and outputting the encoded image embeddings; A semantic prompt encoder module for generating the target segmentation content prompt based on a large language model semantic prompt encoder; A mask decoder module for performing mask decoding using the mask decoder according to the encoded image patch embeddings and the target segmentation content prompt to obtain the diseased image after semantic segmentation.

9. A terminal, characterized in that, Including: A processor and a memory, the memory stores a disease segmentation program based on a language large model, and when the disease segmentation program based on the language large model is executed by the processor, it is used to implement the operations of the disease segmentation method based on the language large model according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a disease segmentation program based on a language large model, and when the disease segmentation program based on the language large model is executed by a processor, it is used to implement the operations of the disease segmentation method based on the language large model according to any one of claims 1-7.

Citation Information

Cited By

  • Steel surface defect detection method and system based on CLIP model cross-domain learning

    CN120782754A

  • A steel surface defect detection method and system based on CLIP model cross-domain learning

    CN120782754B