Medical image segmentation method and system based on evidence-driven visual language model

Through evidence-driven visual language model, the evidence view of aggregated medical images and text information is solved, and the problem that images and text information cannot be deeply integrated in the prior art is achieved, achieving more accurate medical image segmentation and lesion area recognition.

CN120125818AActive Publication Date: 2025-06-10SHANDONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510197527.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In the existing medical image segmentation technology, image and text information cannot be deeply mined, resulting in inaccurate image segmentation and inaccurate acquisition of lesion areas, which affects disease judgment.

Method used

An evidence-driven visual language model is adopted to learn aggregated image-text perspectives through evidence, estimate modal gaps, and improve the accuracy and robustness of medical image segmentation through multimodal fusion.

Benefits of technology

It realizes more accurate medical image segmentation, improves the deep fusion ability of image and text information, enhances the ability to identify lesions, and improves the accuracy of disease judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125818A_ABST
    Figure CN120125818A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image segmentation method and system based on an evidence-driven visual language model, and relates to the technical field of medical image segmentation. The method comprises the following steps: acquiring a medical original image to be segmented; a multi-modal segmentation model is constructed, a deviation variance decomposition method is executed based on a difference matrix to calculate similarity loss, visual evidence embedding and text evidence embedding are respectively expressed as visual viewpoints and text viewpoints by introducing uncertainty viewpoints, viewpoint loss is calculated, and the similarity loss is calculated. The segmentation loss is calculated by utilizing the segmentation difference between the visual text fusion evidence and the real mask; and performing parameter optimization on the multi-modal segmentation model by using the total loss, and performing image segmentation on the medical original image to be segmented by using the multi-modal segmentation model. According to the method, evidence learning is introduced into the visual language model, and the modal gap between the image and the text is estimated by aggregating the image-text viewpoints after evidence conversion, so that multi-modal fusion is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and particularly to a medical image segmentation method and system based on an evidence-driven vision-language model. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] In the medical field, there is a large amount of image information and text information in medical images. In the prior art, the image and text information in medical images cannot be deeply mined and corresponded, resulting in inaccurate image segmentation, inability to obtain accurate lesion regions, and affecting doctors' judgment of the condition. Therefore, accurate segmentation of medical images is crucial for disease judgment.

[0004] Vision-language models (VLMs) can learn vision-language representations by adopting architectures similar to BERT and contrastive learning paradigms. Specifically, vision-language models use contrastive learning to optimize the model by distinguishing the similarities between different images and texts, enabling the model to better capture the complex correlation relationships between images and texts. However, current vision-language models have a modality gap problem, resulting in poor performance in downstream tasks. The modality gap specifically refers to the difference between images and texts. In the medical image segmentation task based on vision-language models, the representations of images and texts often converge into two different groups. Due to the essential differences between image and text data, there is a significant modality gap between them. This modality gap makes it difficult to accurately measure the similarity between images and texts, thereby affecting the effect of modality fusion and resulting in poor segmentation performance. Existing methods mainly use two independent encoders to extract cross-modal features to bridge the modality gap. After that, these features are embedded into a latent common space through a carefully designed objective function to achieve modality invariance. However, these methods have deficiencies: they rely on shallow interactions between modalities and cannot fully eliminate the gap between modalities. The narrow embedding space limits the expressive ability of complex semantic relationships between images and texts. Therefore, this shallow modality interaction method cannot effectively capture and fuse the deep correlations between images and texts. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a medical image segmentation method and system based on an evidence-driven vision-language model. By introducing evidence learning into the vision-language model, the modality gap between images and texts is estimated by aggregating the image-text viewpoints after evidence transformation, thereby realizing multi-modal fusion to improve the accuracy and robustness of medical image segmentation.

[0006] To achieve the above object, the present invention is implemented through the following technical solutions:

[0007] In a first aspect of the present invention, a medical image segmentation method based on an evidence-driven vision-language model is provided, including the following steps:

[0008] Obtain the original medical image to be segmented, where the original medical image contains image information and text information;

[0009] Construct a multi-modal segmentation model, and use an evidence learning algorithm to calculate the overall loss in the image segmentation process using a vision-language model. Among them, the overall loss includes similarity loss, view loss, and segmentation loss. Encode the image information for training into visual evidence embeddings and encode the text information into text evidence embeddings. Calculate the similarity loss between the visual evidence embeddings and the text evidence embeddings using a method based on the differential matrix to perform bias-variance decomposition. Represent the visual evidence embeddings and the text evidence embeddings as visual views and text views respectively by introducing uncertain views, and calculate the view loss according to the visual views and the text views. Fuse the visual evidence embeddings and the text evidence embeddings to obtain visual-text fusion evidence, and calculate the segmentation loss using the segmentation difference between the visual-text fusion evidence and the ground truth mask;

[0010] Optimize the parameters of the multi-modal segmentation model using the overall loss, and use the finally trained multi-modal segmentation model to perform image segmentation on the original medical image to be segmented.

[0011] Further, the specific steps of encoding the image information for training into visual evidence embeddings and encoding the text information into text evidence embeddings are:

[0012] Use a visual encoder to encode the image information to obtain visual evidence embeddings;

[0013] Use a text encoder to encode the text information to obtain text token embeddings;

[0014] Combine the visual evidence embeddings and the text token embeddings through a cross-attention module to obtain text evidence embeddings.

[0015] Even further, after obtaining the visual evidence embeddings and the text evidence embeddings, refine the evidence embeddings of the visual evidence embeddings and the text evidence embeddings. The specific steps are:

[0016] Use a non-local self-attention block to learn the cross-modal evidence affinity of the visual evidence embeddings and the text evidence embeddings to obtain a visual evidence affinity graph and a text evidence affinity graph;

[0017] Use a self-attention module to synthesize the visual evidence affinity graph and the text evidence affinity graph to obtain a global cross-modal affinity graph;

[0018] Affine the global cross-modal affinity graph on the visual evidence affinity graph and the text evidence affinity graph respectively to obtain the refined video cross-modal evidence embedding and the refined text cross-modal evidence embedding.

[0019] Furthermore, the specific steps for calculating the similarity loss between the visual evidence embedding and the text evidence embedding by using the deviation-variance decomposition method based on the difference matrix are as follows:

[0020] Introduce the uncertainty of the view, and use the deviation-variance decomposition method based on the difference matrix to calculate the inconsistency between the visual evidence embedding and the text evidence embedding, and obtain the inconsistency difference loss;

[0021] Calculate the InfoNCE loss of the similarity matrix of the visual evidence embedding and the similarity matrix of the text evidence embedding;

[0022] Add the inconsistency difference loss and the InfoNCE loss to obtain the similarity loss.

[0023] Even further, the specific steps for introducing the uncertainty of the view and using the deviation-variance decomposition method based on the difference matrix to calculate the inconsistency between the visual evidence embedding and the text evidence embedding are as follows:

[0024] Calculate the similarity between the visual evidence embedding and the text evidence embedding based on the uncertainty of the view to obtain the similarity matrix of the visual evidence embedding and the similarity matrix of the text evidence embedding;

[0025] Use the deviation-variance decomposition method based on the difference matrix to perform evidence difference similarity learning on the similarity matrix of the visual evidence embedding and the similarity matrix of the text evidence embedding respectively, and obtain the inconsistency difference loss.

[0026] Furthermore, the specific steps for calculating the segmentation loss by using the segmentation difference between the visual-text fusion evidence and the real mask are as follows:

[0027] Use the decoder to decode the refined video cross-modal evidence embedding and the refined text cross-modal evidence embedding to obtain the decoded visual evidence and the decoded text evidence;

[0028] Fuse the decoded visual evidence and the decoded text evidence to obtain the visual-text fusion evidence;

[0029] Introduce a segmentation loss between the visual-text fusion evidence and the real mask.

[0030] Even further, the specific steps for calculating the view loss according to the visual view and the text view are as follows:

[0031] Use the Dirichlet distribution mapping method to represent the decoded visual evidence and the decoded text evidence as the visual view and the text view respectively;

[0032] Fuse the visual view and the text view to obtain an aggregated view of the visual and text views;

[0033] Calculate the visual view loss based on the visual view, calculate the text view loss based on the text view, calculate the aggregated loss based on the aggregated view of the visual and text views, and add the visual view loss, the text view loss, and the aggregated loss to obtain the view loss.

[0034] The second aspect of the present invention provides a medical image segmentation system based on an evidence-driven vision-language model, including:

[0035] A data acquisition module configured to acquire the original medical image to be segmented, where the original medical image includes image information and text information;

[0036] A model training module configured to build a multi-modal segmentation model, calculate the overall loss in the image segmentation process using a vision-language model based on an evidence learning algorithm, where the overall loss includes a similarity loss, a view loss, and a segmentation loss, encode the image information for training into visual evidence embeddings and encode the text information into text evidence embeddings, calculate the similarity loss between the visual evidence embeddings and the text evidence embeddings using a method for performing bias-variance decomposition based on a difference matrix, represent the visual evidence embeddings and the text evidence embeddings as visual views and text views respectively by introducing an uncertainty view, calculate the view loss based on the visual view and the text view, fuse the visual evidence embeddings and the text evidence embeddings to obtain visual-text fusion evidence, and calculate the segmentation loss using the segmentation difference between the visual-text fusion evidence and the ground truth mask;

[0037] An image segmentation module configured to optimize the parameters of the multi-modal segmentation model using the overall loss and perform image segmentation on the original medical image to be segmented using the finally trained multi-modal segmentation model.

[0038] The third aspect of the present invention provides a medium on which a program is stored, and when the program is executed by a processor, it implements the steps in the medical image segmentation method based on an evidence-driven vision-language model as described in the first aspect of the present invention.

[0039] The fourth aspect of the present invention provides a device including a memory, a processor, and a program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the medical image segmentation method based on an evidence-driven vision-language model as described in the first aspect of the present invention.

[0040] The above one or more technical solutions have the following beneficial effects:

[0041] The present invention discloses a medical image segmentation method and system based on an evidence-driven vision-language model, and proposes an Evidence Affinity Map Generator (EAMG) to refine modality-specific evidence embeddings by learning a global cross-modal affinity map, thereby collecting complementary cross-modal evidence. In addition, Evidence Discrepancy Similarity Learning (EDSL) is proposed to collect consistent cross-modal evidence through the inconsistency of the bidirectional similarity matrix changes based on bias-variance decomposition. Finally, subjective logic is used to map the collected evidence into opinions, and a combination rule based on Dempster-Shafer theory is introduced for opinion aggregation to measure the modality gap for sufficient modality fusion.

[0042] In the present invention, a new vision-language model paradigm is proposed, and for the first time, an attempt is made to introduce evidence learning into the vision-language model to estimate the modality gap by aggregating cross-modal opinions and bridge the modality gap problem in cross-modal fusion between images and texts. Through simulation, in some embodiments, the present invention can achieve accurate segmentation of pneumonia lesion regions, lung infection regions, and breast cancer lesion regions.

[0043] In the present invention, an Evidence Affinity Map Generator is proposed to collect complementary cross-modal evidence by learning a global cross-modal affinity map to refine the evidence embeddings of two modalities.

[0044] In the present invention, an Evidence Discrepancy Similarity Learning is proposed to improve the consistency between cross-modal evidence embeddings by measuring the inconsistency of the changes in the similarity matrix.

[0045] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0047] Figure 1 is a flowchart of the medical image segmentation method based on the evidence-driven vision-language model in Embodiment 1 of the present invention;

[0048] Figure 2 is an overall framework diagram of the medical image segmentation method based on the evidence-driven vision-language model in Embodiment 1 of the present invention;

[0049] Figure 3 is a framework diagram of the Evidence Affinity Map Generator for cross-modal evidence affinity learning and integration in Embodiment 1 of the present invention;

[0050] Figure 4It is a framework diagram of evidence difference similarity learning based on deviation-variance decomposition of the difference matrix in the first embodiment of the present invention. Detailed implementation manners

[0051] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0052] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present invention. As used herein, unless otherwise clearly specified by the context, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof;

[0053] Embodiment 1:

[0054] The first embodiment of the present invention provides a medical image segmentation method based on an evidence-driven vision-language model, which bridges the cross-modal gap between images and texts by introducing evidence learning. Specifically, an evidence affinity graph generator (EAMG) is proposed, which refines modality-specific evidence embeddings by aggregating cross-modal evidence and learning a global cross-modal affinity graph, thereby collecting complementary cross-modal evidence. In addition, evidence difference similarity learning (EDSL) is proposed, which collects consistent cross-modal evidence based on the inconsistency of the bidirectional similarity matrix change of deviation-variance decomposition, and uses the inconsistency of the change of the similarity matrix to improve the consistency of cross-modal evidence embeddings, thereby further enhancing the fusion effect between modalities. Finally, subjective logic is used to map the collected evidence into viewpoints, and a combination rule based on Dempster-Shafer theory is introduced for viewpoint aggregation to measure the modal gap to achieve sufficient modal fusion.

[0055] The overall inventive concept of this embodiment: improving the segmentation accuracy by combining image and text information. Its core concept is to efficiently fuse visual and text evidence through multiple modules for cross-modal reasoning and optimization. As Figure 1 and Figure 2As shown in the figure, first, the image-text evidence embedding is extracted. The original image extracts the visual evidence embedding through a visual encoder based on the U-Net architecture, while the text description extracts the text evidence embedding through a text encoder based on the BioClinicalBERT architecture. To achieve the effective fusion of images and text, a cross-attention module is used to combine the visual and text embeddings. Subsequently, the generated evidence affinity graph is utilized to enhance the complementarity between the image-text evidence. The evidence affinity graph generator module learns the affinity of visual and text evidence through a self-attention mechanism, generating a global cross-modal pixel-level affinity graph to further refine the evidence embedding. After that, a differential evidence similarity learning loss is designed to enhance the consistency between the image-text evidence. The evidence difference similarity learning module optimizes the alignment of image and text evidence through a differential matrix to improve the cross-modal alignment accuracy. Then, the image-text evidence is decoded, and the refined visual and text evidence is fused through a decoding module. A segmentation loss is introduced between the fused visual-text evidence and the ground truth mask to optimize the visual-text joint inference. Then, the image-text evidence is converted into image-text views. The image-text views are aggregated according to Dempster-Shafer theory, and the aggregated view loss is calculated. The model aggregates visual and text views, quantifies the modality gap, and further enhances the reasoning ability. Finally, the differential evidence similarity learning loss, the aggregated view loss, and the segmentation loss are jointly optimized to train the multi-modal segmentation model. The entire framework improves the image segmentation accuracy by optimizing the loss functions of multiple modules and is applied to the medical image segmentation task, achieving the efficient collaborative processing of image and text information.

[0056] Specifically, it includes the following steps:

[0057] Step 1: Obtain the original medical image to be segmented, which contains image information and text information.

[0058] Step 2: Construct a multi-modal segmentation model and calculate the overall loss in the image segmentation process using a visual language model based on an evidence learning algorithm.

[0059] Step 2.1: Encode the image information used for training into visual evidence embeddings and encode the text information into text evidence embeddings.

[0060] Step 2.1.1: Use a visual encoder to encode the image information to obtain visual evidence embeddings.

[0061] In a specific implementation, for an image, the original image is fed into a visual encoder based on the U-Net architecture to directly extract visual evidence embeddings. In the visual encoder, the image is processed through multiple convolutional layers and pooling layers, gradually extracting deeper visual features, which are encoded as visual evidence embeddings. These embeddings can effectively represent the key information in the image.

[0062] Step 2.1.2: Encode the text information using a text encoder to obtain text token embeddings.

[0063] In a specific implementation, for the text, the corresponding text description is fed into a text encoder based on the BioClinicalBERT architecture, and the visual evidence embeddings are made to process the text token embeddings encoded by the text encoder using cross-attention mechanism. These text embeddings contain the syntactic and semantic information of the text description, which helps to understand the context of the medical image.

[0064] Step 2.1.3: Combine the visual evidence embeddings and the text token embeddings through a cross-attention module to obtain text evidence embeddings.

[0065] In a specific implementation, finally, to achieve effective fusion between the image and the text, the visual evidence embeddings and the text token embeddings are combined through a cross-attention module. The cross-attention module allows the model to dynamically refer to and utilize the relevant text information while processing the image, thereby generating more accurate text evidence embeddings.

[0066] Specifically, the image and the text are respectively fed into a visual encoder and a text encoder to extract image evidence embeddings and text evidence embeddings. The specific process of Step 2.1 is as follows:

[0067] Given an image-text pair {V, T}, the encoding paths of U-Net and BioClinicalBERT are used as the visual encoder and the text encoder to extract visual evidence embeddings and text evidence embeddings respectively, where For it is directly extracted by that is: For it is obtained through a cross-attention module (CA). Formally, for the visual evidence embeddings let process the text token embeddings x encoded by and then calculate its corresponding cross-modal text evidence embeddings T that is: that is:

[0068]

[0069] where They are all learnable matrices, and d is the eigen - dimension of the Q, K, and V matrices. is the set of real numbers, and α V2T refers to the attention matrix from image to text, and V2T (vision to text) represents image - to - text.

[0070] Step 2.2: After obtaining the visual evidence embedding and the text evidence embedding, refine the visual evidence embedding and the text evidence embedding. The specific steps are as follows:

[0071] Step 2.2.1: Use the non - local self - attention block to learn the cross - modal evidence affinity of the visual evidence embedding and the text evidence embedding, and obtain the visual evidence affinity graph and the text evidence affinity graph.

[0072] Step 2.2.2: Use the self - attention (Self - Attention, SA) module to synthesize the visual evidence affinity graph and the text evidence affinity graph to obtain the global cross - modal affinity graph.

[0073] In a specific implementation, in order to learn the multi - modal evidence affinity, a self - attention (Self - Attention, SA) module is used to integrate the two modality - specific affinity graphs with self - attention to synthesize two modality - specific affinity graphs. By integrating the two learned visual - perception evidence affinity and text - perception evidence affinity, a global cross - modal pixel - level affinity graph is generated for evidence embedding refinement, thereby strengthening the complementarity between the image and the text.

[0074] Step 2.2.3: Affine the global cross - modal affinity graph on the visual evidence affinity graph and the text evidence affinity graph respectively to obtain the refined video cross - modal evidence embedding and the refined text cross - modal evidence embedding.

[0075] Specifically, the evidence affinity graph generation module integrates the two learned visual - perception and text - perception evidence affinities to generate a global cross - modal pixel - level affinity graph for evidence embedding refinement, thereby strengthening the complementarity between the image and the text. As Figure 3 shown, the specific process of Step 2.2 is as follows:

[0076] Use the non - local self - attention block (NonLocal) to capture the semantic correlation of spatial positions according to the similarity between the feature vectors of any two positions, and learn the inter - modal correlation.

[0077]

[0078] Particularly, for The input image features are a tensor of size H×W×D, which are encoded into a triple Q, K, V through three 1×1 convolutional layers, and then such a triple is flattened to a size of HW×D. The visual evidence affinity graph is calculated through the dot product of Q and K. Each row of represents the similarity between a spatial position and all other spatial positions. The text evidence affinity graph is generated in a similar way to the visual evidence affinity graph. At the same time, in order to learn the multi-modal evidence affinity, a self-attention (SA) module is also used to synthesize two modality-specific affinity graphs, and then two spatial attention graphs are generated by the SA module. The visual evidence affinity graph and the text evidence affinity graph are aggregated into a global cross-modal affinity graph

[0079]

[0080] Among them, are the two learned spatial attention graphs. In order to refine the two modality-specific evidence embeddings respectively on and affine the global cross-modal affinity graph A evi to obtain two refined cross-modal evidence embeddings and

[0081]

[0082]

[0083] where affine(·) refers to the evidence affine operation and the operation used for the evidence embedding is the Hadamard product.

[0084] Through the evidence affine operator, based on the visual evidence embedding and the affinity graph A evi the refined affine visual evidence embedding calculated can be expressed as:

[0085]

[0086] Similarly, based on the text evidence embedding and the affinity graph A evi the refined affine visual evidence embedding calculated can be expressed as:

[0087]

[0088] Among them, h and w refer to the width and height of the affinity graph. The evidence affinity graph generator refines the cross-modal evidence embedding by learning the global pixel-level affinity graph, enhancing the complementarity of cross-modal evidence learning.

[0089] Step 2.3: Calculate the similarity loss between the visual evidence embedding and the text evidence embedding by using the deviation-variance decomposition method based on the difference matrix.

[0090] Step 2.3.1: Introduce the uncertainty of viewpoints, and use the deviation-variance decomposition method based on the difference matrix to calculate the inconsistency between the visual evidence embedding and the text evidence embedding, obtaining the inconsistency difference loss.

[0091] In a specific implementation, the visual evidence embedding and the text evidence embedding are fed into the evidence difference similarity learning module, which performs deviation-variance decomposition based on the difference matrix to learn the change inconsistency between the two cross-modal similarity matrices, thereby improving the alignment between the image and the text evidence embedding. At the same time, in order to further ensure the reliability of the similarity matrix, the visual viewpoint uncertainty and the text viewpoint uncertainty are respectively applied to the visual evidence embedding and the text evidence embedding.

[0092] Step 2.3.1.1: Calculate the similarity between the visual evidence embedding and the text evidence embedding based on the uncertainty of viewpoints, obtaining the similarity matrix of the visual evidence embedding and the similarity matrix of the text evidence embedding.

[0093] Step 2.3.1.2: Use the deviation-variance decomposition method based on the difference matrix to perform evidence difference similarity learning on the similarity matrix of the visual evidence embedding and the similarity matrix of the text evidence embedding respectively, obtaining the inconsistency difference loss.

[0094] Specifically, as Figure 4 shown, Step 2.3.1 includes: The evidence difference similarity learning module performs deviation-variance decomposition based on the difference matrix to learn the change inconsistency between the two bidirectional cross-modal similarity matrices, thereby enhancing the strong alignment between the image and the text evidence embedding. In order to further ensure the reliability of the similarity matrix, the visual viewpoint uncertainty u V and the text viewpoint uncertainty u T are respectively applied to the visual evidence embedding and the text evidence embedding So for and their evidence similarity matrices s ij and can be calculated by the following formula:

[0095]

[0096]

[0097] Among them, s ij represents the similarity matrix of visual evidence embedding, represents the similarity matrix of text evidence embedding, cos(·) measures the cosine similarity, and i and j represent the i-th image and the j-th text respectively. and represent two bidirectional cross-modal similarity matrices respectively, where B is the batch size.

[0098] To measure the change inconsistency between two similarity matrices [S V2T , S T2V , subtract S T2V from S V2T to obtain their difference matrix S diff , and then perform bias-variance decomposition based on S diff .

[0099] S diff = f BVD (S V2T , S T2V ) (10).

[0100] The bias-variance decomposition is defined as follows: Let S V2T and S T2V be the bidirectional similarity matrices respectively. The bias-variance decomposition function f BVD is derived as follows:

[0101]

[0102] Among them, Var(·) and Cov(·) represent the expectation, variance, and covariance operations respectively. The calculated bias and variance measure the change inconsistency between two cross-modal similarity matrices. Therefore, regard as the inconsistency difference loss of the similarity matrix

[0103] Step 2.3.2: Calculate the InfoNCE loss of the similarity matrix of visual evidence embedding and the similarity matrix of text evidence embedding.

[0104] Calculate the InfoNCE (Information Noise-Contrastive Estimation) loss of the above two cross-modal similarity matrices to maximize the mutual information between real image-text pairs.

[0105]

[0106] Among them, is the image-to-text InfoNCE loss, is the text-to-image InfoNCE loss, and τ is the temperature hyperparameter.

[0107] Step 2.3.3: Add the inconsistency difference loss and the InfoNCE loss to obtain the similarity loss which is used to enhance the strong matching between the image and text evidence embeddings.

[0108] Specifically, the overall objective of EDSL is the sum of the InfoNCE loss and the inconsistency loss.

[0109]

[0110] Among them, λ 1 and λ 2 are hyperparameters used to balance these two losses.

[0111] EDSL is based on the differential matrix for bias-variance decomposition, performs consistency learning on the cross-modal similarity matrix, and enhances the consistency of cross-modal evidence.

[0112] Step 2.4: Fuse the visual evidence embedding and the text evidence embedding to obtain the visual-text fusion evidence, and calculate the segmentation loss using the segmentation difference between the visual-text fusion evidence and the ground truth mask.

[0113] In a specific implementation, the refined cross-modal visual evidence embedding and text evidence embedding are respectively fed into the visual decoder and the text decoder to obtain the decoded visual evidence and text evidence, and then the two are fused by addition. A segmentation loss is introduced between the fused visual-text evidence and the ground truth mask, aiming to help the model better perform joint reasoning of images and texts during training and optimize the visual-text fusion of the model, thereby improving the accuracy of the image segmentation task, especially when combining text information.

[0114] Step 2.4.1: Use the decoder to decode the refined video cross-modal evidence embedding and the refined text cross-modal evidence embedding to obtain the decoded visual evidence and the decoded text evidence

[0115] Step 2.4.2: Fuse the decoded visual evidence and the decoded text evidence by addition to obtain the visual-text fusion evidence.

[0116] Step 2.4.3: Introduce a segmentation loss between the visual-text fusion evidence and the pre-annotated ground truth mask.

[0117] Specifically, the segmentation loss between the fused visual-text evidence and the ground truth template is introduced.

[0118] where m represents the m-th pixel in the mask. W and H are the width and height of the image, is the cross-entropy loss, is the evidence of the m-th pixel point in the image, is the evidence of the m-th pixel point in the text, y m is the value of the m-th pixel point in the ground truth mask.

[0119] Step 2.5: Represent the visual evidence embedding and the text evidence embedding as visual perspectives and text perspectives respectively by introducing the uncertainty perspective, and calculate the perspective loss based on the visual perspective and the text perspective.

[0120] In a specific implementation, to measure the modality gap based on the aggregated perspective, the decoded visual evidence and text evidence are represented as visual perspectives and text perspectives to quantify the Dirichlet distribution uncertainty, where each perspective reflects the reasoning result of the modality evidence under the given question. Finally, the visual perspective and the text perspective are aggregated into an overall perspective, that is, the comprehensive reasoning result combining visual and text information.

[0121] Step 2.5.1: Represent the decoded visual evidence and the decoded text evidence as visual perspectives and text perspectives respectively using the Dirichlet distribution mapping method.

[0122] Step 2.5.2: Fuse the visual perspective and the text perspective to obtain the aggregated perspective of the visual and text perspectives.

[0123] Step 2.5.3: Calculate the visual perspective loss according to the visual perspective, calculate the text perspective loss according to the text perspective, calculate the aggregated loss according to the aggregated perspective of the visual and text perspectives, and add the visual perspective loss, the text perspective loss and the aggregated loss to obtain the perspective loss.

[0124] Specifically, Step 2.5 includes: To measure the modality gap based on the aggregated perspective, the decoded visual evidence and the text evidence are represented as the visual perspective o V and the text perspective o T to quantify the distribution uncertainty in Dir(μ V |α V ) and Dir(μ T |α T ).

[0125] In evidence learning, it is often necessary to infer the probability distribution of a multi-class problem based on data. The Dirichlet distribution can be used as the prior distribution of these classes, which helps to smooth the data and enables the model to handle the sparsity problem.

[0126] According to subjective logic, a principle method for probabilistic reasoning under uncertainty, the visual Dirichlet distribution of class probabilities Dir(μ V |α V ) is determined by the evidence , where C is the number of classes. The distribution parameters are derived by α V = e V + 1. Then, the visual Dirichlet distribution is mapped to the visual opinion satisfying:

[0127]

[0128] where, is the belief mass of class c, is the sum of the Dirichlet distribution, measures the uncertainty of the Dirichlet distribution.

[0129] The finally predicted pixel-level probability is the expectation of the Dirichlet distribution, that is where μ V is the original predicted probability. Assuming that the pixel-level sample cannot provide any evidence for the decision, that is, e V = 0. According to the definitions of a V , S V and μ V , the uncertainty μ V is negatively correlated with the sum of the evidence. Therefore, such pixel-level samples will produce high uncertainty. The reasoning process of the text opinion o = {b 1 , b 2 ,..., b C , u} = o V ⊕ o T is the same as that of vision.

[0130] To obtain the aggregated opinion based on visual and text opinions, following the Dempster-Shafer evidence theory, we use the belief fusion operator to aggregate the visual opinion o V and the text opinion o T . Specifically, for and the aggregated opinion o = {b 1 , b 2 , …, b C , u} = o V ⊕ oT Derived from the following formula:

[0131]

[0132] where, is the normalization factor.

[0133] To calculate the losses of vision, text, and aggregated views, the integrated cross-entropy loss is used. The loss of the vision view is given by the following formula:

[0134]

[0135] where, y is the one-hot label, and m is the digamma function. The overall view loss of vision, text, and integrated views is given by the following formula:

[0136]

[0137] where, represents the view loss, and is implemented in the same way as In cross-modal joint reasoning, is used to measure the modality gap. The modality gap refers to the difference between two modality features, such as the difference between image features and text features. Measuring and reducing the modality gap through evidence learning helps to enhance the consistency between modalities and improve the multi-modal fusion performance.

[0138] Step 2.6: The overall loss includes the similarity loss, view loss, and segmentation loss. Therefore, the overall loss is expressed as:

[0139]

[0140] where, is the overall loss, ω 1 , ω 2 and ω 3 are hyperparameters for balancing the three objective functions.

[0141] Step 3: Use the overall loss to optimize the parameters of the multi-modal segmentation model, and use the finally trained multi-modal segmentation model to perform image segmentation on the original medical images to be segmented. Among them, applying the model trained in this embodiment to medical images such as pneumonia, pulmonary infection, and breast cancer for segmentation tasks can achieve accurate segmentation of pneumonia lesion areas, pulmonary infection areas, and breast cancer lesion areas.

[0142] In a specific embodiment, the parameters of the learnable visual encoder, visual decoder, and text decoder are optimized according to the visual perspective loss, text perspective loss, aggregated perspective loss, segmentation loss, and the loss of the evidence difference similarity learning module. Specifically, optimization means using the overall loss for backpropagation to optimize the parameters of the encoder and decoder. It should be noted that the text encoder uses a pre-trained encoder, and during training, the parameters of the text encoder are frozen, so the parameters of the text encoder are not optimized.

[0143] The finally trained multi-modal segmentation model is used for medical image segmentation tasks. The evidence-driven visual language model network can be trained in an end-to-end manner.

[0144] The method of the present invention proposes a new visual-language model paradigm - evidence-driven visual language model. Traditional cross-modal learning methods often face differences between different modalities, which makes the effective fusion of information more difficult. This method aims to solve the problem of the modality gap between current images and texts by innovatively introducing evidence learning technology, thereby achieving more efficient cross-modal fusion. It not only enhances the connection between images and texts during the learning process but also makes the integration of cross-modal data closer and more efficient.

[0145] The method of the present invention proposes an evidence affinity graph generation method, which aims to collect and integrate cross-modal evidence. By learning a global cross-modal affinity graph, this method further refines the evidence embeddings specific to the two modalities. Specifically, through a detailed mapping relationship, the features of different modalities can complement each other, thereby improving the quality of cross-modal fusion. In addition, by strengthening the relevance of evidence, the inconsistency of information between different modalities is effectively reduced, and the accuracy of the fusion result is improved.

[0146] The method of the present invention proposes an evidence difference similarity learning method, which aims to ensure the consistency of evidence between different modalities. By measuring the change inconsistency of the similarity matrix, this method enhances the alignment degree between cross-modal evidence embeddings. In other words, this method guarantees the semantic consistency of image and text information by comparing the evidence similarities of different modalities. Finally, the collected cross-modal evidence is transformed into an estimate of the modality gap, providing an accurate basis for further modality integration.

[0147] Embodiment 2:

[0148] Embodiment 2 of the present invention provides a medical image segmentation system based on an evidence-driven visual language model, including:

[0149] A data acquisition module, configured to acquire a medical original image to be segmented, where the medical original image includes image information and text information;

[0150] A model training module, configured to build a multimodal segmentation model, calculate the overall loss in the image segmentation process using a vision-language model based on an evidence learning algorithm, where the overall loss includes a similarity loss, an opinion loss, and a segmentation loss, encode the image information for training into visual evidence embeddings and encode the text information into text evidence embeddings, calculate the similarity loss between the visual evidence embeddings and the text evidence embeddings using a method for performing bias-variance decomposition based on a differential matrix, represent the visual evidence embeddings and the text evidence embeddings as visual opinions and text opinions respectively by introducing uncertain opinions, calculate the opinion loss according to the visual opinions and the text opinions, fuse the visual evidence embeddings and the text evidence embeddings to obtain visual-text fusion evidence, and calculate the segmentation loss using the segmentation difference between the visual-text fusion evidence and the ground truth mask;

[0151] An image segmentation module, configured to optimize the parameters of the multimodal segmentation model using the overall loss, and perform image segmentation on the medical original image to be segmented using the finally trained multimodal segmentation model.

[0152] Embodiment III:

[0153] Embodiment III of the present invention provides a medium, on which a program is stored, and when the program is executed by a processor, the steps in the medical image segmentation method based on an evidence-driven vision-language model as described in Embodiment I of the present invention are implemented.

[0154] Embodiment IV:

[0155] Embodiment IV of the present invention provides a device, including a memory, a processor, and a program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the medical image segmentation method based on an evidence-driven vision-language model as described in Embodiment I of the present invention are implemented.

[0156] The steps involved in Embodiments II, III, and IV above correspond to those in Method Embodiment I, and the specific implementation manners can refer to the relevant description part of Embodiment I.

[0157] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0158] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A medical image segmentation method based on evidence-driven visual language model, characterized in that: The following steps are involved: Acquire an original medical image to be segmented, where the original medical image contains image information and text information; A multimodal segmentation model is constructed. The overall loss in the image segmentation process is calculated using a visual language model based on an evidence learning algorithm. The overall loss includes similarity loss, viewpoint loss, and segmentation loss. The image information used for training is encoded into visual evidence embedding and the text information is encoded into text evidence embedding. The similarity loss of visual evidence embedding and text evidence embedding is calculated using a bias-variance decomposition method based on a difference matrix. The visual evidence embedding and text evidence embedding are represented as visual viewpoint and text viewpoint, respectively, by introducing uncertain viewpoints. The viewpoint loss is calculated based on the visual viewpoint and text viewpoint. The visual evidence embedding and text evidence embedding are fused to obtain visual-text fusion evidence. The segmentation loss is calculated using the segmentation difference between the visual-text fusion evidence and the true mask. The overall loss is used to optimize the parameters of the multimodal segmentation model, and the finally trained multimodal segmentation model is used to perform image segmentation on the original medical images to be segmented.

2. The medical image segmentation method based on the evidence-driven visual language model according to claim 1, characterized in that: The specific steps of encoding image information for training into visual evidence embedding and text information into text evidence embedding are: Encode the image information using a visual encoder to get visual evidence embedding; Encode the text information using a text encoder to obtain text tag embedding; The visual evidence embedding is combined with the textual tag embedding through a cross-attention module to obtain the textual evidence embedding.

3. The medical image segmentation method based on the evidence-driven visual language model according to claim 2, characterized in that: After obtaining the visual evidence embedding and the text evidence embedding, the visual evidence embedding and the text evidence embedding are refined. The specific steps are as follows: The non-local self-attention block is used to learn the cross-modal evidence affinity of visual evidence embedding and textual evidence embedding, and the visual evidence affinity graph and textual evidence affinity graph are obtained; The visual evidence affinity graph and the textual evidence affinity graph are synthesized using the self-attention module to obtain a global cross-modal affinity graph. The global cross-modal affinity graph is affine-mapped on the visual evidence affinity graph and the textual evidence affinity graph to obtain the refined video cross-modal evidence embedding and the refined text cross-modal evidence embedding respectively.

4. The medical image segmentation method based on the evidence-driven visual language model according to claim 1, characterized in that: The specific steps of calculating the similarity loss of visual evidence embedding and text evidence embedding using the difference matrix-based bias-variance decomposition method are as follows: The uncertainty of opinions is introduced, and the difference matrix is ​​used to perform the bias-variance decomposition method to calculate the inconsistency between the visual evidence embedding and the text evidence embedding, and the inconsistency difference loss is obtained; Compute the InfoNCE loss of the similarity matrix of visual evidence embedding and the similarity matrix of textual evidence embedding; The similarity loss is obtained by adding the inconsistency difference loss and the InfoNCE loss.

5. The medical image segmentation method based on the evidence-driven visual language model according to claim 4, characterized in that: The specific steps of introducing the uncertainty of opinions and using the difference matrix to perform the bias-variance decomposition method to calculate the inconsistency between visual evidence embedding and text evidence embedding are as follows: The similarity between visual evidence embedding and text evidence embedding is calculated based on the uncertainty of viewpoints, and the similarity matrix of visual evidence embedding and the similarity matrix of text evidence embedding are obtained; The difference matrix is ​​used to perform bias-variance decomposition method to learn the evidence difference similarity of the similarity matrix of visual evidence embedding and the similarity matrix of textual evidence embedding, respectively, and the inconsistency difference loss is obtained.

6. The medical image segmentation method based on evidence-driven visual language model according to claim 1, characterized in that: The specific steps of calculating the segmentation loss using the segmentation difference between the visual text fusion evidence and the true mask are: Using a decoder to decode the refined video cross-modal evidence embedding and the refined text cross-modal evidence embedding to obtain decoded visual evidence and decoded text evidence; The decoded visual evidence and the decoded text evidence are fused to obtain visual-text fusion evidence; and a segmentation loss is introduced between the visual-text fusion evidence and the true mask.

7. The medical image segmentation method based on the evidence-driven visual language model according to claim 6, characterized in that: The specific steps for calculating the opinion loss based on visual opinion and textual opinion are: The decoded visual evidence and the decoded textual evidence are represented as visual opinions and textual opinions respectively using Dirichlet distribution mapping method; Fusing visual and textual views to obtain an aggregated view of visual and textual views; The visual opinion loss is calculated according to the visual opinion, the textual opinion loss is calculated according to the textual opinion, the aggregate loss is calculated according to the aggregate opinion of the visual and textual opinions, and the visual opinion loss, textual opinion loss and aggregate loss are added together to get the opinion loss.

8. A medical image segmentation system based on evidence-driven visual language model, characterized in that: include: A data acquisition module is configured to acquire an original medical image to be segmented, where the original medical image contains image information and text information; A model training module is configured to build a multimodal segmentation model, calculate the overall loss in the image segmentation process using a visual language model based on an evidence learning algorithm, wherein the overall loss includes similarity loss, viewpoint loss and segmentation loss, encode image information used for training into visual evidence embedding and encode text information into text evidence embedding, calculate the similarity loss of visual evidence embedding and text evidence embedding using a deviation variance decomposition method based on a difference matrix, express the visual evidence embedding and text evidence embedding as visual viewpoint and text viewpoint respectively by introducing uncertainty viewpoints, calculate the viewpoint loss based on the visual viewpoint and text viewpoint, fuse the visual evidence embedding and the text evidence embedding to obtain visual-text fusion evidence, and calculate the segmentation loss using the segmentation difference between the visual-text fusion evidence and the true mask; The image segmentation module is configured to optimize the parameters of the multimodal segmentation model using the overall loss, and to perform image segmentation on the original medical image to be segmented using the finally trained multimodal segmentation model.

9. A computer-readable storage medium, characterized in that: A plurality of instructions are stored therein, and the instructions are suitable for being loaded by a processor of a terminal device and executing the medical image segmentation method based on an evidence-driven visual language model as described in any one of claims 1-7.

10. A terminal device, characterized in that: It includes a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded by the processor and executing the medical image segmentation method based on the evidence-driven visual language model described in any one of claims 1-7.

Citation Information

Patent Citations

  • Medical image segmentation method and system based on double-branch embedded attention mechanism

    CN116309650A

  • Semi-supervised medical image segmentation method and system based on visual language model

    CN118115516A

  • Visual intention understanding method and system based on uncertainty cross-granularity evidence feature fusion network, and storage medium

    CN118196592A

  • Medical bill image text detection method based on feature enhancement and multi-scale feature fusion

    CN118887685A

  • Training of text and image models

    EP4266195A1