Three-dimensional medical image visual language model pre-training method for reconstructing hybrid strategy
By adopting reconstruction hybrid strategy and semantic perception fusion strategy in the medical visual language model, combining large language models to extract medical report information, and performing multi-task joint learning of image and text features, the problem of poor text and feature fusion in the existing medical visual language model is solved, and the performance of downstream tasks is significantly improved.
Patent Information
- Application Number
- CN202510119149.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing medical visual language models are difficult to generate high-quality text in medical reports, and the lack of effective fusion between image and text features leads to poor performance in downstream tasks.
The three-dimensional medical image visual language model pre-training method adopts the reconstruction hybrid strategy, and the diagnostic and attribute information in medical reports is extracted through large language models, and combined with semantic perception fusion strategy and advanced mask reconstruction tasks to perform multi-task joint learning of image and text features.
It significantly improves the performance of downstream medical visual or language tasks, generates high-quality text, and the alignment between image and text features is more effective, improving the model's performance in downstream tasks.
Smart Images

Figure CN119943252A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image computing technology, and in particular relates to a three-dimensional medical image visual language model pre-training method of a reconstruction hybrid strategy. Background Art
[0002] Visual language model is widely defined as a multimodal model that can be learned on large-scale image-text pairs to improve multi-granularity downstream vision and language tasks. Visual language model usually consists of three elements: image encoder, text encoder and learning strategy to fuse the information of the two encoders. Since the loss function is designed around the above model structure and learning strategy, it is necessary to tightly couple the above key elements together. Among the traditional methods, the most representative method is CLIP, which shows great potential for learning mutual information between visual and language data. More recent studies have shown that fine-grained context alignment is conducive to the model learning more representative representations. Among them, the BLIP method reconstructs text by leveraging visual semantic context. However, in more challenging medical fields, such as medical reports, the accuracy requirements are more stringent, and the above methods are difficult to meet this demand.
[0003] To address this problem, recent medical visual language models have improved the efficiency of model learning through different pre-training methods, such as contrastive learning methods for medical visual representations of image-text pairs and few-shot self-supervised contrastive learning pre-training methods, which pre-train the model by directly maximizing the mutual information between global representations. SAT proposes fine-grained features that align paired image patches and words. BioVIL uses paired data samples to try to understand complex medical reports. MedKLIP uses a triple extraction module as an additional supervisory signal to extract medically relevant information.
[0004] Patent document with application number 202210903886.5 discloses a pre-training method and device for a medical multimodal model. The target medical image and text sample data bootstrapping method cannot generate high-quality medical-text pairs, and the multimodal hybrid codec MED does not fuse features well, and lacks perception between image and text features. This makes the pre-trained model have no performance advantage in downstream tasks.
[0005] Patent document No. 202410135051.9 discloses a context-aware medical visual language pre-training method, system and application. The distillation report does not generate high-quality prompt text, and the multi-scale context fusion method that integrates visual features and text embedding does not align visual and text features well in the embedding space, which makes the pre-trained model have no performance advantage in downstream tasks.
[0006] Therefore, a pre-training method is needed that can generate high-quality text and better integrate image features and text features. Summary of the invention
[0007] The purpose of the present invention is to provide a visual language model pre-training method based on a large language model and a reconstruction hybrid strategy, using a pre-training method that includes a large language model to extract text information strategy, a semantic perception fusion strategy and a high-order mask reconstruction task joint learning to improve the performance of various downstream medical vision or language tasks, so as to overcome the shortcomings of the prior art that image and text features are not well integrated and there is a lack of effective means of perception between image and text features.
[0008] In order to achieve the above object, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for pre-training a 3D medical image visual language model using a reconstruction hybrid strategy, which specifically comprises the following steps: S1, constructing a medical image-text pair dataset, wherein the medical image-text pair dataset includes three-dimensional medical images and medical field reports; S2, extract and generate text information: extract the diagnosis and attribute information of the medical field report of the medical image text pair dataset in S1 to generate the text corresponding to the template; S3, text feature generation: perform a random mask operation on the text generated by S2 to obtain a text mask, and input the obtained text mask into the text encoder to generate text features; S4, image feature generation: preprocessing the three-dimensional medical image of the medical image text pair data set in S1, and performing a random mask operation on the preprocessed three-dimensional medical image to obtain a three-dimensional medical mask image, and inputting the obtained three-dimensional medical mask image into an image encoder to generate image features; S5, input the text features obtained in S3 into the text decoder to obtain the reconstructed text; S6, inputting the image features obtained in S4 into the image decoder to obtain a reconstructed image; S7, semantic-aware fusion strategy: Use the cross-attention mechanism to fuse the text features and image features obtained in S2 and S3 to generate new text features; S8, text reconstruction task: jointly calculate the text reconstruction loss by combining the text features generated in S3 and the reconstructed text obtained in S5; S9, image reconstruction task: jointly calculate the image reconstruction loss by combining the image features generated in S4 and the reconstructed image obtained in S6; S10, image-text pairing task: the image features in S4 and the new text features obtained in S7 are combined for comparative learning, and the image-text pairing loss is calculated; S11, multi-task joint learning: weighted sum of the losses of different tasks in S8, S9 and S10.
[0009] Furthermore, since 3D medical datasets are difficult to obtain and have high annotation costs, there is currently no large-scale medical image-text pair dataset. Therefore, we first collect a large number of publicly available 3D medical image datasets and medical field report data, and then pair the image data with the medical reports to build a large-scale medical image-text pair dataset. The specific steps of S1 to build a large-scale medical image-text pair dataset include: S11, collect and integrate different publicly available 3D medical image datasets; S12, collect medical domain reports from major medical domain knowledge sources; S13, checking the objects in each 3D medical image dataset and assigning a medical field report to it, first checking each object in each dataset and assigning a medical report to it, which ensures the accuracy and clarity of the medical text between the datasets.
[0010] Furthermore, due to the complexity, diversity, and professionalism of medical report data, it is difficult to select appropriate and efficient prompts from medical reports by manual methods. Therefore, it is necessary to use the powerful generalization of large language models to generate high-quality diagnostic prompts based on medical reports and templates. Secondly, considering that medical reports not only contain diagnostic information, but also contain description information of corresponding medical image textures and features, the present invention uses a large language model to extract the description information of medical image attributes in medical reports, and generates usable attribute prompts based on templates. The specific steps of extracting and generating text information and generating text features in S2 and S3 are: Use a large language model and fine-tune it on a publicly available large-scale medical language dataset; S22, inputting the medical field report and the corresponding template information into the trained and fine-tuned large language model, so that the trained and fine-tuned large language model can extract the diagnosis and attribute information of the corresponding medical field report and generate the diagnosis text and attribute text representation corresponding to the template; S23, using the diagnostic text and attribute text corresponding to the generated template as prompts to perform a random masking operation, where the masking probability is fixed at 20%, to obtain a text mask, and input the text mask into the text encoder to generate text features.
[0011] Furthermore, due to the characteristics of sparse data, poor imaging quality, and imbalanced categories of 3D medical image data, 3D medical image data cannot be used directly for training. It is often necessary to preprocess and generate data with rich visual information to improve the speed of model training and enable the model to have good generalization ability. The specific steps of preprocessing 3D medical images and generating image features in S4 are as follows: S41, setting a central area based on the center of the three-dimensional medical image; S42, selecting a point in the central area as the central point for cropping to obtain a three-dimensional medical image block x, where x represents a single image block; S43, performing a random masking operation on the obtained three-dimensional medical image block x, wherein the masking probability is fixed at 20%, to obtain a three-dimensional medical masked image, and inputting the obtained three-dimensional medical masked image into an image encoder to generate image features.
[0012] Furthermore, the specific steps for obtaining text features in S3 are: The obtained diagnosis text and attribute text are segmented into words; Add special words at the beginning and end of the diagnostic text sequence and attribute text sequence respectively; Convert the segmented diagnostic text sequence and attribute text sequence into index identifiers corresponding to the vocabulary in the large language model; Create attention masks to indicate actual words, special words, and fillers in diagnosis text sequences and attribute text sequences; Generate a type mask and assign different type identifiers to the diagnostic text sequence and attribute text sequence in each sentence; Create positional encodings to provide the large language model with information about the position of words in the sentence; In the text encoder, each word is mapped to a vector space through a linear projection layer to extract text features.
[0013] Furthermore, the specific steps of obtaining image features in S4 are as follows: Segmenting the preprocessed three-dimensional medical image into a plurality of small image blocks; Each image patch is mapped to a higher dimensional space through a linear layer; The representation of each image patch will be added with a position code, which can indicate the position of each image patch in the original image; The small image blocks are serialized and input into the image encoder to obtain image features.
[0014] Furthermore, the specific steps of fusing image features and text features in S7 are as follows: In the cross-attention layer, image features are input as queries and text features are input as keys and values; Attention weights are obtained by calculating the dot product between the query and the key, which are used to weight the values and generate the fused feature representation; The fused features are passed through a feed-forward network to generate the final text features.
[0015] The loss function of the new features formed after fusion is defined as follows: The image features in S4 and the new text features obtained in S7 are jointly contrastively learned, and the loss function for calculating the image-text pairing loss is defined as follows: The information noise contrast estimation loss function is a loss function for self-supervised learning, which is particularly suitable for contrastive learning tasks. It learns useful representations by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs. The goal of the information noise contrast estimation loss function is to maximize the similarity between matching images and texts, while minimizing the similarity between unmatched images and texts; this loss function is expressed as follows:
[0016] in, represents the information noise contrast estimation loss, is the feature vector of the image, is with the image The matched text feature vector, is with the image Unmatched text feature vector, sim is the similarity function, is the temperature parameter, N is the total number of vectors in the batch; The text features generated in S3 and the reconstructed text obtained in S5 are jointly calculated. In the text reconstruction task, the commonly used loss function is the log-likelihood loss function, also known as the cross entropy loss function. This loss function is used to measure the difference between the text sequence generated by the model and the real text sequence. The specific text reconstruction loss function is recorded as:
[0017] in, represents the log-likelihood loss, N represents the length of the text sequence, represents the tth word in the real text sequence, Represents the sequence generated by the model given the input x and before Under the condition of probability.
[0018] The image features generated in S4 and the reconstructed image obtained in S6 are jointly calculated. The commonly used loss function in the image mask reconstruction task is the mean square error, which is used to measure the difference between the reconstructed image and the original image. The loss function of the image reconstruction loss is recorded as:
[0019] in, represents the mean square error loss, N represents the total number of pixels in the image, represents the i-th pixel value in the original image, Represents the i-th pixel value in the reconstructed image.
[0020] The weighted sum of the loss functions of the three tasks is defined as follows:
[0021] in, is the final total loss function, represents the weight of the information noise contrast estimation loss, represents the weight of the log-likelihood loss, Represents the weight of the mean squared error loss.
[0022] In a second aspect, the present invention proposes an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the three-dimensional medical image visual language model pre-training method of the reconstruction hybrid strategy when executing the computer program.
[0023] In a third aspect, the present invention proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the three-dimensional medical image visual language model pre-training method of the reconstruction hybrid strategy.
[0024] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention provides a pre-training method for a three-dimensional medical image visual language model with a reconstruction hybrid strategy, fine-tunes a large language model, uses the fine-tuned large language model to extract diagnosis and attribute information in medical reports and generate efficient prompts, and the large language model has a strong generalization ability, which greatly saves the cost of manual annotation. The mask reconstruction strategy of the present invention includes mask reconstruction of visual images and mask reconstruction of language texts, which not only enriches the semantic knowledge of network learning, but also improves the efficiency of pre-training. The semantic perception fusion strategy of the present invention is to fuse the text features and image features obtained by the text encoder to generate new text features, so that the text can perceive the diagnosis and attribute features of the image in advance, and further optimize the alignment of the image and text in the embedding space, so that the text can perceive the image features in advance and align, thereby improving the efficiency of pre-training. The present invention designs two branch networks and three branch tasks when constructing a network. The three branch tasks are: jointly calculating the text reconstruction loss by combining text features and reconstructed text, jointly calculating the image reconstruction loss by combining image features and reconstructed images, and jointly performing comparative learning on image features and fused new text features to calculate the image-text pairing loss; weighted summing of the loss functions of the three branches is performed to achieve multi-task joint learning, and multiple objective functions are jointly trained during the training process. The network can learn more relevant semantic knowledge, enrich the content reserve of the network, facilitate the network to be used for downstream multi-task fine-tuning, and greatly improve the performance of downstream visual or language tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 Schematic diagram of the process of pre-training a 3D medical image visual language model using a hybrid reconstruction strategy in an embodiment of the present invention.
[0026] Figure 2 Schematic diagram of the model architecture in an embodiment of the present invention.
[0027] Figure 3 It is a schematic diagram of identifying a target that needs to be distinguished between directions according to the left and right of a human body in an embodiment of the present invention.
[0028] Figure 4 This is a schematic diagram showing that the naming of the same anatomical target from different data sets is consistent in an embodiment of the present invention.
[0029] Figure 5 A schematic diagram of merging fine-grained classes and generating additional classes as supplements in an embodiment of the present invention.
[0030] Figure 6 It is a schematic diagram showing that the large language model in an embodiment of the present invention can extract corresponding diagnosis and attribute information and generate a text representation corresponding to the template.
[0031] Figure 7Schematic diagram of a text encoder in an embodiment of the present invention.
[0032] Figure 8 Schematic diagram of an image encoder in an embodiment of the present invention.
[0033] Fig. 9 Schematic diagram of the semantic-aware fusion strategy in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0036] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0037] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0038] See also Figure 1 ,The 3D medical image visual language model pre-training method of reconstruction hybrid strategy mainly includes the following steps: S1, constructing a medical image-text pair dataset, wherein the medical image-text pair dataset includes three-dimensional medical images and medical field reports; S2, extract and generate text information: extract the diagnosis and attribute information of the medical field report of the medical image text pair dataset in S1 to generate the text corresponding to the template; S3, text feature generation: perform a random mask operation on the text generated by S2 to obtain a text mask, and input the obtained text mask into the text encoder to generate text features; S4, image feature generation: preprocessing the three-dimensional medical image of the medical image text pair data set in S1, and performing a random mask operation on the preprocessed three-dimensional medical image to obtain a three-dimensional medical mask image, and inputting the obtained three-dimensional medical mask image into an image encoder to generate image features; S5, input the text features obtained in S3 into the text decoder to obtain the reconstructed text; S6, inputting the image features obtained in S4 into the image decoder to obtain a reconstructed image; S7, semantic-aware fusion strategy: Use the cross-attention mechanism to fuse the text features and image features obtained in S2 and S3 to generate new text features; S8, text reconstruction task: jointly calculate the text reconstruction loss by combining the text features generated in S3 and the reconstructed text obtained in S5; S9, image reconstruction task: jointly calculate the image reconstruction loss by combining the image features generated in S4 and the reconstructed image obtained in S6; S10, image-text pairing task: the image features in S4 and the new text features obtained in S7 are combined for comparative learning, and the image-text pairing loss is calculated; S11, multi-task joint learning: weighted sum of the losses of different tasks in S8, S9 and S10.
[0039] Among them, due to the characteristics of 3D medical image datasets being difficult to obtain and having high annotation costs, there is currently no large-scale 3D medical image-text pair dataset. Therefore, the present invention first collects a large number of publicly available 3D medical image datasets and medical field report data, and then pairs the image data with the medical reports to construct a large-scale medical image-text pair dataset. Figures 2 to 5 , the specific implementation is as follows: S11: This paper collects and integrates 45 different publicly available 3D medical image segmentation datasets, totaling 13,918 CT scan data, including 252,243 segmentation annotations across 8 major regions of the human body. The detailed information of the dataset is shown in Table 1; Table 1: Details of 45 medical segmentation datasets
[0040] S12: The present invention collects medical report data from two main medical field knowledge sources, mainly including the concepts, definitions and relationships of segmented organs, such as Figure 2 As shown, the detailed information of the knowledge source is shown in Table 2; S13: The present invention pairs the 3D medical image data with the medical report and performs a variety of procedures to ensure the unity of the semantic information of the organs in the image and the semantic information of the medical report. The present invention first checks each anatomical target in each 3D medical image data set and assigns a medical report to it, which ensures the accuracy and clarity of the medical text between the data sets. For example, the target that needs to be distinguished between the positions, such as the left and right lungs, is always identified according to the left and right sides of the human body. Figure 3 Second, the naming of the same anatomical targets from different datasets is consistent. For example, the i-th lumbar vertebra in the TotalSegmentator dataset and the MRSpineSeg dataset are both named in the format of “lumbar vertebra i”, as shown in Figure 4 Finally, the same anatomical structure may be annotated with different levels in different datasets. In this case, the present invention merges fine-grained classes and generates additional classes as supplements to narrow the gap between datasets. For example, the liver subregions in the CouinaudLiver dataset are merged and added as a new class “liver”, as shown in Figure 5 shown.
[0041] In some implementations, due to the complexity, diversity, and professionalism of medical report data, it is difficult to select appropriate and efficient prompts from medical reports by manual methods. Therefore, it is necessary to use the powerful generalization of large language models to generate high-quality diagnostic prompts based on medical reports and templates. Secondly, considering that medical reports contain not only diagnostic information but also description information of corresponding medical image textures and features, the present invention uses large language models to extract description information of medical image attributes in medical reports and generates usable attribute prompts based on templates. Figure 6 to Figure 7 , the specific implementation of extracting and generating text information and generating text features in S2 and S3 is as follows: Use the open source large language model BioBERT and fine-tune it on a publicly available large-scale medical language dataset; train the model to extract diagnosis and attribute information from medical reports; Inputting medical field reports and corresponding template information into the trained and fine-tuned large language model, the trained and fine-tuned large language model can extract the diagnosis and attribute information of the corresponding medical field reports and generate the diagnosis text and attribute text representation corresponding to the template; The diagnostic text and attribute text corresponding to the generated template are used as prompts for random masking operations, where the masking probability is fixed at 20%, to obtain the text mask, which is then input into the text encoder to generate text features.
[0042] In some implementations, due to the characteristics of sparse data, poor imaging quality, and imbalanced categories, 3D medical image data cannot be used directly for training. It is often necessary to preprocess the data to generate data with rich visual information, improve the speed of model training, and enable the model to have good generalization ability. Figure 8 , the specific implementation is as follows: In some embodiments, the three-dimensional medical image data is preprocessed, including a variety of image transformation operations, to improve the accuracy and efficiency of image analysis. The detailed information of the preprocessing is shown in Table 3.
[0043] The specific steps of preprocessing 3D medical images and generating image features in S4 are: S31: setting a central area based on the image center; S32: Select a point in the central area as the central point for cropping, and finally obtain a three-dimensional medical image block x, where x represents a single image block, R represents the image dimension, W represents the width of the image, H represents the height of the image, and D represents the depth of the image. The resolution of the three-dimensional medical image block used in the pre-training stage of the present invention is 32×256×256, that is, D=32, H=256, and W=256; S33: The present invention performs a random masking operation on the obtained three-dimensional medical image block x, wherein the masking probability is fixed at 20%, and generates visual features through an image encoder for the masked image block. The visual features are input into the image encoder to obtain a reconstructed image. The structure of the image encoder is as follows: Figure 8 As shown, the encoder image patch size is 4×16×16.
[0044] The semantic perception fusion strategy of the present invention combines text features with image features to enhance the transferability of the visual language model, so that the text can perceive the image features in advance and align them. The specific implementation is as follows: exist Figure 8 In the image encoder shown, the acquisition of image features goes through multiple steps: (1) Image segmentation: The input image is segmented into multiple small patches, which are usually rectangular blocks, such as 4×16×16 pixels.
[0045] (2) Linear mapping: each image patch is mapped to a higher-dimensional space through a linear layer (usually a convolutional layer).
[0046] (3) Adding positional encoding: In order to enable the model to understand the spatial information in the image, the representation of each patch will be added with positional encoding, so that the model can know the position of each image patch in the original image.
[0047] (4) Serialization: These encoded patches form a sequence, which will be input into the encoder.
[0048] In some embodiments, the size of the three-dimensional medical image input to the image encoder is 32×256×256. After image segmentation, the three-dimensional medical image is divided into multiple image blocks of size 4×16×16. Then the three-dimensional medical image will be divided into 8×16×16 image blocks. Each image block is mapped to a dimension of 768 through a linear layer. Subsequently, a position code is added to each image block to indicate its position in the original image. Finally, the image blocks are serialized and input into the image encoder. Then, the dimension of the three-dimensional medical image after the image encoder extracts visual features is 2048×768.
[0049] exist Figure 7 In the text encoder shown, the acquisition of text features goes through multiple steps: (1) Word segmentation: segment the obtained diagnostic text and attribute text into words, usually words or subwords; (2) Add special words, such as "classification" and "segmentation", at the beginning and end of the diagnosis text sequence and attribute text sequence respectively. "Classification" is usually used for classification tasks, while "segmentation" is used to separate sentences; (3) Convert words into unique identifiers, and convert the diagnostic text sequence and attribute text sequence after word segmentation into index identifiers corresponding to the vocabulary in the open source large language model BioBERT; (4) Create an attention mask. Create a mask to indicate which parts of the text sequence are actual words and which are special words or fillers. This helps the model ignore the filled parts when processing the sequence. (5) Generating type masks: When processing two-sentence tasks (such as question answering or sentence pair classification), it is necessary to distinguish the text sequences of the two sentences, which is usually achieved by assigning different type identifiers to the text sequences in each sentence. (6) Create position encoding. Since the Transformer architecture itself does not have the ability to capture sequence order, adding position encoding provides the model with information about the position of words in the sentence. In some embodiments, when the text encoder inputs a medical report, it first divides the report into words through word segmentation, then adds special words to segment the sentences, and pads or truncates to ensure that the text length is 2048 (the text encoder requires a text length of 2048, and the text length refers to the number of words. If the text length is less than 2048, it is padded, and if the text length is greater than 2048, it is truncated), and then each word is converted into a unique identifier in the word list, and an attention mask, a type mask, and a position encoding are created. In the text encoder, each word is mapped to a vector space with a dimension of 768 through a linear projection layer, so the dimension of the medical report after the text encoder extracts semantic features is 2048×768.
[0050] In some embodiments, a cross-attention mechanism is used to fuse image features and text features. The output dimension of the known image encoder is 2048×768, and the output dimension of the text encoder is 2048×768. The feature dimensions of image features and text features are the same. In the cross-attention layer, the image features are input as queries, and the text features are input as keys and values. In this way, the model can learn the alignment representation between images and texts. The attention weights are obtained by calculating the dot product between the query and the key, and these weights are used to weight the values to generate the fused feature representation. The fused features are further processed and passed through a feed-forward network to generate the final text feature representation. The final fused feature dimension is 2048×768.
[0051] In some embodiments, the image features in S4 and the new text features obtained in S7 are jointly subjected to contrastive learning, and the loss function for calculating the image-text pairing loss is defined as follows: The information noise contrast estimation loss function is a loss function for self-supervised learning, which is particularly suitable for contrastive learning tasks. It learns useful representations by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs. The goal of the information noise contrast estimation loss function is to maximize the similarity between matching images and texts, while minimizing the similarity between unmatched images and texts; this loss function is expressed as follows:
[0052] in, represents the information noise contrast estimation loss, is the feature vector of the image, is with the image The matched text feature vector, is with the image Unmatched text feature vector, sim is the similarity function, is the temperature parameter, N is the total number of vectors in the batch; The text features generated in S3 and the reconstructed text obtained in S5 are jointly calculated. In the text reconstruction task, the commonly used loss function is the log-likelihood loss function, also known as the cross entropy loss function. This loss function is used to measure the difference between the text sequence generated by the model and the real text sequence. The specific text reconstruction loss function is recorded as:
[0053] in, represents the log-likelihood loss, N represents the length of the text sequence, represents the tth word in the real text sequence, Represents the sequence generated by the model given the input x and before Under the condition of probability.
[0054] The image features generated in S4 and the reconstructed image obtained in S6 are jointly calculated. The commonly used loss function in the image mask reconstruction task is the mean square error, which is used to measure the difference between the reconstructed image and the original image. The loss function of the image reconstruction loss is recorded as:
[0055] in, represents the mean square error loss, N represents the total number of pixels in the image, represents the i-th pixel value in the original image, Represents the i-th pixel value in the reconstructed image.
[0056] The weighted sum of the loss functions of the three tasks is defined as follows:
[0057] in, represents the weighted sum of the losses of the three tasks, represents the weight of the information noise contrast estimation loss, represents the weight of the log-likelihood loss, Represents the weight of the mean squared error loss.
[0058] The provided method is further applied in model training as follows: The network’s performance on medical organ segmentation was tested by fine-tuning the pre-trained visual encoder on three standard medical image segmentation datasets (BTCV, TotalSegmentaor v2, and TotalSegmentaorMRI), using 1%, 10%, and 100% of the training data, respectively.
[0059] As shown in Table 2, under different training data ratios of the three data sets, the method of the present invention significantly outperforms all models based on CNN, Transformer, Mamba, LSTM and SAM in performance.
[0060] Table 2 Experimental results of downstream tasks (organ segmentation of 3D medical images)
[0061] It is noted that in the BTCV dataset (which contains 50 abdominal CT scans from patients with metastatic liver cancer or postoperative abdominal wall hernia. Each scan in the dataset was performed during the portal contrast phase with different volume and field of view parameters. The in-image resolution in the dataset varies from 0.54 x 0.54 mm² to 0.98 x 0.98 mm², and the slice thickness ranges from 2.5 mm to 5.0 mm) the method of the present invention outperforms the strongest opponent by 0.3% using only 1% of the labeled data, by 0.3% using 10% of the labeled data, and by 0.2% when using 100% of the data.
[0062] It is worth noting that on the TotalSegmentaor v2 dataset (this dataset is the largest publicly annotated CT segmentation dataset. The first version of the data was released in July 2022. The official dataset was significantly updated in September 2023, with a slight increase in the number of images and the number of annotated categories. The total number of images was increased from 1204 to 1228 (only the number of test sets was increased), and the number of categories was increased from 104 to 117. The currently public dataset is divided into 1082 training sets, 57 validation sets, and 89 test sets (the number of v1 versions is 65), all of which are publicly annotated), the method of the present invention achieved SOTA results when using 1% and 10% of the labeled data, but did not exceed the strongest method when using 100% of the labeled data, which indicates that the model pre-trained by the method of the present invention has fast convergence.
[0063] Amazingly, on the TotalSegmentaor MRI dataset (the dataset contains 298 MR images and provides segmentation annotations of up to 56 different common anatomical structures. Among them, 251 MR images are from the picture archiving and communication system (PACS) of the University Hospital of Basel between 2011 and 2023, and the other 47 MR images are from the Imaging Data Commons (IDC) to increase the diversity of images. This dataset is derived from random sampling in daily clinical work and represents a real-world dataset that can be generalized to clinical applications. It covers a variety of different lesions, scanners, imaging sequences, and data from different medical institutions. It is worth noting that although the official paper mentions that it contains 59 categories, only 56 categories of annotations are provided in the public dataset, which is slightly different.) The method of the present invention achieves the accuracy of the SOTA method when using 1% of the labeled data, and exceeds all methods when using 10% and 100% labeled data, indicating that the visual model pre-trained by this method has excellent transfer performance.
[0064] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0065] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0066] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0068] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the attached claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any figure mark in the claims should not be regarded as limiting the claims involved.
[0069] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art. The above content is only to illustrate the technical ideas of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical ideas proposed by the present invention shall fall within the protection scope of the claims of the present invention. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A three-dimensional medical image visual language model pre-training method based on a hybrid reconstruction strategy, characterized in that: The following steps are involved: S1, constructing a medical image-text pair dataset, wherein the medical image-text pair dataset includes three-dimensional medical images and medical field reports; S2, extract and generate text information: extract the diagnosis and attribute information of the medical field report of the medical image text pair dataset in S1 to generate the text corresponding to the template; S3, text feature generation: perform a random mask operation on the text generated by S2 to obtain a text mask, and input the obtained text mask into the text encoder to generate text features; S4, image feature generation: preprocessing the three-dimensional medical image of the medical image text pair data set in S1, and performing a random mask operation on the preprocessed three-dimensional medical image to obtain a three-dimensional medical mask image, and inputting the obtained three-dimensional medical mask image into an image encoder to generate image features; S5, input the text features obtained in S3 into the text decoder to obtain the reconstructed text; S6, inputting the image features obtained in S4 into the image decoder to obtain a reconstructed image; S7, semantic-aware fusion strategy: Use the cross-attention mechanism to fuse the text features and image features obtained in S2 and S3 to generate new text features; S8, text reconstruction task: jointly calculate the text reconstruction loss by combining the text features generated in S3 and the reconstructed text obtained in S5; S9, image reconstruction task: jointly calculate the image reconstruction loss by combining the image features generated in S4 and the reconstructed image obtained in S6; S10, image-text pairing task: the image features in S4 and the new text features obtained in S7 are combined for comparative learning, and the image-text pairing loss is calculated; S11, multi-task joint learning: weighted sum of the losses of different tasks in S8, S9 and S10.
2. The method for pre-training a 3D medical image visual language model using a hybrid reconstruction strategy according to claim 1, characterized in that: The specific steps of S1 to build a large-scale medical image-text pair dataset include: S11, collect and integrate different publicly available 3D medical image datasets; S12, collect medical domain reports from major medical domain knowledge sources; S13, inspecting objects in each three-dimensional medical image dataset and assigning a medical domain report thereto.
3. The method for pre-training a 3D medical image visual language model using a hybrid reconstruction strategy according to claim 1, characterized in that: The specific steps of extracting and generating text information and generating text features in S2 and S3 are: Use a large language model and fine-tune it on a publicly available large-scale medical language dataset; Inputting medical field reports and corresponding template information into the trained and fine-tuned large language model, the trained and fine-tuned large language model can extract the diagnosis and attribute information of the corresponding medical field reports and generate the diagnosis text and attribute text representation corresponding to the template; The diagnostic text and attribute text corresponding to the generated template are used as prompts for random masking operations, where the masking probability is fixed at 20%, to obtain the text mask, which is then input into the text encoder to generate text features.
4. The method for pre-training a 3D medical image visual language model using a hybrid reconstruction strategy according to claim 1, characterized in that: The specific steps of preprocessing 3D medical images and generating image features in S4 are: S31, setting a central area based on the center of the three-dimensional medical image; S32, selecting a point in the central area as the central point for cropping to obtain a three-dimensional medical image block x, where x represents a single image block; S33, performing a random masking operation on the obtained three-dimensional medical image block x, wherein the masking probability is fixed at 20%, to obtain a three-dimensional medical masked image, and inputting the obtained three-dimensional medical masked image into an image encoder to generate image features.
5. The method for pre-training a 3D medical image visual language model using a hybrid reconstruction strategy according to claim 2, characterized in that: The specific steps for obtaining text features in S3 are: The obtained diagnosis text and attribute text are segmented into words; Add special words at the beginning and end of the diagnostic text sequence and attribute text sequence respectively; Convert the segmented diagnostic text sequence and attribute text sequence into index identifiers corresponding to the vocabulary in the large language model; Create attention masks to indicate actual words, special words, and fillers in diagnosis text sequences and attribute text sequences; Generate a type mask and assign different type identifiers to the diagnostic text sequence and attribute text sequence in each sentence; Create positional encodings to provide the large language model with information about the position of words in the sentence; In the text encoder, each word is mapped to a vector space through a linear projection layer to extract text features.
6. The method for pre-training a 3D medical image visual language model using a hybrid reconstruction strategy according to claim 4, characterized in that: The specific steps for obtaining image features in S4 are as follows: Segmenting the preprocessed three-dimensional medical image into a plurality of small image blocks; Each image patch is mapped to a higher dimensional space through a linear layer; The representation of each image patch will be added with a position code, which can indicate the position of each image patch in the original image; The small image blocks are serialized and input into the image encoder to obtain image features.
7. The method for pre-training a 3D medical image visual language model using a hybrid reconstruction strategy according to claim 1, characterized in that: The specific steps of fusing image features and text features in S7 are as follows: In the cross-attention layer, image features are input as queries and text features are input as keys and values; Attention weights are obtained by calculating the dot product between the query and the key, which are used to weight the values and generate the fused feature representation; The fused features are used to generate new text features through a feed-forward network.
8. The method for pre-training a 3D medical image visual language model using a hybrid reconstruction strategy according to claim 1, characterized in that: The image features in S4 and the new text features obtained in S7 are jointly compared and learned, and the loss function for calculating the image-text pairing loss is defined as follows: in, represents the information noise contrast estimation loss, is the feature vector of the image, is with the image The matched text feature vector, is with the image Unmatched text feature vector, sim is the similarity function, is the temperature parameter, N is the total number of vectors in the batch; The loss function of jointly calculating the text reconstruction loss of the text features generated in S3 and the reconstructed text obtained in S5 is recorded as: in, represents the log-likelihood loss, N represents the length of the text sequence, represents the tth word in the real text sequence, Represents the sequence generated by the model given the input x and before Under the condition of probability; The loss function of jointly calculating the image reconstruction loss of the image features generated in S4 and the reconstructed image obtained in S6 is recorded as: in, represents the mean square error loss, N represents the total number of pixels in the image, represents the i-th pixel value in the original image, represents the i-th pixel value in the reconstructed image; The weighted sum of the loss functions of the three tasks is defined as follows: in, represents the weight of the information noise contrast estimation loss, represents the weight of the log-likelihood loss, Represents the weight of the mean squared error loss.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Pre-training method and device for medical multi-modal model
CN114972929A
Context awareness medical vision language model pre-training method, system and application
CN118039056A
Training method and device of multi-language multi-mode pre-training model and electronic equipment
CN114970721A
Ultrasonic image pre-training method based on vision-language multi-mode contrast learning
CN118821900A
Cancer auxiliary diagnosis and treatment method based on vision-language large model
CN119170257A
Cited By
Chest radiograph report generation method and system based on cross-modal alignment and significant semantic region
CN120748604A
A chest radiograph report generation method and system based on cross-modal alignment and salient semantic regions
CN120748604B
Image processing system and method and storage medium
CN121034602A
Medical report generation method based on anti-factual reasoning
CN121075616A
A medical report generation method based on counterfactual reasoning
CN121075616B