Medical image analysis method based on prior position information
By using a priori position information-based method in medical imaging analysis, initialize learning Queries, extract and interact features, generate segmented masks and textual descriptions, and input large language models, the problem of poor target position learning in the existing technology is solved, and more accurate medical imaging text report generation is achieved.
Patent Information
- Application Number
- CN202411974152.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In the extraction of 3D medical image information, especially the learning of target locations, it is difficult to effectively supervise, resulting in insufficient accuracy in the generation of medical image text reports.
Through a medical image analysis method based on prior location information, learnable Queries in different regions are initialized, 3D visual features and local features are extracted, feature interactive learning is performed, segmented masks and textual descriptions of the target region are generated, and input them into a large language model LLM to generate text reports.
This method effectively reduces the amount of calculation, improves detection accuracy, enhances the semantic interaction between image features and target positions, and can generate more accurate target area text reports.
Smart Images

Figure CN119941648A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to medical image analysis, and in particular to a medical image analysis method based on priori position information. Background Art
[0002] Medical scenarios contain a large amount of multimodal information, including patient information, diagnostic reports, and medical images of various imaging modalities. The diagnostic reports combined with medical images can provide accurate and detailed descriptions, findings, and diagnoses, and are considered high-quality annotations. These diagnostic reports and medical images are saved in the database on a large scale together with the doctor's diagnostic workflow without any additional cost, and how to make full use of these data to accurately analyze medical images has become a key issue.
[0003] The existing technical solution mainly uses the ViT visual encoder to encode the CT image, and then inputs it into the large language model LLM after simple resampling and dimensionality transformation. The large language model LLM then outputs a text report. The training method usually adopts the self-supervised training mode of predicting the next time.
[0004] The above methods are difficult to learn effectively in extracting 3D medical image information, especially the location of the target. The main reason is that there is no direct and effective supervision between features and text. For the task of generating medical image text reports, the association between location description and image features is crucial to the accuracy of the text report, so the problem of strengthening the information of the target location needs to be solved urgently.
[0005] In addition, most of these methods use a series model of "visual encoder + feature sampler + LLM". The visual encoder is pre-trained in an independent stage, and then connected with the feature sampler to complete the final stage of fine-tuning. In this training mode, based on universal considerations, the pre-training process of the visual encoder does not conduct effective query information learning and guidance for specific areas. Although it can be easily inserted into various LLM frameworks, the extracted features lack the characteristics of emphasizing regional features, and the effect is not good for lesion detection that relies more on regional features. Summary of the invention
[0006] 1. Technical issues to be solved
[0007] In view of the above-mentioned shortcomings of the prior art, the present invention provides a medical image analysis method based on prior position information, which can effectively overcome the defect of the prior art that it is difficult to accurately generate a text report of the target area.
[0008] (II) Technical solution
[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0010] A medical image analysis method based on prior position information comprises the following steps:
[0011] S1. Initialize the learnable Queries for different areas according to the importance of different parts;
[0012] S2, input 3D medical images, extract 3D visual features in the corresponding area according to the prior position information, and extract local features of the specific area from all 3D visual features;
[0013] S3, perform feature interactive learning on the learnable Queries of each region and the local features of each region to obtain the visual interactive features of each region;
[0014] S4, using the segmentation head to perform region segmentation on the input image based on the visual interaction features of the target region to obtain a segmentation mask of the target region;
[0015] S5, using the classification head to perform classification based on the visual interaction features of the target area, and performing text processing on the classification results to obtain a text description of the classification results of the target area;
[0016] S6. Input the visual interaction features of the target area, the textual description of the classification results and the user instruction text into the large language model LLM to obtain a text report of the target area, and output it together with the segmentation mask of the target area.
[0017] Preferably, in S1, learnable Queries for different areas are initialized according to the importance of different parts, including:
[0018] According to the types of lesions in different parts, initialize the learnable Queries in charge of different areas:
[0019] Queries = {Q 左肺 ;Q 右肺 ;Q 纵膈};
[0020] Among them, Q 左肺 For the left lung query, considering that there are many types of lung lesions, the number of assigned queries is 20. 右肺 For the right lung query, considering that there are many types of lung lesions, the number of assigned queries is 20. 纵膈 For mediastinal query, the number of allocated queries is 10 due to the large number of lung lesions.
[0021] Preferably, in S2, a 3D medical image is input, 3D visual features are extracted in a corresponding area according to the prior position information, and local features of a specific area are extracted from all 3D visual features, including:
[0022] S21, input 3D medical images, align and adjust the input images, where the positions of each part are fixed and known, and use the pre-trained SAT model to extract the region mask of the input image to obtain the prior position information proposal;
[0023] S22, using the ViT visual encoder to extract 3D visual features in the corresponding area based on the prior position information proposal;
[0024] S23, using ROIAlign method to extract local features of specific areas from all 3D visual features:
[0025] F 左肺 =ROIAlign[SAT(left lung); F];
[0026] Among them, SAT(left lung) means that the SAT model extracts the region mask of the left lung in the input image, F is the 3D visual feature of all regions in the input image, and F 左肺 is the local feature of the left lung in the input image;
[0027] S24. According to the local features of all specific regions, the local features of the input image are obtained:
[0028] F region ={F 左肺 ; F 右肺 ; F 纵膈};
[0029] Among them, F 左肺 、F 右肺 、F 纵膈 are the local features of the left lung, right lung, and mediastinum in the input image, respectively. region is the local feature of the input image.
[0030] Preferably, in S3, feature interactive learning is performed on the learnable Queries of each region and the local features of each region to obtain visual interactive features of each region, including:
[0031] S31. Perform feature interactive learning on the learnable Queries of each region and the local features of each region to obtain the visual interactive features of each region:
[0032] T 左肺 =CrossAttention(Q 左肺 ; F 左肺 );
[0033] Among them, CrossAttention represents the cross attention mechanism, T 左肺is the visual interaction feature of the left lung in the input image;
[0034] S32. Obtain visual interaction features of the input image based on the visual interaction features of all regions:
[0035] T region ={T 左肺 ; T 右肺 ; T 纵膈};
[0036] Among them, T 左肺 、T 右肺 、T 纵膈 are the visual interaction features of the left lung, right lung, and mediastinum in the input image, respectively. region is the visual interaction feature of the input image.
[0037] Preferably, in S4, the segmentation head is used to perform region segmentation on the input image based on the visual interaction features of the target region to obtain a segmentation mask of the target region, including:
[0038] S41, pre-training the segmentation head using the segmentation task, directly guiding the segmentation head to learn the location and characteristics of the region, and obtaining a pre-trained segmentation head;
[0039] S42, inputting the visual interaction features of the target area into the pre-trained segmentation head, and the segmentation head performs region segmentation on the input image to obtain a segmentation mask of the target area.
[0040] Preferably, in S5, the classification head is used to perform classification based on the visual interaction features of the target area, and the classification result is processed into text to obtain a text description of the classification result of the target area, including:
[0041] S51. Pre-train the classification head using the classification task, extract the classification labels from the text report of the medical image, and then use the visual interaction features of each region in the medical image for classification prediction. The loss function is calculated based on the classification prediction results and the classification labels to guide the classification head to learn, so as to emphasize the regionality of the visual interaction features, so that the specified learnable queries focus more on their own regional semantics.
[0042] S52. Pre-train the classification head using the image-text comparison task, taking the visual interaction features of each region in the medical image and the corresponding text report paragraph as positive sample pairs, and taking different paragraphs in the same text report as negative sample pairs, to form positive and negative sample pairs for image-text comparison, and guide the classification head to perform comparative learning, so as to guide the visual interaction features to extend their own semantics to the region-related descriptions. Under the premise of ensuring the regional semantics, more fine-grained semantic guidance is added to obtain the pre-trained classification head;
[0043] S53, inputting the visual interaction features of the target area into a pre-trained classification head, and the classification head performs classification to obtain a classification result;
[0044] S54, textually process the classification results to obtain a textual description of the classification results of the target area.
[0045] Preferably, in S6, the visual interaction features of the target area, the textual description of the classification result and the user instruction text are input into the large language model LLM to obtain a text report of the target area, and output it together with the segmentation mask of the target area, including:
[0046] S61, pre-training the large language model LLM by using the LLM self-supervision task, so as to obtain a broader semantic guidance than the image-text comparison task through the text guidance provided by the large language model LLM, and obtain a pre-trained large language model LLM;
[0047] S62, inputting the visual interaction features of the target area, the textual description of the classification result, and the user instruction text Prompt into the pre-trained large language model LLM, and the large language model LLM generates a text report of the target area;
[0048] S63, outputting the text report of the target area together with the segmentation mask of the target area;
[0049] Among them, when the LLM self-supervision task is used to pre-train the large language model LLM, the obtained semantic guidance has strong noise, so the loss function of the LLM self-supervision task has the smallest weight in the overall loss function of pre-training.
[0050] Preferably, the overall loss function of the pre-training includes the loss functions of the segmentation task, the classification task, the image-text comparison task and the LLM self-supervision task:
[0051] Loss = αLoss DiceLoss +βLoss BCE +γLoss ContractiveLoss +δLoss LLM ;
[0052] Among them, Loss DiceLoss 、Loss BCE 、Loss ContractiveLoss 、Loss LLM are the loss functions of the segmentation task, classification task, image-text comparison task, and LLM self-supervision task, respectively. α, β, γ, and δ are the weights of the loss functions of the segmentation task, classification task, image-text comparison task, and LLM self-supervision task, respectively. Loss is the overall loss function of pre-training.
[0053] (III) Beneficial effects
[0054] Compared with the prior art, the medical image analysis method based on prior position information provided by the present invention has the following beneficial effects:
[0055] 1) Extracting local features of specific areas based on prior position information can effectively reduce the amount of calculation, improve detection accuracy, and enhance the semantic interaction between image features and target locations;
[0056] 2) Perform feature interactive learning on the learnable queries of each region and the local features of each region, and perform pre-training through segmentation tasks, classification tasks, image-text comparison tasks, and LLM self-supervision tasks. Although these tasks are common downstream tasks, the information they guide is very consistent with the features and purposes designed by the present invention, and the model performance is improved through purposeful downstream task learning;
[0057] 3) Different from the existing training mode, the present invention performs pre-training together with the visual encoder, and relying on this training mode, the present invention has the ability to generate a text report of a specified area. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0059] Figure 1 It is a schematic diagram of the process of the present invention;
[0060] Figure 2 Schematic diagram of the process of generating a segmentation mask and a text report for the left lung in the present invention. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0062] A medical image analysis method based on prior position information, such as Figure 1 As shown, S1, according to the importance of different parts, initialize the learnable Queries in charge of different areas, including:
[0063] According to the types of lesions in different parts, initialize the learnable Queries in charge of different areas:
[0064] Queries = {Q 左肺 ;Q 右肺 ;Q 纵膈};
[0065] Among them, Q 左肺 For the left lung query, considering that there are many types of lung lesions, the number of assigned queries is 20. 右肺 For the right lung query, considering that there are many types of lung lesions, the number of assigned queries is 20. 纵膈 For mediastinal query, the number of allocated queries is 10 due to the large number of lung lesions.
[0066] S2. Input 3D medical images, extract 3D visual features in the corresponding area according to the prior position information, and extract local features of specific areas from all 3D visual features, including:
[0067] S21, input 3D medical images, align and adjust the input images, where the positions of each part are fixed and known, and use the pre-trained SAT model to extract the region mask of the input image to obtain the prior position information proposal;
[0068] S22, using the ViT visual encoder to extract 3D visual features in the corresponding area based on the prior position information proposal;
[0069] S23, using ROIAlign method to extract local features of specific areas from all 3D visual features:
[0070] F 左肺 =ROIAlign[SAT(left lung); F];
[0071] Among them, SAT(left lung) means that the SAT model extracts the region mask of the left lung in the input image, F is the 3D visual feature of all regions in the input image, and F 左肺 is the local feature of the left lung in the input image;
[0072] S24. According to the local features of all specific regions, the local features of the input image are obtained:
[0073] F region ={F 左肺 ; F 右肺 ; F 纵膈};
[0074] Among them, F 左肺 、F右肺 、F 纵膈 are the local features of the left lung, right lung, and mediastinum in the input image, respectively. region is the local feature of the input image.
[0075] In the technical solution of this application, if Figure 1 As shown in the figure, the Target Model is suitable for a large multimodal model of medical images. After pre-training, the model output features can be input into the large language model LLM as visual features, including the Visual Encoder (which can be a visual encoder with spatial semantics such as ViT and Conv) that calculates image features, and the Cross-Attention module that extracts regional features, and are trained as a whole. After accessing the large language model LLM, only a small amount of data is needed for fine-tuning to complete the empowerment.
[0076] S3: Perform feature interactive learning on the learnable Queries of each region and the local features of each region to obtain the visual interactive features of each region, including:
[0077] S31. Perform feature interactive learning on the learnable Queries of each region and the local features of each region to obtain the visual interactive features of each region:
[0078] T 左肺 =CrossAttention(Q 左肺 ; F 左肺 );
[0079] Among them, CrossAttention represents the cross attention mechanism, T 左肺 is the visual interaction feature of the left lung in the input image;
[0080] S32. Obtain visual interaction features of the input image based on the visual interaction features of all regions:
[0081] T region ={T 左肺 ; T 右肺 ; T 纵膈};
[0082] Among them, T 左肺 、T 右肺 、T 纵膈 are the visual interaction features of the left lung, right lung, and mediastinum in the input image, respectively. region is the visual interaction feature of the input image.
[0083] S4. Use the segmentation head to segment the input image based on the visual interaction features of the target area to obtain a segmentation mask of the target area, which specifically includes:
[0084] S41, pre-training the segmentation head using the segmentation task, directly guiding the segmentation head to learn the location and characteristics of the region, and obtaining a pre-trained segmentation head;
[0085] S42, inputting the visual interaction features of the target area into the pre-trained segmentation head, and the segmentation head performs region segmentation on the input image to obtain a segmentation mask of the target area.
[0086] S5. Classify the target area using the classification head based on the visual interaction features, and textually process the classification results to obtain a textual description of the classification results of the target area, including:
[0087] S51. Pre-train the classification head using the classification task, extract the classification labels from the text report of the medical image, and then use the visual interaction features of each region in the medical image for classification prediction. The loss function is calculated based on the classification prediction results and the classification labels to guide the classification head to learn, so as to emphasize the regionality of the visual interaction features, so that the specified learnable queries focus more on their own regional semantics.
[0088] S52. Pre-train the classification head using the image-text comparison task, taking the visual interaction features of each region in the medical image and the corresponding text report paragraph as positive sample pairs, and taking different paragraphs in the same text report as negative sample pairs, to form positive and negative sample pairs for image-text comparison, and guide the classification head to perform comparative learning, so as to guide the visual interaction features to extend their own semantics to the region-related descriptions. Under the premise of ensuring the regional semantics, more fine-grained semantic guidance is added to obtain the pre-trained classification head;
[0089] S53, inputting the visual interaction features of the target area into a pre-trained classification head, and the classification head performs classification to obtain a classification result;
[0090] S54, textually process the classification results to obtain a textual description of the classification results of the target area.
[0091] S6. Input the visual interaction features of the target area, the textual description of the classification results and the user instruction text into the large language model LLM to obtain a text report of the target area, and output it together with the segmentation mask of the target area, including:
[0092] S61, pre-training the large language model LLM by using the LLM self-supervision task, so as to obtain a broader semantic guidance than the image-text comparison task through the text guidance provided by the large language model LLM, and obtain a pre-trained large language model LLM;
[0093] S62, inputting the visual interaction features of the target area, the textual description of the classification result, and the user instruction text Prompt into the pre-trained large language model LLM, and the large language model LLM generates a text report of the target area;
[0094] S63, outputting the text report of the target area together with the segmentation mask of the target area;
[0095] Among them, when the LLM self-supervision task is used to pre-train the large language model LLM, the obtained semantic guidance has strong noise, so the loss function of the LLM self-supervision task has the smallest weight in the overall loss function of pre-training.
[0096] In the technical solution of this application, the overall loss function of pre-training includes the loss functions of the segmentation task, classification task, image-text comparison task and LLM self-supervision task:
[0097] Loss = αLoss DiceLoss +βLoss BCE +γLoss ContractiveLoss +δLoss LLM ;
[0098] Among them, Loss DiceLoss 、Loss BCE 、Loss ContractiveLoss 、Loss LLM are the loss functions of the segmentation task, classification task, image-text comparison task, and LLM self-supervision task, respectively. α, β, γ, and δ are the weights of the loss functions of the segmentation task, classification task, image-text comparison task, and LLM self-supervision task, respectively. Loss is the overall loss function of pre-training.
[0099] In the technical solution of this application, for each case in each task, there are multiple lesion areas and corresponding multiple text reports. During pre-training, not all pairs are used, and pairs are randomly discarded. On the one hand, the information relevance between the image and text pairs can be strengthened, and on the other hand, it can learn to output text reports under different regional combinations according to different input queries (queries in each region).
[0100] In order to better illustrate the technical solution of this application, the following is an example of querying left lung disease. Figure 2 As shown:
[0101] To query left lung diseases, you only need to input the left lung Query for feature query, concatenate the corresponding left lung visual interaction features with the text description of the left lung classification results and the user command text Prompt (Are there any lung diseases?), and then input them into the large language model LLM. Finally, a text report on the left lung and the accompanying lung segmentation mask are generated and output.
[0102] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A medical image analysis method based on prior position information, characterized in that: The following steps are involved: S1. Initialize the learnable Queries for different areas according to the importance of different parts; S2, input 3D medical images, extract 3D visual features in the corresponding area according to the prior position information, and extract local features of the specific area from all 3D visual features; S3, perform feature interactive learning on the learnable Queries of each region and the local features of each region to obtain the visual interactive features of each region; S4, using the segmentation head to perform region segmentation on the input image based on the visual interaction features of the target region to obtain a segmentation mask of the target region; S5, using the classification head to perform classification based on the visual interaction features of the target area, and performing text processing on the classification results to obtain a text description of the classification results of the target area; S6. Input the visual interaction features of the target area, the textual description of the classification results and the user instruction text into the large language model LLM to obtain a text report of the target area, and output it together with the segmentation mask of the target area.
2. The medical image analysis method based on prior position information according to claim 1, characterized in that: In S1, learnable queries for different areas are initialized according to the importance of different parts, including: According to the types of lesions in different parts, initialize the learnable Queries in charge of different areas: Queries={Q 左肺 ;Q 右肺 ;Q 纵膈 }; Among them, Q 左肺 For the left lung query, considering that there are many types of lung lesions, the number of assigned queries is 20. 右肺 For the right lung query, considering that there are many types of lung lesions, the number of assigned queries is 20. 纵膈 For mediastinal query, the number of allocated queries is 10 due to the large number of lung lesions.
3. The medical image analysis method based on prior position information according to claim 2, characterized in that: S2 inputs 3D medical images, extracts 3D visual features in the corresponding area based on the prior position information, and extracts local features of specific areas from all 3D visual features, including: S21, input 3D medical images, align and adjust the input images, where the positions of each part are fixed and known, and use the pre-trained SAT model to extract the region mask of the input image to obtain the prior position information proposal; S22, using the ViT visual encoder to extract 3D visual features in the corresponding area based on the prior position information proposal; S23, using ROIAlign method to extract local features of specific areas from all 3D visual features: F 左肺 =ROIAlign[SAT(left lung); F]; Among them, SAT(left lung) means that the SAT model extracts the region mask of the left lung in the input image, F is the 3D visual feature of all regions in the input image, and F 左肺 is the local feature of the left lung in the input image; S24. According to the local features of all specific regions, the local features of the input image are obtained: F region ={F 左肺 ;F 右肺 ;F 纵膈 }; Among them, F 左肺 、F 右肺 、F 纵膈 are the local features of the left lung, right lung, and mediastinum in the input image, respectively. region is the local feature of the input image.
4. The medical image analysis method based on prior position information according to claim 3, characterized in that: In S3, the learnable queries of each region and the local features of each region are interactively learned to obtain the visual interaction features of each region, including: S31. Perform feature interactive learning on the learnable Queries of each region and the local features of each region to obtain the visual interactive features of each region: T 左肺 =CrossAttention(Q 左肺 ;F 左肺 ); Among them, CrossAttention represents the cross attention mechanism, T 左肺 is the visual interaction feature of the left lung in the input image; S32. Obtain visual interaction features of the input image based on the visual interaction features of all regions: T region ={T 左肺 ;T 右肺 ;T 纵膈 }; Among them, T 左肺 , T 右肺 , T 纵膈 are the visual interaction features of the left lung, right lung, and mediastinum in the input image, respectively. region is the visual interaction feature of the input image.
5. The medical image analysis method based on prior position information according to claim 4, characterized in that: In S4, the segmentation head is used to segment the input image based on the visual interaction features of the target area to obtain the segmentation mask of the target area, including: S41, pre-training the segmentation head using the segmentation task, directly guiding the segmentation head to learn the location and characteristics of the region, and obtaining a pre-trained segmentation head; S42, inputting the visual interaction features of the target area into the pre-trained segmentation head, and the segmentation head performs region segmentation on the input image to obtain a segmentation mask of the target area.
6. The medical image analysis method based on prior position information according to claim 5, characterized in that: In S5, the classification head is used to perform classification based on the visual interaction features of the target area, and the classification results are processed into text to obtain a text description of the classification results of the target area, including: S51. Pre-train the classification head using the classification task, extract the classification labels from the text report of the medical image, and then use the visual interaction features of each region in the medical image for classification prediction. The loss function is calculated based on the classification prediction results and the classification labels to guide the classification head to learn, so as to emphasize the regionality of the visual interaction features, so that the specified learnable queries focus more on their own regional semantics. S52. Pre-train the classification head using the image-text comparison task, taking the visual interaction features of each region in the medical image and the corresponding text report paragraph as positive sample pairs, and taking different paragraphs in the same text report as negative sample pairs, to form positive and negative sample pairs for image-text comparison, and guide the classification head to perform comparative learning, so as to guide the visual interaction features to extend their own semantics to the region-related descriptions. Under the premise of ensuring the regional semantics, more fine-grained semantic guidance is added to obtain the pre-trained classification head; S53, inputting the visual interaction features of the target area into a pre-trained classification head, and the classification head performs classification to obtain a classification result; S54, textually process the classification results to obtain a textual description of the classification results of the target area.
7. The medical image analysis method based on prior position information according to claim 6, characterized in that: In S6, the visual interaction features of the target area, the textual description of the classification results, and the user instruction text are input into the large language model LLM to obtain a text report of the target area, which is output together with the segmentation mask of the target area, including: S61, pre-training the large language model LLM by using the LLM self-supervision task, so as to obtain a broader semantic guidance than the image-text comparison task through the text guidance provided by the large language model LLM, and obtain a pre-trained large language model LLM; S62, inputting the visual interaction features of the target area, the textual description of the classification result, and the user instruction text Prompt into the pre-trained large language model LLM, and the large language model LLM generates a text report of the target area; S63, outputting the text report of the target area together with the segmentation mask of the target area; Among them, when the LLM self-supervision task is used to pre-train the large language model LLM, the obtained semantic guidance has strong noise, so the loss function of the LLM self-supervision task has the smallest weight in the overall loss function of pre-training.
8. The medical image analysis method based on prior position information according to claim 7, characterized in that: The overall loss function of the pre-training includes the loss functions of the segmentation task, classification task, image-text comparison task, and LLM self-supervision task: Loss=αLoss DiceLoss +βLoss BCE +γLoss ContractiveLoss +δLoss LLM ; Among them, Loss DiceLoss 、Loss BCE 、Loss ContractiveLoss 、Loss LLM are the loss functions of the segmentation task, classification task, image-text comparison task, and LLM self-supervision task, respectively. α, β, γ, and δ are the weights of the loss functions of the segmentation task, classification task, image-text comparison task, and LLM self-supervision task, respectively. Loss is the overall loss function of pre-training.
Citation Information
Patent Citations
Target detection method and device, equipment and storage medium
CN115953665A
Single-stage three-dimensional scene point cloud target positioning and description method
CN116664983A
Small sample semantic segmentation method based on cross attention self-guidance
CN117611808A
Model training and application method and device for target detection and storage medium
CN118799608A
Brain tumor diagnosis and report generation method and system based on large language model
CN119049638A