Medical image analysis method based on fine-grained features

By dividing 3D medical images into multi-layer 2D images and using hierarchical feature samplers and fine-grained feature samplers, the problem of difficulty in extracting fine-grained feature and global position information in the prior art is solved, and more accurate and efficient lesion recognition and analysis are achieved.

CN119942185APending Publication Date: 2025-05-06安徽影联云享医疗科技有限公司 +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411974150.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract fine-grained features and global location information in 3D medical imaging analysis, resulting in weak lesions recognition capabilities, affecting the accuracy and efficiency of the analysis.

Method used

By imitating the doctor's reading mode, 3D medical images are divided into multi-layer 2D images, and hierarchical features are extracted using ViT encoding and hierarchical feature sampler. Combined with fine-grained feature sampler, multi-output classification head and large language model LLM are constructed for analysis.

Benefits of technology

It improves the detection ability of fine-grained lesions, enhances the model's understanding and analysis ability of 3D medical images, and improves the accuracy and efficiency of analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942185A_ABST
    Figure CN119942185A_ABST
Patent Text Reader

Abstract

The invention relates to medical image analysis, in particular to a medical image analysis method based on fine-grained features, and the method comprises the steps: simulating a reading mode when a doctor actually watches a medical image, dividing a 3D medical image into multiple layers of 2D medical images, and inputting the 2D medical images; performing ViT coding on each layer of 2D medical image to obtain a ViT coding result of each layer of 2D medical image; performing up-sampling and down-sampling on the ViT coding result of each layer of 2D medical image to obtain corresponding up-sampling features and down-sampling features; according to the down-sampling features of each layer of 2D medical image, extracting hierarchical features of each layer of 2D medical image by using a hierarchical feature sampler; according to the hierarchical features of each layer of 2D medical image, using a hierarchical feature aggregator to extract hierarchical aggregation features of each layer of 2D medical image; according to the technical scheme provided by the invention, the defect that the 3D medical image is difficult to accurately and efficiently analyze in the prior art can be effectively overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to medical image analysis, and in particular to a medical image analysis method based on fine-grained features. Background Art

[0002] The existing technical solution mainly uses a ViT-based visual encoder to encode the CT image, and then inputs it into the large language model LLM after simple resampling and dimension transformation, and then the large language model LLM outputs a report. The training paradigm adopted is usually a self-supervised training mode that predicts the next time. The input is in the form of text prompt + visual feature tokens, and the powerful information processing capabilities of the large language model LLM are used to complete feature interaction and generate a report.

[0003] In the current field of 3D medical image analysis, the aforementioned large language model LLM faces some challenges in extracting image information, especially in terms of effective extraction of details and capture of global position information. Existing models often lack efficient processing methods. This limitation results in the model's weak ability to identify lesions and the inability to fully utilize the interaction between fine-grained features and global features in the image, thus affecting the accuracy and efficiency of medical image analysis. Summary of the invention

[0004] 1. Technical issues to be resolved

[0005] In view of the above-mentioned shortcomings of the prior art, the present invention provides a medical image analysis method based on fine-grained features, which can effectively overcome the defect of the prior art that it is difficult to accurately and efficiently analyze 3D medical images.

[0006] (II) Technical solution

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0008] A medical image analysis method based on fine-grained features comprises the following steps:

[0009] S1, imitating the reading mode of doctors when actually viewing medical images, dividing the 3D medical images into multiple layers of 2D medical images for input;

[0010] S2, performing ViT encoding on each layer of 2D medical images to obtain a ViT encoding result of each layer of 2D medical images;

[0011] S3, upsampling and downsampling the ViT encoding results of each layer of 2D medical images respectively to obtain corresponding upsampling features and downsampling features;

[0012] S4, extracting hierarchical features of each layer of 2D medical images using a hierarchical feature sampler according to the down-sampled features of each layer of 2D medical images;

[0013] S5, extracting hierarchical aggregation features of each layer of 2D medical images using a hierarchical feature aggregator according to the hierarchical features of each layer of 2D medical images;

[0014] S6, extracting fine-grained features of each layer of 2D medical images using a fine-grained feature sampler according to the upsampled features and hierarchical aggregation features of each layer of 2D medical images;

[0015] S7. According to the hierarchical aggregation features of each layer of 2D medical images, a multi-output classification head is constructed, and the corresponding disease classification label is obtained according to the multi-output results of the classification head;

[0016] S8. Convert the multi-output results of the classification head into corresponding disease classification descriptions, and concatenate them with the text description and fine-grained features of each layer of 2D medical images and input them into the large language model LLM to obtain the disease text output.

[0017] Preferably, S1 imitates the reading mode of doctors actually viewing medical images, and divides the 3D medical images into multiple layers of 2D medical images for input, including:

[0018] Input 3D medical images of size (B, Z, 3, H, W). Each 3D medical image in each batch contains Z layers of 3×H×W 2D medical images. These 2D medical images are input to simulate the reading mode of doctors when viewing medical images.

[0019] Among them, B is the batch, Z is the number of layers of 2D medical images contained in each 3D medical image, and H and W are the height and width of each layer of 2D medical images, respectively.

[0020] Preferably, performing ViT encoding on each layer of 2D medical images in S2 to obtain a ViT encoding result of each layer of 2D medical images includes:

[0021] ViT encoding of each layer of 2D medical images:

[0022] X z =ViT(S z );

[0023] Among them, S z is the z-th layer 2D medical image, ViT represents ViT coding, X z is the ViT encoding result of the z-th layer 2D medical image, z = 1, 2, ..., Z;

[0024] For each layer of 2D medical images, after ViT encoding, a ViT encoding result containing L visual tokens is obtained, and each visual token contains C feature dimensions; for each 3D medical image, the size of the ViT encoding result obtained after ViT encoding is Z×L×C.

[0025] Preferably, in S3, the ViT encoding result of each layer of 2D medical image is upsampled and downsampled respectively to obtain corresponding upsampled features and downsampled features, including:

[0026] S31. Upsample the ViT encoding result of each layer of 2D medical image to obtain the corresponding upsampled features:

[0027] F z up =Deconv2D(X z );

[0028] Among them, Deconv2D represents upsampling using two-dimensional deconvolution. In two-dimensional deconvolution, the step size is 2, the convolution kernel size is 3, and F z up is the upsampled feature of the z-th layer 2D medical image;

[0029] S32, downsample the ViT encoding result of each layer of 2D medical image to obtain the corresponding downsampled features:

[0030] F z down =Conv2D(X z );

[0031] Among them, Conv2D represents downsampling using two-dimensional convolution. In the two-dimensional convolution, the step size is 2, the convolution kernel size is 3, and F z down is the down-sampled feature of the z-th layer 2D medical image.

[0032] Preferably, in S4, according to the down-sampled features of each layer of the 2D medical image, the hierarchical features of each layer of the 2D medical image are extracted using a hierarchical feature sampler, including:

[0033] S41, downsample the feature F of the z-th layer 2D medical image z down As the sum of native information T of the visual features of this layer z ori , and initialize multiple learnable tokens for this layer as T z learn , to learn layer information to solve the problem that the same encoder lacks understanding of visual features when processing different layers of information;

[0034] S42. Extract hierarchical features of each layer of 2D medical images using a hierarchical feature sampler:

[0035] T z layer =CrossAttn(Concat(T z ori ; T z learn );F z down );

[0036] Among them, Concat represents concatenation in the length dimension, CrossAttn represents the cross attention mechanism, and T z layer It is the hierarchical feature of the z-th layer 2D medical image, containing T+1 global visual information tokens.

[0037] Preferably, in S5, according to the hierarchical features of each layer of 2D medical images, a hierarchical feature aggregator is used to extract hierarchical aggregated features of each layer of 2D medical images, including:

[0038] The hierarchical feature aggregator is used to extract hierarchical aggregation features of each layer of 2D medical images:

[0039] T z layer_fusion =Conv(T z layer );

[0040] Among them, Conv represents feature aggregation using convolution. If the 2D information of the layer is not lost, 3D convolution is used for feature aggregation. In 3D convolution, the step size is 1×1×1 and the window size is 3×3×3. If the 2D information of the layer is lost, the positions of the global visual information tokens in the hierarchical features between different layers still correspond one to one, so 2D convolution is used for feature aggregation. In order to ensure that the dimension of the tokens in the hierarchical aggregation features remains unchanged, 2D convolution / 3D convolution with padding is used. T z layer_fusion is the hierarchical aggregation feature of the z-th layer 2D medical image.

[0041] Preferably, in S6, the fine-grained features of each layer of 2D medical images are extracted using a fine-grained feature sampler according to the up-sampled features and hierarchical aggregation features of each layer of 2D medical images, including:

[0042] S61, upsample the feature F of the z-th layer 2D medical image z up As keys and values, the hierarchical aggregation features T of the z-th layer 2D medical image are zlayer_fusion As queries, perform cross-attention interactions;

[0043] S62. Extract fine-grained features of each layer of 2D medical images using a fine-grained feature sampler:

[0044] T z fine =CrossAttn(F z up ; T z layer_fusion );

[0045] Among them, T z fine is the fine-grained feature of the z-th layer 2D medical image.

[0046] Preferably, in S7, a multi-output classification head is constructed according to the hierarchical aggregation features of each layer of 2D medical images, and corresponding disease classification labels are obtained according to the multi-output results of the classification head, including:

[0047] S71, due to the hierarchical aggregation feature T of the z-th layer 2D medical image z layer_fusion The features of this layer and its adjacent layers are aggregated, and the high-level semantic features of the original image are extracted. Therefore, a multi-output classification head is constructed through pooling and MLP layer mapping:

[0048]

[0049] Among them, Pool represents pooling, MLP represents MLP layer mapping, Y is the multi-output result of the classification head, and each output result is the disease classification feature of a specified type of disease;

[0050] S72, obtaining corresponding disease classification labels according to the multiple output results of the classification head;

[0051] Among them, the loss function of classification head training is BCE Loss.

[0052] Preferably, in S8, the multiple output results of the classification head are converted into corresponding disease classification descriptions, and are concatenated with the text descriptions and fine-grained features of each layer of 2D medical images and then input into the large language model LLM to obtain disease text output, including:

[0053] S81. Convert the multiple output results of the classification head into corresponding disease classification descriptions, and concatenate them with the text description and fine-grained features of each layer of 2D medical images to obtain the input of the large language model LLM:

[0054]

[0055] Among them, T Prompt is the text description of each layer of 2D medical image, T Class_Prompt It converts the multi-output results of the classification head into corresponding disease classification descriptions. LLM_Input is the input of the large language model LLM.

[0056] S82, inputting the concatenation result into the large language model LLM to obtain the disease text output;

[0057] Among them, the loss function of the large language model LLM training is LLMLoss.

[0058] Preferably, the overall loss function of the classification head and the large language model LLM training is:

[0059] L=α·L BCE +L LLM ;

[0060] Among them, L BCE is the loss function BCE Loss of the classification head training, α is the weight coefficient of the loss function of the classification head training, L LLM LLMLoss is the loss function for training the large language model LLM, and L is the overall loss function.

[0061] (III) Beneficial effects

[0062] Compared with the prior art, the medical image analysis method based on fine-grained features provided by the present invention has the following beneficial effects:

[0063] 1) Module design of hierarchical feature sampler: The core of this module is to build a feature pyramid. These features are the sum of the original information of visual features T z ori , combined with learnable hierarchical knowledge T z learn ,Using the cross-attention mechanism for information interaction, this design aims to achieve two goals: one is to efficiently extract global information, and the other is to achieve rapid dimension reduction of features. Through the cross-attention mechanism, information at different levels can complement each other and improve the expressiveness of features;

[0064] 2) Module design of hierarchical feature aggregator: This module uses convolution or other window-based algorithms to aggregate multi-level global information. Since there is a dependency relationship between the inter-layer information of medical images, the design of the hierarchical feature aggregator fully considers this and realizes cross-layer information interaction. In this way, the model can better understand the structure and relationship of different levels in medical images, thereby improving the ability to identify large-area lesions;

[0065] 3) Module design of fine-grained feature sampler: Based on hierarchical feature aggregation, the aggregated hierarchical features are used to guide the sampling of fine-grained features. This method can improve the model's ability to detect fine-grained lesions, especially when identifying small or early lesions. Through refined feature sampling, the model can more accurately locate lesions and provide richer information for subsequent recognition and classification;

[0066] 4) Classification feature supervision and information enhancement based on classification results: Aggregated hierarchical features are used for classification prediction, and the binary cross entropy (BCE) loss function is used for loss calculation to allow the model to fully learn disease knowledge. The classification results are not only used to guide model learning, but also used as prompts to input into the large language model (LLM). Since the hierarchical feature sampler and hierarchical feature aggregator provide high-level semantic information, the classification supervision and prediction information can more effectively enhance the model's disease judgment ability. In addition, the loss guidance of the large language model (LLM) enables the model to make greater use of existing report data and learn a variety of report output styles, thereby improving the generalization ability and practicality of the model.

[0067] In summary, the present invention fully considers the characteristics of 3D medical images, and through efficient feature extraction, aggregation and sampling, as well as supervision and information enhancement based on classification results, a powerful 3D medical image analysis framework is formed, which can identify various types of lesions more accurately and efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0069] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0071] A medical image analysis method based on fine-grained features, such as Figure 1As shown, S1, imitating the reading mode of doctors when actually viewing medical images, dividing the 3D medical image into multiple layers of 2D medical images for input, specifically including:

[0072] Input 3D medical images of size (B, Z, 3, H, W). Each 3D medical image in each batch contains Z layers of 3×H×W 2D medical images. These 2D medical images are input to simulate the reading mode of doctors when viewing medical images.

[0073] Among them, B is the batch, Z is the number of layers of 2D medical images contained in each 3D medical image, and H and W are the height and width of each layer of 2D medical images, respectively.

[0074] S2. Perform ViT encoding on each layer of 2D medical images to obtain the ViT encoding result of each layer of 2D medical images, specifically including:

[0075] ViT encoding of each layer of 2D medical images:

[0076] X z =ViT(S z );

[0077] Among them, S z is the z-th layer 2D medical image, ViT represents ViT coding, X z is the ViT encoding result of the z-th layer 2D medical image, z = 1, 2, ..., Z;

[0078] For each layer of 2D medical images, after ViT encoding, a ViT encoding result containing L visual tokens is obtained, and each visual token contains C feature dimensions; for each 3D medical image, the size of the ViT encoding result obtained after ViT encoding is Z×L×C.

[0079] S3. Upsample and downsample the ViT encoding results of each layer of 2D medical images to obtain corresponding upsampled features and downsampled features, including:

[0080] S31. Upsample the ViT encoding result of each layer of 2D medical image to obtain the corresponding upsampled features:

[0081] F z up =Deconv2D(X z );

[0082] Among them, Deconv2D represents upsampling using two-dimensional deconvolution. In two-dimensional deconvolution, the step size is 2, the convolution kernel size is 3, and F z up is the upsampled feature of the z-th layer 2D medical image;

[0083] S32, downsample the ViT encoding result of each layer of 2D medical image to obtain the corresponding downsampled features:

[0084] F z down =Conv2D(X z );

[0085] Among them, Conv2D represents downsampling using two-dimensional convolution. In the two-dimensional convolution, the step size is 2, the convolution kernel size is 3, and F z down is the down-sampled feature of the z-th layer 2D medical image.

[0086] S4. According to the down-sampled features of each layer of 2D medical images, the hierarchical features of each layer of 2D medical images are extracted using a hierarchical feature sampler, specifically including:

[0087] S41, downsample the feature F of the z-th layer 2D medical image z down As the sum of native information T of the visual features of this layer z ori (You can use [CLS], [REG] or global-pooling of all features, or directly use image features as the sum of native information T of the visual features of this layer z ori , the technical solution of this application directly uses the down-sampling feature F of the z-th layer 2D medical image z down , retaining spatial information for the subsequent level feature aggregator), and initialize multiple learnable tokens for this layer as T z learn , to learn layer information to solve the problem that the same encoder will lack understanding of visual features when processing different layers of information (no prior information belonging to the layer position will lead to lack of visual ability);

[0088] S42. Extract hierarchical features of each layer of 2D medical images using a hierarchical feature sampler:

[0089] T z layer =CrossAttn(Concat(T z ori ; T z learn );F z down );

[0090] Among them, Concat represents concatenation in the length dimension, CrossAttn represents the cross attention mechanism, and Tz layer It is the hierarchical feature of the z-th layer 2D medical image, containing T+1 global visual information tokens.

[0091] S5. According to the hierarchical features of each layer of 2D medical images, a hierarchical feature aggregator is used to extract hierarchical aggregation features of each layer of 2D medical images, specifically including:

[0092] The hierarchical feature aggregator is used to extract hierarchical aggregation features of each layer of 2D medical images:

[0093] T z layer_fusion =Conv(T z layer );

[0094] Among them, Conv represents feature aggregation using convolution. If the 2D information of the layer is not lost, 3D convolution is used for feature aggregation. In 3D convolution, the step size is 1×1×1 and the window size is 3×3×3. If the 2D information of the layer is lost, the positions of the global visual information tokens in the hierarchical features between different layers still correspond one to one, so 2D convolution is used for feature aggregation. In order to ensure that the dimension of the tokens in the hierarchical aggregation features remains unchanged, 2D convolution / 3D convolution with padding is used. T z layer_fusion is the hierarchical aggregation feature of the z-th layer 2D medical image.

[0095] In the technical solution of the present application, after the hierarchical features of each layer of 2D medical images are extracted by using a hierarchical feature sampler, the global visual information tokens at this time only have the information of the layer to which they belong, and it is necessary to perform cross-layer information interaction on these global visual information tokens. The input medical image represents 3D information to a certain extent (for medical images, the acquisition method is slice acquisition, which cannot be completely equivalent to a 3D model, and the inter-layer information is not obtained at all; if the input is a 3D medical image, the technical solution of the present application can also be applied), and cross-layer information interaction has a great benefit on the robustness of visual features.

[0096] S6. According to the upsampled features and hierarchical aggregation features of each layer of 2D medical images, a fine-grained feature sampler is used to extract fine-grained features of each layer of 2D medical images, specifically including:

[0097] S61, upsample the feature F of the z-th layer 2D medical image z up As keys and values, the hierarchical aggregation features T of the z-th layer 2D medical image are z layer_fusionAs queries, perform cross-attention interactions;

[0098] S62. Extract fine-grained features of each layer of 2D medical images using a fine-grained feature sampler:

[0099] T z fine =CrossAttn(F z up ; T z layer_fusion );

[0100] Among them, T z fine is the fine-grained feature of the z-th layer 2D medical image.

[0101] In the technical solution of the present application, the fine-grained features of all layers of 2D medical images can be post-fused by addition / concatenation and the like (in the input expression of the large language model LLM below, the post-fusion is performed by addition), and then input into the large language model LLM as implicit visual features.

[0102] In order to solve the problem in the prior art that existing models often lack efficient processing methods in terms of effective extraction of details and capture of global position information, the technical solution of the present application constructs a feature pyramid in the hierarchical feature sampler to extract the hierarchical features of each layer of 2D medical images. The feature pyramid can effectively capture image features at different scales, thereby improving the ability to recognize details. However, the introduction of the feature pyramid inevitably causes a sharp increase in the number of visual features, which undoubtedly increases the difficulty of other modules, especially large models, in capturing effective visual features, and also brings a significant increase in computational complexity.

[0103] In order to overcome the above difficulties, the technical solution of this application constructs an importance sampling module based on high-dimensional feature aggregation based on hierarchical feature samplers, hierarchical feature aggregators and fine-grained feature samplers. The design concept of this module is to achieve rapid dimensionality reduction of visual features while ensuring the effective extraction of important visual features. In this way, the amount of calculation can be effectively reduced while retaining key information in the image. In addition, the module can use these effective global features to guide the extraction of fine-grained features, thereby improving the detection rate of lesions.

[0104] S7. Based on the hierarchical aggregation features of each layer of 2D medical images, a multi-output classification head is constructed, and the corresponding disease classification labels are obtained according to the multi-output results of the classification head, including:

[0105] S71, due to the hierarchical aggregation feature T of the z-th layer 2D medical image zlayer_fusion The features of this layer and its adjacent layers are aggregated, and the high-level semantic features of the original image are extracted. Therefore, a multi-output classification head is constructed through pooling and MLP layer mapping:

[0106]

[0107] Among them, Pool represents pooling, MLP represents MLP layer mapping, Y is the multi-output result of the classification head, and each output result is the disease classification feature of a specified type of disease;

[0108] S72, obtaining corresponding disease classification labels according to the multiple output results of the classification head;

[0109] Among them, the loss function of classification head training is BCE Loss.

[0110] In the technical solution of the present application, the corresponding disease classification labels are obtained according to the multi-output results of the classification head. For example, the multi-output results of the classification head are: pleural effusion and lung nodules are positive, and enlarged cardiac shadow is negative; the multi-output results of the classification head are transcribed into corresponding disease classification labels through templated phrase construction: pleural effusion-positive, lung nodules-positive, enlarged cardiac shadow-negative.

[0111] S8. Convert the multiple output results of the classification head into corresponding disease classification descriptions, and concatenate them with the text descriptions and fine-grained features of each layer of 2D medical images and input them into the large language model LLM to obtain disease text output, including:

[0112] S81. Convert the multiple output results of the classification head into corresponding disease classification descriptions, and concatenate them with the text description and fine-grained features of each layer of 2D medical images to obtain the input of the large language model LLM:

[0113]

[0114] Among them, T Prompt is the text description of each layer of 2D medical image, T Class_Prompt It converts the multi-output results of the classification head into corresponding disease classification descriptions. LLM_Input is the input of the large language model LLM.

[0115] S82, inputting the concatenation result into the large language model LLM to obtain the disease text output;

[0116] Among them, the loss function of the large language model LLM training is LLMLoss.

[0117] In the technical solution of this application, the overall loss function of the classification head and the large language model LLM training is:

[0118] L=α·LBCE +L LLM ;

[0119] Among them, L BCE is the loss function BCE Loss of the classification head training, α is the weight coefficient of the loss function of the classification head training, L LLM LLMLoss is the loss function for training the large language model LLM, and L is the overall loss function.

[0120] In the technical solution of this application, both implicit and explicit methods are used to improve the detection capability of fine-grained lesions. The implicit method automatically identifies the features that contribute most to lesion identification by learning the importance of features in a high-dimensional feature space; while the explicit method directly extracts and strengthens the fine-grained features related to lesions by designing specific algorithms. The combination of these two methods not only improves the performance of the model, but also ensures that the model can more accurately and efficiently identify potential lesions when processing complex 3D medical images.

[0121] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A medical image analysis method based on fine-grained features, characterized in that: The following steps are involved: S1, imitating the reading mode of doctors when viewing medical images, dividing the 3D medical images into multiple layers of 2D medical images for input; S2, performing ViT encoding on each layer of 2D medical images to obtain a ViT encoding result of each layer of 2D medical images; S3, upsampling and downsampling the ViT encoding results of each layer of 2D medical images respectively to obtain corresponding upsampling features and downsampling features; S4, extracting hierarchical features of each layer of 2D medical images using a hierarchical feature sampler according to the down-sampled features of each layer of 2D medical images; S5, extracting hierarchical aggregation features of each layer of 2D medical images using a hierarchical feature aggregator according to the hierarchical features of each layer of 2D medical images; S6, extracting fine-grained features of each layer of 2D medical images using a fine-grained feature sampler according to the upsampled features and hierarchical aggregation features of each layer of 2D medical images; S7. According to the hierarchical aggregation features of each layer of 2D medical images, a multi-output classification head is constructed, and the corresponding disease classification label is obtained according to the multi-output results of the classification head; S8. Convert the multi-output results of the classification head into corresponding disease classification descriptions, and concatenate them with the text description and fine-grained features of each layer of 2D medical images and input them into the large language model LLM to obtain the disease text output.

2. The medical image analysis method based on fine-grained features according to claim 1, characterized in that: S1 imitates the reading mode of doctors when viewing medical images, and divides 3D medical images into multiple layers of 2D medical images for input, including: Input 3D medical images of size (B, Z, 3, H, W). Each 3D medical image in each batch contains Z layers of 3×H×W 2D medical images. These 2D medical images are input to simulate the reading mode of doctors when viewing medical images. Among them, B is the batch, Z is the number of layers of 2D medical images contained in each 3D medical image, and H and W are the height and width of each layer of 2D medical images, respectively.

3. The medical image analysis method based on fine-grained features according to claim 2, characterized in that: In S2, each layer of 2D medical image is ViT-encoded to obtain the ViT encoding result of each layer of 2D medical image, including: ViT encoding of each layer of 2D medical images: X z =ViT(S z ); Among them, S z is the z-th layer 2D medical image, ViT represents ViT coding, X z is the ViT encoding result of the z-th layer 2D medical image, z = 1, 2, ..., Z; For each layer of 2D medical images, after ViT encoding, a ViT encoding result containing L visual tokens is obtained, and each visual token contains C feature dimensions; for each 3D medical image, the size of the ViT encoding result obtained after ViT encoding is Z×L×C.

4. The medical image analysis method based on fine-grained features according to claim 3, characterized in that: In S3, the ViT encoding results of each layer of 2D medical images are upsampled and downsampled respectively to obtain the corresponding upsampled features and downsampled features, including: S31. Upsample the ViT encoding result of each layer of 2D medical image to obtain the corresponding upsampled features: F z up =Deconv2D(X z ); Among them, Deconv2D represents upsampling using two-dimensional deconvolution. In two-dimensional deconvolution, the step size is 2, the convolution kernel size is 3, and F z up is the upsampled feature of the z-th layer 2D medical image; S32, downsample the ViT encoding result of each layer of 2D medical image to obtain the corresponding downsampled features: F z down =Conv2D(X z ); Among them, Conv2D represents downsampling using two-dimensional convolution. In the two-dimensional convolution, the step size is 2, the convolution kernel size is 3, and F z down is the down-sampled feature of the z-th layer 2D medical image.

5. The medical image analysis method based on fine-grained features according to claim 4, characterized in that: In S4, based on the down-sampled features of each layer of 2D medical images, the hierarchical features of each layer of 2D medical images are extracted using a hierarchical feature sampler, including: S41, downsample the feature F of the z-th layer 2D medical image z down As the sum of native information T of the visual features of this layer z ori , and initialize multiple learnable tokens for this layer as T z learn , to learn layer information to solve the problem that the same encoder lacks understanding of visual features when processing different layers of information; S42. Extract hierarchical features of each layer of 2D medical images using a hierarchical feature sampler: T z layer =CrossAttn(Concat(T z ori ;T z learn );F z down ); Among them, Concat represents concatenation in the length dimension, CrossAttn represents the cross attention mechanism, and T z layer It is the hierarchical feature of the z-th layer 2D medical image, containing T+1 global visual information tokens.

6. The medical image analysis method based on fine-grained features according to claim 5, characterized in that: In S5, based on the hierarchical features of each layer of 2D medical images, a hierarchical feature aggregator is used to extract the hierarchical aggregation features of each layer of 2D medical images, including: The hierarchical feature aggregator is used to extract hierarchical aggregation features of each layer of 2D medical images: T z layer_fusion =Conv(T z layer ); Among them, Conv represents feature aggregation using convolution. If the 2D information of the layer is not lost, 3D convolution is used for feature aggregation. In 3D convolution, the step size is 1×1×1 and the window size is 3×3×3. If the 2D information of the layer is lost, the positions of the global visual information tokens in the hierarchical features between different layers still correspond one to one, so 2D convolution is used for feature aggregation. In order to ensure that the dimension of the tokens in the hierarchical aggregation features remains unchanged, 2D convolution / 3D convolution with padding is used. T z layer_fusion is the hierarchical aggregation feature of the z-th layer 2D medical image.

7. The medical image analysis method based on fine-grained features according to claim 6, characterized in that: In S6, the fine-grained features of each layer of 2D medical images are extracted using a fine-grained feature sampler based on the up-sampled features and hierarchical aggregation features of each layer of 2D medical images, including: S61, upsample the feature F of the z-th layer 2D medical image z up As keys and values, the hierarchical aggregation features T of the z-th layer 2D medical image are z layer_fusion As queries, perform cross-attention interactions; S62. Extract fine-grained features of each layer of 2D medical images using a fine-grained feature sampler: T z fine =CrossAttn(F z up ;T z layer_fusion ); Among them, T z fine is the fine-grained feature of the z-th layer 2D medical image.

8. The medical image analysis method based on fine-grained features according to claim 7, characterized in that: In S7, a multi-output classification head is constructed based on the hierarchical aggregation features of each layer of 2D medical images, and the corresponding disease classification labels are obtained based on the multi-output results of the classification head, including: S71, due to the hierarchical aggregation feature T of the z-th layer 2D medical image z layer_fusion The features of this layer and its adjacent layers are aggregated, and the high-level semantic features of the original image are extracted. Therefore, a multi-output classification head is constructed through pooling and MLP layer mapping: Among them, Pool represents pooling, MLP represents MLP layer mapping, Y is the multi-output result of the classification head, and each output result is the disease classification feature of a specified type of disease; S72, obtaining corresponding disease classification labels according to the multiple output results of the classification head; Among them, the loss function of classification head training is BCE Loss.

9. The medical image analysis method based on fine-grained features according to claim 8, characterized in that: In S8, the multiple output results of the classification head are converted into corresponding disease classification descriptions, which are then concatenated with the text descriptions and fine-grained features of each layer of 2D medical images and input into the large language model LLM to obtain disease text outputs, including: S81. Convert the multiple output results of the classification head into corresponding disease classification descriptions, and concatenate them with the text description and fine-grained features of each layer of 2D medical images to obtain the input of the large language model LLM: Among them, T Prompt is the text description of each layer of 2D medical image, T Class_Prompt It converts the multi-output results of the classification head into corresponding disease classification descriptions. LLM_Input is the input of the large language model LLM. S82, inputting the concatenation result into the large language model LLM to obtain the disease text output; Among them, the loss function of the large language model LLM training is LLM Loss.

10. The medical image analysis method based on fine-grained features according to claim 9, characterized in that: The overall loss function for the classification head and large language model LLM training is: L=α·L BCE +L LLM ; Among them, L BCE is the loss function BCE Loss of the classification head training, α is the weight coefficient of the loss function of the classification head training, L LLM LLM Loss is the loss function for training the large language model LLM, and L is the overall loss function.