Focus positioning method based on large model
Through a large-model-based lesion positioning method, combined with image coding and large language model, the problem that doctors in the prior art needs to manually compare images and reports is solved, and diagnostic efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202411974161.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing medical imaging analysis techniques require doctors to manually compare the location of lesions in the images and the text descriptions in the report, resulting in inefficient diagnosis, increasing the workload of doctors, and may lead to misjudgment or omissions.
The lesion positioning method based on the large model is adopted, and the global features are extracted through an image encoder, combined with the lesion detection and segmentation modules, and organized into a feature map with the same size, and input it into a large language model, integrating the content to output the lesion positioning image and text description.
The steps of manual comparison are avoided, which improves diagnostic efficiency, reduces the workload of doctors, and enhances the accuracy and consistency of diagnosis.
Smart Images

Figure CN119941649A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image analysis, and in particular relates to a lesion localization method based on a large model. Background Art
[0002] In the field of medical image analysis, computer-aided diagnosis (CAD) technology has gradually become an important tool for clinicians, assisting doctors in lesion detection and diagnosis by automatically processing image data. Existing image analysis technologies mainly rely on algorithms such as target detection and segmentation in deep learning. The development of these technologies provides a basis for the automated analysis of medical images.
[0003] Existing image analysis technology is limited to providing lesion detection or segmentation results, and these results lack direct relevance to the final diagnostic report written by the doctor. In actual operation, the doctor needs to manually compare the lesion location in the image with the text description in the report. This fragmented workflow makes it impossible to seamlessly connect the image analysis results with the clinical report, which increases the workload of the doctor and may lead to misjudgment or omissions in the diagnosis process. Especially in complex cases or with multiple lesions, doctors need to frequently switch between images and reports, which is not only time-consuming, but also increases the workload and reduces the efficiency of diagnosis.
[0004] To this end, we proposed a lesion localization method based on a large model to solve the above problems. Summary of the invention
[0005] The purpose of the present invention is to solve the problem of low diagnostic efficiency in the prior art, which requires manual comparison of the lesion location in the image and the text description in the report, and proposes a lesion localization method based on a large model.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for lesion localization based on a large model comprises the following steps:
[0008] S1: Input the original image into the image encoder in the large model processing module, and the image encoder extracts global features;
[0009] S2: inputting the global features processed by the image encoder into the lesion detection module and the lesion segmentation module of the input processing module;
[0010] S3: The lesion detection module detects the approximate location of the lesion and marks the area frame;
[0011] S4: The lesion segmentation module performs fine segmentation on the lesion based on the region frame of the lesion detection module as a prompt;
[0012] S5: the regional encoder in the input processing module processes each lesion feature and arranges it into feature maps of the same size;
[0013] S6: The lesion feature code in the region encoder is input into the knowledge base module for comparison;
[0014] S7: The global features in the image encoder, the lesion feature encoding in the region encoder, the similar lesion features in the knowledge base module, and the text prompt input by the user are all input into the large language model (LLM) in the large model processing module;
[0015] S8: The large language model (LLM) integrates content through multimodal task adaptation and outputs lesion localization images and text descriptions.
[0016] Preferably, step S3 comprises:
[0017] S31: The lesion detection module marks a regional frame for the approximate location of the detected lesion;
[0018] S32: Use non-maximum suppression and knowledge filtering methods to retain the 20 region boxes with the highest confidence.
[0019] Preferably, step S4 comprises:
[0020] S41: the lesion segmentation module segments the lesion and submits it to the user for review;
[0021] S42: The user inputs prompt information, and the lesion segmentation module performs further fine division.
[0022] Preferably, step S5 comprises:
[0023] S51: the region encoder extracts four feature layers of the encoding process of the image encoder;
[0024] S52: the region encoder extracts according to the region frame detected by the lesion detection module;
[0025] S53: the region encoder performs attention enhancement for extracting the potential region according to the weight information of the mask of the target lesion, while retaining some background information;
[0026] S54: The region encoder uses the ROI Align method to organize each target lesion into a feature map of the same size, 14x14 <pn>.
[0027] Preferably, in step S7, similar lesion features in the knowledge base are synchronously provided to the user, and the user provides similarity feedback.
[0028] Preferably, the lesion detection module is based on a dinov2 model retrained in medical imaging data and shares an image encoder with the lesion segmentation module.
[0029] Preferably, the lesion detection module is composed of a SAM2 model adjusted image encoder.
[0030] To sum up, the technical effects and advantages of the present invention are as follows: compared with the existing devices, the lesion localization method based on the large model processes each lesion feature through a regional encoder and organizes it into a feature map of consistent size, which is then input into the large language model. The large language model integrates the feature map, global features, similar cases and the user's prompt, and outputs a medical imaging report. Compared with the existing technology, it avoids the need to manually compare the lesion location in the image and the text description in the report, thereby improving the diagnostic efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a structural block diagram of the present invention. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0033] Reference Figure 1 A lesion localization system based on a large model includes an input processing module, a large model processing module and a knowledge base module. The input processing module includes a lesion detection module, a lesion segmentation module and a region encoder. The large model processing module includes a large model module (LLM) and an image encoder.
[0034] The knowledge base module is used to store global features, regional features, and medical report annotations for easy subsequent use.
[0035] Image encoders are used to extract global features from medical images and provide high-quality input for subsequent analysis.
[0036] The lesion detection module is used for preliminary detection to detect the approximate location of the lesion in the medical image.
[0037] The lesion segmentation module is used to further refine the segmentation of the output of the lesion detection module.
[0038] The regional encoder is used to encode the lesion features so that they can be subsequently fed into the large model for use.
[0039] The Large Model Module (LLM) is used to integrate content and output inference results.
[0040] A method for lesion localization based on a large model comprises the following steps:
[0041] S1: The original image is input into the image encoder in the large model processing module, and the image encoder extracts global features.
[0042] S2: Input the global features processed by the image encoder into the lesion detection module and the lesion segmentation module of the input processing module.
[0043] S3: The lesion detection module detects the approximate location of the lesion and marks the area box.
[0044] S4: The lesion segmentation module performs fine segmentation of the lesion based on the region box of the lesion detection module as a hint.
[0045] S5: The regional encoder in the input processing module processes each lesion feature and organizes it into feature maps of consistent size so that it can be subsequently fed into the large model for use.
[0046] S6: The lesion feature encoding in the region encoder is input into the knowledge base module for comparison.
[0047] S7: The global features in the image encoder, the lesion feature encoding in the region encoder, the similar lesion features in the knowledge base module, and the text prompt input by the user are all input into the large language model in the large model processing module.
[0048] S8: The Large Language Model (LLM) integrates content through multimodal task adaptation and outputs lesion localization images and text descriptions. The Large Language Model (LLM) can complete tasks such as medical report generation, lesion detection, and imaging diagnosis.
[0049] The large model processing module includes a large model module (LLM) and an independent image encoder, as well as a conventional projection layer module and a parsing module.
[0050] The projection layer module maps the output of the image encoder to adapt to the feature requirements of subsequent tasks. This module maps global features and regional features to an encoding format suitable for large model processing through dimensionality reduction and feature transformation. It mainly includes: feature transformation and adaptation of global and regional features;<global image> The label is replaced by the encoded global image feature; <p1> 、 <p2>The region labels are replaced with the corresponding region encoding features.
[0051] The parsing module is responsible for decoding and presenting the task results. This module decodes the output of the large model and performs structural processing. This includes parsing the encoded information into structured report information and <p1>The regional labels are matched to the specific regional image labels, and the visual images are generated based on the global information.
[0052] Since the projection layer module and the analysis module are relatively common in the field of large models, they are not described in detail and are not drawn in the schematic diagram.
[0053] This large model-based lesion localization method provides doctors with the ability to automatically generate reports and accurately locate lesions by combining large model technology with a variety of interactive tools.
[0054] Step S3 comprises.
[0055] S31: The lesion detection module marks a regional frame for the approximate location of the detected lesion;
[0056] S32: Use non-maximum suppression and knowledge filtering methods to retain the 20 region boxes with the highest confidence.
[0057] Step S4 comprises.
[0058] S41: The lesion segmentation module segments the lesion and submits it to the user for review;
[0059] S42: The user inputs prompt information, and the lesion segmentation module performs further fine division.
[0060] As a promptable segmentation module, the lesion segmentation module provides more modal choices for user input. Instead of just using the area box in the above-mentioned lesion detection module as a prompt, the module can also directly accept the user's input of points, drawing areas, box selections and other prompt information, perform more detailed divisions based on user needs, interact with users, and determine more precise segmentation targets.
[0061] Step S5 comprises.
[0062] S51: The region encoder extracts four feature layers of the encoding process of the image encoder. On the image encoder in the large model processing module, four feature layers are extracted from the encoding process to form a hierarchical feature pyramid.
[0063] S52: The region encoder extracts the region frame according to the region frame detected by the lesion detection module.
[0064] S53: The region encoder performs attention enhancement to extract the potential region based on the weight information of the mask of the target lesion, while retaining some background information.
[0065] S54: The regional encoder uses the ROI Align method to organize each target lesion into a feature map of the same size, 14x14 <pn>, so that it can be sent to the large model for subsequent use.
[0066] In step S7, similar lesion features in the knowledge base are synchronously provided to the user, and the user provides similarity feedback. After the doctor provides feedback, the features are added to the modification list for subsequent updates to the knowledge base, thereby improving the accuracy and adaptability of lesion positioning.
[0067] The lesion detection module is based on the dinov2 model retrained from medical imaging data, and shares the image encoder with the lesion segmentation module. The training data of the lesion detection module is mainly the detection box in the medical image, which is expanded by 1.2 times as the GT box. During the training process, it is first pre-trained based on large-scale medical imaging data, and then fine-tuned using the fine data of the corresponding scene.
[0068] The lesion detection module is composed of an image encoder adjusted by the SAM2 model. After the lesion detection module is trained, the image encoding module of the dinov2 model is extracted as the basic parameters of the image encoding module of the SAM2 model.
[0069] Medical imaging report generation (including report generation and question corresponding box, segmentation) example:
[0070] Prompt:
[0071] [Detail Report Generation]<global image> For the above image, the following locations <p1> , <p2> , <p3> ,…… <pn>There may be lesions, please analyze them and write a detailed medical imaging report.
[0072] Res:
[0073] Findings:
[0074] The thorax is symmetrical, the trachea is in the middle, and the bony structure of the ribs above the diaphragm is intact. <d> Increased texture of both lungs< / d> <p1> 。 <d> High-density nodules are seen in the right upper lung field, left lower lung field, and the outer zone of the left lower lung field< / d> <p4> <p5>, the boundaries are clearer. <d> Hypertrophy of both lung hilar areas< / d> <p6> 。 <d> The heart is boot-shaped, the waist is sunken, and the apex is bulging and enlarged to the left.< / d> <p7> , <d> Aortic knot protrusion with calcification, and local calcification shifted inward< / d> <p8> 。 <d> Blunt right costophrenic angle< / d> <p9>。
[0075] Impression:
[0076] 1. <d> Old lesions in the right upper lung and left lower lung< / d> <p4> <p5>
[0077] 2. <d> Aortic heart shadow< / d> <p7> , <d> Calcification of aortic junction< / d> <p8>
[0078] 3. <d> Pleural effusion< / d> <p9>
[0079] This allows us to identify problematic descriptions, and based on these tags, retrieve previously generated detection boxes and segmentation information, and then provide a visual display analysis based on these.
[0080] Example of regional questions and answers:
[0081] Prompt:
[0082] [VQA]<global image> For the above image, the following locations <p4> , <p9>Please analyze whether there is a lesion.
[0083] Res:
[0084] The following problems were found
[0085] 1. <d> Old lesion in right upper lung< / d> <p4>
[0086] 2. <d> Blunt right costophrenic angle< / d> <p9>
[0087] For the above prompt,<global image> will be replaced by the full-image encoding information, and <p4> , <p9>These two labels will be followed by the local information encoded by the region encoder.
[0088] After large model inference, the above Res information can be output.
[0089] For the output Res information, we can use the <p4>or <p9>Labels, retrieve the coordinate values of the corresponding areas, and use color correspondence and visualization information to better display the corresponding relationship between the problem and the area.
[0090] Other general questions and answers:
[0091] With the support of multimodal large models, other common problems can also be answered, such as:
[0092] Report Correction Prompt:
[0093] [Correction] Please correct the following description based on the image and report content.
[0094] The large model compares the imaging features with the original report section by section, corrects inaccurate or missing diagnostic information, and generates a new revised report.
[0095] Multi-period comparative analysis Prompt:
[0096] [Multi-phase comparison] Please compare the changes in images A and B.
[0097] Extract multi-phase image codes, and then identify the progression or changes of lesions through feature matching and time series reasoning to generate comparative analysis results.
[0098] Prompt for generating diagnosis and treatment suggestions:
[0099] [Recommendations] Provide diagnosis and treatment plans based on the patient's images and reports.
[0100] Utilize the medical knowledge base of the big model and combine the disease characteristics in the images and reports to automatically generate personalized diagnosis and treatment recommendations.
[0101] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention. < / p4> < / p4> < / p4> < / p8> < / p7> < / p5> < / p4> < / p8> < / p7> < / p6> < / p4> < / p1> < / pn> < / p3> < / p2> < / p1> < / pn> < / p1> < / pn>
Claims
1. A lesion localization method based on a large model, characterized in that: Includes steps: S1: Input the original image into the image encoder in the large model processing module, and the image encoder extracts global features; S2: inputting the global features processed by the image encoder into the lesion detection module and the lesion segmentation module of the input processing module; S3: The lesion detection module detects the approximate location of the lesion and marks the area frame; S4: The lesion segmentation module performs fine segmentation on the lesion based on the region frame of the lesion detection module as a prompt; S5: the regional encoder in the input processing module processes each lesion feature and arranges it into feature maps of the same size; S6: The lesion feature code in the region encoder is input into the knowledge base module for comparison; S7: The global features in the image encoder, the lesion feature codes in the region encoder, the similar lesion features in the knowledge base module, and the text prompt input by the user are all input into the large language model in the large model processing module; S8: The large language model is adapted through multimodal tasks, integrates content, and outputs lesion localization images and text descriptions.
2. A lesion localization method based on a large model according to claim 1, characterized in that: The step S3 comprises: S31: The lesion detection module marks a regional frame for the approximate location of the detected lesion; S32: Use non-maximum suppression and knowledge filtering methods to retain the 20 region boxes with the highest confidence.
3. A lesion localization method based on a large model according to claim 1, characterized in that: The step S4 comprises: S41: the lesion segmentation module segments the lesion and submits it to the user for review; S42: The user inputs prompt information, and the lesion segmentation module performs further fine division.
4. A method for locating lesions based on a large model according to claim 1, characterized in that: The step S5 comprises: S51: the region encoder extracts four feature layers of the encoding process of the image encoder; S52: the region encoder extracts according to the region frame detected by the lesion detection module; S53: the region encoder performs attention enhancement for extracting the potential region according to the weight information of the mask of the target lesion, while retaining some background information; S54: The region encoder uses the ROIAlign method to organize each target lesion into a feature map of the same size, 14x14.
5. The method for locating lesions based on a large model according to claim 1, characterized in that: In step S7, similar lesion features in the knowledge base are synchronously provided to the user, and the user provides similarity feedback.
6. The method for locating lesions based on a large model according to claim 1, characterized in that: The lesion detection module is based on the dinov2 model retrained in medical imaging data and shares the image encoder with the lesion segmentation module.
7. A method for locating lesions based on a large model according to claim 6, characterized in that: The lesion detection module is composed of a SAM2 model adjusted image encoder.
Citation Information
Patent Citations
Medical model evaluation method and device, electronic equipment and storage medium
CN117407682A
Method, device and system for medical diagnosis and medical equipment
CN117637179A
Medical image report generation method and system and computer storage medium
CN117954041A
Medical report information extraction method and system, electronic equipment and readable storage medium
CN118675713A
Pathological diagnosis report generation method based on information retrieval
CN118782203A