Medical image report generation method based on multi-modal large model retrieval enhancement
By combining multimodal large models and dynamic retrieval technology, the problems of lack of historical case references and information omissions in imaging report generation in existing technologies are solved, and more accurate and efficient imaging report generation is achieved.
Patent Information
- Application Number
- CN202510541999.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent medical imaging diagnostic report generation method lacks reference to similar historical cases, which easily leads to overlooking key lesion characteristics and causing diagnostic errors. In addition, the fixed retrieval strategy leads to interference from irrelevant historical data or omission of effective information.
A medical imaging report generation method based on a multimodal large model is adopted. By combining the image vector model module, the vector database and similarity detection module, the multimodal large model module, the dynamic retrieval module and the retrieval enhancement generation prompt engineering module, image-text joint reasoning and diagnosis report generation are realized.
It improves the robustness of medical image feature representation, enhances the accuracy and efficiency of image report generation, reduces the risk of missing key information, and improves the credibility of diagnostic results.
Smart Images

Figure CN120673959A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a method for generating medical image reports based on multimodal large model retrieval enhancement. Background Art
[0002] With the development of artificial intelligence technology, large language models have demonstrated powerful text processing and generation capabilities. However, in addition to text modal data, the real world also has a large amount of image modal data. Therefore, multimodal large models have also developed rapidly. Multimodal large models can process image and text inputs at the same time, understand image output responses based on text instructions, and the field of medical imaging is a strong multimodal scenario. Doctors need to assist in diagnosis based on patient images. Writing reports based on images is one of the most important links. Writing imaging reports relies on extremely high medical imaging knowledge and takes a lot of doctors' time. At the same time, due to the uneven level of doctors, the quality of written reports is also uncontrollable. Multimodal large models have powerful image-text understanding capabilities. Therefore, medical imaging report generation methods based on multimodal large models have become a research hotspot.
[0003] The current method of generating intelligent medical imaging diagnostic reports lacks reference to similar historical cases, and is prone to overlooking key lesion characteristics, resulting in diagnostic errors; and the fixed retrieval strategy leads to interference from irrelevant historical data or omission of effective information; and currently, the image pictures are directly input into the multimodal large model to generate the corresponding imaging report, but the generated imaging report may contain errors. Summary of the Invention
[0004] In order to solve the problems of insufficient generation accuracy and difficulty in balancing efficiency and precision raised in the above background technology, the purpose of the present invention is to provide a medical imaging report generation method based on multimodal large model retrieval enhancement.
[0005] To achieve the above objectives, the present invention provides the following technical solutions: a method for generating medical imaging reports based on multimodal large model retrieval enhancement, comprising an image vector model module, a vector database and similarity detection module, a multimodal large model module, a dynamic retrieval module, and a retrieval enhancement generation prompt engineering module;
[0006] The multimodal large model module includes a multimodal input splicing unit, an inference information processing unit, and a diagnosis report set generation unit;
[0007] The multimodal input splicing unit performs model splicing on the new image and the reference report, the inference information processing unit selects the top K indexes with the highest similarity through the multimodal large model, and the diagnosis report set generation unit generates reference diagnosis report set information based on the reports in the selected large model;
[0008] Reference diagnostic report set selection formula:
[0009] (3);
[0010] In the above formula (3), is the total number of images in the database, For Select the top K indexes with the highest similarity from the samples. is the diagnostic report text corresponding to the k-th similar image, is the collection of retrieved Top-K reference reports;
[0011] The dynamic retrieval module includes a similarity threshold analysis unit and a retrieval quantity analysis unit;
[0012] The similarity threshold analysis unit is used to analyze the maximum similarity value between the new image and the database, and the retrieval quantity analysis unit can determine the number of retrieved diagnostic reports based on the maximum similarity value;
[0013] The number of retrieved diagnostic reports is calculated as follows:
[0014] (5);
[0015] In the above formula (5), is the maximum similarity value between the new image and the database [i.e. max(s i )], is the preset threshold, The number of searches (1 / 3 / 5) is dynamically adjusted based on similarity.
[0016] Preferably, The preset threshold is: =0.7-0.9, =0.3-0.6.
[0017] Preferably, the image vector model module includes a graphics preprocessing unit, a feature extraction unit and an L2 normalization unit;
[0018] The image preprocessing unit processes the image into a predetermined resolution, generates an original feature vector by calling the CLIP encoder through the feature extraction unit, and normalizes the vector through the L2 normalization unit;
[0019] The unit vector is calculated as follows:
[0020] (1);
[0021] In the above formula (1), is the i-th medical image (input image), is the image encoding function of the CLIP model (mapping the image into a vector), is the L2 norm (Euclidean length of the vector), is the normalized image feature vector, is a trainable pixel encoder, (H is height; W is width; C is the number of channels).
[0022] Preferably, The dimension of the normalized image feature vector is 512.
[0023] Preferably, the vector database and similarity monitoring module includes a vector database unit, a sorting and screening unit, an index unit and a similarity comparison unit;
[0024] The vector database unit stores historical image feature vectors The sorting and screening unit sorts and screens according to the balance accuracy and speed, and obtains the cosine similarity value of the similar vector according to the similarity comparison unit. ;
[0025] Cosine similarity value Calculated by the following formula:
[0026] (2);
[0027] In the above formula (2), For the new input image I new The normalized eigenvector of is the normalized feature vector of the i-th image in the database, is the cosine similarity value.
[0028] Preferably, The range of cosine similarity value is [-1, 1].
[0029] Preferably, the retrieval enhancement generation prompt engineering module includes a diagnosis report generation unit, a dual-channel verification unit, and a few-sample learning unit;
[0030] The diagnostic report generation unit can generate a final diagnostic report based on the image and the reference diagnostic report set, the dual-channel verification unit is used to verify the diagnostic report, and the few-sample learning unit can add the verified diagnostic report to the vector database;
[0031] The final diagnostic report is generated as follows:
[0032] (4);
[0033] In the above formula (4), For the final diagnosis report, For newly input medical images, for multimodal generative models (such as CogVLM), is the generated report text sequence (T is the text length), Based on historical words , the conditional probabilities of the new image and the reference report, Generate text sequences that maximize probability through autoregression.
[0034] Preferably, For newly input medical images, CT, MRI or X-ray images are preferred.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. This invention utilizes image vectorization processing technology based on the CLIP-ViT model, combined with a department-adaptive dynamic resolution adjustment mechanism (configurable from 244×244 to 488×488). Through feature fusion and L2 normalization, it significantly improves the robustness of medical image feature representation. It utilizes the cosine similarity algorithm to construct a high-dimensional vector space, enabling precise cross-modal matching in a million-level distributed vector database (such as Milvus). Its HNSW indexing technology optimizes the balance between retrieval accuracy and speed to millisecond-level response, ensuring a more than 40% increase in similar case matching efficiency. This enables the system to accurately capture subtle lesion features and reduce the risk of missing critical information.
[0037] 2. This invention introduces a dynamic threshold adjustment algorithm that effectively adjusts the number of searches (1-5) based on a similarity threshold. When the maximum similarity between a new image and a historical case exceeds 0.9, the system automatically focuses on the top-1, highest-quality reference report. When the similarity falls below 0.6, the system expands to the top-5 to enhance contextual understanding. This flexible mechanism reduces the interference of irrelevant data while increasing effective information coverage, addressing the rigidity of traditional fixed search strategies.
[0038] 3. This invention leverages the cross-modal attention mechanism of a multimodal large model (e.g., CogVLM) to deeply align new image features with the textual semantics of the top-K reference reports. This method implements joint image-text reasoning based on vector data and a probability generation strategy calculated by a similarity monitoring module. A unique dual-channel verification unit generates two versions of the report through independent two-way reasoning, which are then cross-validated with the initial results. Combined with a three-dimensional error assessment mechanism, this significantly reduces diagnostic error rates, thereby significantly improving clinical credibility.
[0039] The parts not involved in the device are the same as those in the prior art or can be implemented by using the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a system block diagram of the medical imaging report generation method based on multimodal large model retrieval enhancement of the present invention.
[0041] Figure 2 This is a simplified flowchart of the method for generating medical imaging reports based on multimodal large model retrieval enhancement of the present invention. DETAILED DESCRIPTION
[0042] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods.
[0043] like Figure 1 As shown, the medical imaging report generation method based on multimodal large model retrieval enhancement provided by the present invention includes an image vector model module, a vector database and similarity detection module, a multimodal large model module, a dynamic retrieval module and a retrieval enhancement generation prompt engineering module.
[0044] The image vector model module is communicated with the vector database and the similarity monitoring module, the vector database and the similarity monitoring module are communicated with the detection enhancement generation prompt engineering module, the vector database and the similarity monitoring module are communicated with the multimodal large model module, the vector database and the similarity monitoring module are communicated with the dynamic retrieval module, the multimodal large model module is communicated with the dynamic retrieval module, and the retrieval enhancement generation prompt engineering module is communicated with the dynamic detection module.
[0045] The image vector model module includes a graphics preprocessing unit, a feature extraction unit, and an L2 normalization unit. The image preprocessing unit processes the image to a predetermined resolution, and the feature extraction unit calls the CLIP encoder (CLIP-ViT model) to generate the original feature vector. The L2 normalization unit normalizes the vector. In this embodiment, the image needs to be normalized to a predetermined value between 244*244 and 488*488, and the pixel value is normalized to [-1, 1].
[0046] The unit vector is calculated as follows:
[0047] (1);
[0048] In the above formula (1), is the i-th medical image (input image), is the image encoding function of the CLIP model (mapping the image into a vector), is the L2 norm (Euclidean length of the vector), is the normalized image feature vector (dimension is 512), is a trainable pixel encoder, (H is height; W is width; C is the number of channels).
[0049] In this embodiment, the present application needs to set a predetermined resolution for medical image processing according to different medical departments. For example, if a higher image resolution effect is not required, such as dental or finger X-ray images, the medical image is standardized to 244*244 through the graphics preprocessing unit. If a higher image resolution effect is required, such as CT images of brain or otolaryngology, the medical image is standardized to 488*488 through the graphics preprocessing unit. By performing feature fusion operations on the pixel encoder and the model encoding function, the stability of normalization is further improved, the robustness of the cosine similarity calculation is improved, and the cost of computing power is reduced.
[0050] The vector database and similarity monitoring module includes a vector database unit, a sorting and screening unit, an index unit and a similarity comparison unit. The vector database unit stores historical image feature vectors. The sorting and screening unit sorts and screens according to the balance accuracy and speed, and obtains the cosine similarity value of the similar vector according to the similarity comparison unit. ;
[0051] Cosine similarity value Calculated by the following formula:
[0052] (2);
[0053] In the above formula (2), For the new input image I new The normalized eigenvector of is the normalized feature vector of the i-th image in the database, is the cosine similarity value (range [-1, 1], the larger the value, the more similar it is).
[0054] In this embodiment, the vector database unit adopts a distributed storage architecture to support real-time query of millions of vectors; the HNSW index is constructed by the index unit to balance the retrieval speed and accuracy, and the sorting and screening unit is used to sort the vectors. Sort in descending order.
[0055] The multimodal large model module includes a multimodal input splicing unit, an inference information processing unit, and a diagnosis report set generation unit. The multimodal input splicing unit performs model splicing on the new image and the reference report. The inference information processing unit selects the top K indexes with the highest similarity through the multimodal large model. The diagnosis report set generation unit generates reference diagnosis report set information based on the reports in the selected large model.
[0056] Reference diagnostic report set selection formula:
[0057] (3);
[0058] In the above formula (3), is the total number of images in the database, For Select the top K indexes with the highest similarity from the samples. is the diagnostic report text corresponding to the k-th similar image, is the collection of retrieved Top-K reference reports.
[0059] The retrieval enhancement generation prompt engineering module includes a diagnosis report generation unit, a dual-channel verification unit, and a few-shot learning unit. The diagnosis report generation unit can generate a final diagnosis report based on the image and reference diagnosis report set. The dual-channel verification unit is used to verify the diagnosis report. The few-shot learning unit can add the verified diagnosis report to the vector database.
[0060] The final diagnostic report is generated as follows:
[0061] (4);
[0062] In the above formula (4), For the final diagnosis report, is a new input medical image (CT, MRI or X-ray image), for multimodal generative models (such as CogVLM), is the generated report text sequence (T is the text length), Based on historical words , the conditional probabilities of the new image and the reference report, Generate text sequences that maximize probability through autoregression.
[0063] The dynamic retrieval module includes a similarity threshold analysis unit and a retrieval quantity analysis unit. The similarity threshold analysis unit is used to analyze the maximum similarity value between the new image and the database, and the retrieval quantity analysis unit can determine the number of diagnostic reports to be retrieved based on the maximum similarity value.
[0064] The number of retrieved diagnostic reports is calculated as follows:
[0065] (5);
[0066] In the above formula (5), is the maximum similarity value between the new image and the database (i.e. max(s i ), For the preset threshold (such as =0.9, =0.6), The number of searches (1 / 3 / 5) is dynamically adjusted based on similarity.
[0067] It should be noted that if Figure 2 As shown, the implementation steps of this application are as follows:
[0068] Step 1: Upon receiving the medical image input, the system immediately initiates a department-adaptive intelligent preprocessing pipeline: based on preset clinical department-resolution mapping rules (e.g., dental X-rays are standardized to 244×244, neurosurgery MRIs are adjusted to 488×488, and cardiovascular CTAs are optimized to 512×512), the feature extraction unit then calls the CLIP encoder (CLIP-ViT model) to generate the original feature vector.
[0069] Step 2: Using the existing image vectors and the high-dimensional vectors output by the feature extraction module, the system constructs a distributed multimodal vector database for medical scenarios (vector databases such as Milvus can be used), establishes key-value pairs of image vectors and corresponding images and reports, and constitutes a vector database information library. The image vector library and the image information library together constitute the vector database, and designs an "image vector library-image information library" dual storage engine architecture.
[0070] Step 3: The original feature vector for generating the image report is merged with the vector database obtained in the previous step into a new model.
[0071] Step 4: Refer to formula (5) above and adjust the number of reference reports according to the similarity.
[0072] Step 5: Compare the original feature vector obtained in step 1 with similar vectors in the vector database, and select several medical images and reports that are most similar to the original feature vector from the vector database.
[0073] Step 6: The sorting and screening unit performs an in-depth comparison analysis with the original feature vector based on the reference medical images and corresponding reports obtained in step 5 that have the highest similarity to the original feature vector, thereby implementing precise screening and scientific sorting. Subsequently, the indexing unit carefully annotates each of the selected and sorted reference medical images with prompt words and conducts in-depth logical reasoning based on these prompt words.
[0074] Step 7: Based on the detailed results of the reasoning analysis in Step 6, the diagnostic report generation unit officially begins its work. Leveraging advanced algorithms and professional medical knowledge, this unit comprehensively integrates and deeply interprets the reasoning results, generating an accurate and comprehensive impact report for the new medical image. This report not only covers the key information presented by the image, but also incorporates medical logic and clinical experience.
[0075] Step 8: To ensure the accuracy and reliability of the diagnosis, the dual-channel verification unit performs the entire process from Step 1 to Step 7 twice. These two verification processes are independent of each other, and each performs reasoning and analysis based on its own data and algorithms, resulting in two different diagnostic results.
[0076] Step 1: After completing the three diagnostic processes, the system will evaluate the error of the three results. When the errors of the three results are all within the preset standard range, it indicates that the three diagnostic results have high accuracy and consistency. At this time, the system will formally incorporate the impact report generated based on the new medical image into the database and build a new vector database based on it. This measure not only enriches the content of the database, but also provides more reference samples for subsequent medical image analysis, which helps to improve the diagnostic capabilities of the entire system; however, if the error of the three results does not meet the preset standard, it means that there is a large uncertainty or deviation in the three diagnostic results. In order to ensure the accuracy and reliability of the data in the database, the system will decisively delete the three results and re-incorporate the new image into the process of step one for comprehensive and in-depth analysis. Through this cyclic iterative method, the diagnostic results are continuously optimized until the expected accuracy requirements are met.
[0077] In this document, relational terms such as first and second, etc., are used solely to distinguish one entity or operation from another entity or operation and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed or that are inherent to such process, method, article, or apparatus.
[0078] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A medical imaging report generation method based on multimodal large model retrieval enhancement, characterized in that: It includes image vector model module, vector database and similarity detection module, multimodal large model module, dynamic retrieval module and retrieval enhancement generation prompt engineering module; The multimodal large model module includes a multimodal input splicing unit, an inference information processing unit, and a diagnosis report set generation unit; The multimodal input splicing unit performs model splicing on the new image and the reference report, the inference information processing unit selects the top K indexes with the highest similarity through the multimodal large model, and the diagnosis report set generation unit generates reference diagnosis report set information based on the reports in the selected large model; Reference diagnostic report set selection formula: (3); In the above formula (3), is the total number of images in the database, For Select the top K indexes with the highest similarity from the samples. is the diagnostic report text corresponding to the k-th similar image, is the collection of retrieved Top-K reference reports; The dynamic retrieval module includes a similarity threshold analysis unit and a retrieval quantity analysis unit; The similarity threshold analysis unit is used to analyze the maximum similarity value between the new image and the database, and the retrieval quantity analysis unit can determine the number of retrieved diagnostic reports based on the maximum similarity value; The number of retrieved diagnostic reports is calculated as follows: (5); In the above formula (5), is the maximum similarity value between the new image and the database [i.e. max(s i )], is the preset threshold, The number of searches (1 / 3 / 5) is dynamically adjusted based on similarity.
2. The method for generating medical imaging reports based on multimodal large model retrieval enhancement according to claim 1, characterized in that: The preset threshold is: =0.7-0.9, =0.3-0.
6.
3. The method for generating medical imaging reports based on multimodal large model retrieval enhancement according to claim 1, characterized in that: The image vector model module includes a graphics preprocessing unit, a feature extraction unit and an L2 normalization unit; The image preprocessing unit processes the image into a predetermined resolution, generates an original feature vector by calling the CLIP encoder through the feature extraction unit, and normalizes the vector through the L2 normalization unit; The unit vector is calculated as follows: (1); In the above formula (1), is the i-th medical image (input image), is the image encoding function of the CLIP model (mapping the image into a vector), is the L2 norm (Euclidean length of the vector), is the normalized image feature vector, is a trainable pixel encoder, (H is height; W is width; C is the number of channels).
4. The method for generating medical imaging reports based on multimodal large model retrieval enhancement according to claim 3, characterized in that: The dimension of the normalized image feature vector is 512.
5. The method for generating medical imaging reports based on multimodal large model retrieval enhancement according to claim 3, characterized in that: The vector database and similarity monitoring module includes a vector database unit, a sorting and screening unit, an index unit and a similarity comparison unit; The vector database unit stores historical image feature vectors The sorting and screening unit sorts and screens according to the balance accuracy and speed, and obtains the cosine similarity value of the similar vector according to the similarity comparison unit. ; Cosine similarity value Calculated by the following formula: (2); In the above formula (2), For the new input image I new The normalized eigenvector of is the normalized feature vector of the i-th image in the database, is the cosine similarity value.
6. The method for generating medical imaging reports based on multimodal large model retrieval enhancement according to claim 5, characterized in that: The range of cosine similarity value is [-1, 1].
7. The method for generating medical imaging reports based on multimodal large model retrieval enhancement according to claim 5, characterized in that: The retrieval enhancement generation prompt engineering module includes a diagnosis report generation unit, a dual-channel verification unit, and a few-sample learning unit; The diagnostic report generation unit can generate a final diagnostic report based on the image and the reference diagnostic report set, the dual-channel verification unit is used to verify the diagnostic report, and the few-sample learning unit can add the verified diagnostic report to the vector database; The final diagnostic report is generated as follows: (4); In the above formula (4), For the final diagnosis report, For newly input medical images, for multimodal generative models (such as CogVLM), is the generated report text sequence (T is the text length), Based on historical words , the conditional probabilities of the new image and the reference report, Generate text sequences that maximize probability through autoregression.
8. The method for generating medical imaging reports based on multimodal large model retrieval enhancement according to claim 7, characterized in that: For newly input medical images, CT, MRI or X-ray images are preferred.
Citation Information
Cited By
Power report generation method, system and equipment based on adaptive learning and medium
CN120873028A
Multi-modal ultrasonic data processing and report generating method and system based on retrieval enhancement
CN121415975A