Search-enhanced multimodal ultrasound data processing and report generation method, system

By decomposing the fetal ultrasound multimodal report generation task into single-organ sub-tasks using the Organ-Aware Routing Retrieval Enhanced Generation Framework (ORM-RAG), the many-to-many mapping problem was solved, and high accuracy and reliability of fetal ultrasound report generation were achieved.

CN121415975BActive Publication Date: 2026-04-17HUNAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN UNIV
Filing Date
2025-12-29
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing multimodal learning methods face the challenge of "many-to-many" mapping in fetal ultrasound report generation, leading to information contamination and the masking of abnormality detection, resulting in reports that lack specificity and diagnostic value.

Method used

The Organ-Aware Routing Retrieval Enhanced Generation Framework (ORM-RAG) is adopted. Through organ-level retrieval space decoupling and dynamic routing mechanism, the multi-organ retrieval task is decomposed into single-organ sub-tasks, and a multimodal large language model is combined to generate structured reports.

Benefits of technology

It improves the accuracy and clinical reliability of report generation, enabling precise understanding of multi-organ scenarios and the generation of standardized structured reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415975B_ABST
    Figure CN121415975B_ABST
Patent Text Reader

Abstract

This application relates to a method, system, computer device, storage medium, and computer program product for multimodal ultrasound data processing and report generation based on retrieval enhancement. The method involves acquiring an ultrasound image set containing at least two ultrasound images of different organs; analyzing the ultrasound images in the set and classifying them according to organ type; obtaining a report set for each classified ultrasound image; and finally generating a structured ultrasound report based on the report sets. This method solves the problem of existing technologies forcibly injecting a large amount of irrelevant information into the generation process, causing serious semantic pollution. It achieves the technical effect of decomposing the complex multi-organ retrieval task into multiple single-organ sub-tasks through organ-level retrieval space decoupling and dynamic routing mechanisms, thereby improving the accuracy and clinical reliability of report generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data processing and human-computer interaction technology, and in particular to a method, system, storage medium, and computer program product for multimodal ultrasound data processing and report generation based on retrieval enhancement. Background Technology

[0002] Fetal ultrasound, as the only safe, non-invasive, and widely used medical imaging method in prenatal screening, occupies an irreplaceable core position in modern perinatal medicine. Unlike single-organ imaging examinations for adults or children, fetal ultrasound screening requires the simultaneous evaluation of multiple key organ systems—including six major structural domains: brain, heart, face, abdomen, limbs, and placenta—during a single examination. This multi-organ assessment model is not only crucial for the early identification of severe congenital malformations (such as complex congenital heart disease, neural tube defects, and craniofacial developmental abnormalities), but also makes the diagnostic process far more complex than conventional radiological examinations. In clinical practice, a complete fetal ultrasound report is typically written by an experienced sonographer based on dozens or even hundreds of images, with detailed descriptions for each organ, combined with a comprehensive judgment of anatomical structures, hemodynamics, and growth parameters. This highly specialized, lengthy, and structured report generation process is extremely reliant on human labor, and the continued global shortage of qualified obstetric sonographers further exacerbates the clinical burden, highlighting the urgent need for automated report generation systems.

[0003] Against this backdrop, the introduction of artificial intelligence (AI) technology into the field of fetal ultrasound, especially the end-to-end automation from multimodal image input to structured natural language report output, has become an important frontier in medical AI research in recent years. However, the unique characteristics of fetal ultrasound present it with unique technical challenges. First, its "many-to-many" information correspondence constitutes a core bottleneck: a unified report contains multiple organ segments, and the information in each segment may come from multiple ultrasound cross-sectional images corresponding to that organ; conversely, each image only reflects a small part of the overall report. This fuzzy and misaligned mapping relationship makes traditional modeling paradigms based on single image-report pairs difficult to apply. Second, fetal ultrasound images themselves are highly heterogeneous—the anatomical morphology of different organs varies greatly, and the appearance of the same organ varies greatly at different gestational weeks and scanning angles; more importantly, pathological signs are often extremely subtle, and the visual difference between normal and abnormal images may be only a fraction of a millimeter, which places extremely high demands on the model's discrimination ability. In addition, medical report generation not only pursues fluent language but also emphasizes the accuracy of clinical facts and the reliability of diagnostic conclusions; any "illusion" or factual error may lead to serious clinical consequences. Therefore, how to achieve accurate understanding and structured representation of complex multi-organ scenes while ensuring high fidelity is a key issue that urgently needs to be addressed.

[0004] The rise of multimodal learning offers a new approach to solving these challenges. In the context of medical imaging, multimodal learning primarily refers to the deep fusion and collaborative modeling of visual information (ultrasound images) and linguistic information (diagnostic reports). In recent years, advanced technologies, represented by large-scale multimodal language models (MLLMs), have made significant progress in general-domain image-text understanding tasks. These models, through large-scale visual-linguistic pre-training, have initially acquired the ability to extract semantics from images and generate coherent text. However, when directly transferred to the highly specialized clinical scenario of fetal ultrasound, performance often suffers a significant drop. The fundamental reason lies in the huge "domain gap" between general pre-training data and specialized medical data: the model lacks prior knowledge of fetal anatomy, pathological features, and clinical terminology, making it difficult to capture those fine-grained local abnormalities crucial for diagnosis. To bridge this gap, Retrieval-Augmented Generation (RAG) technology has emerged. RAG (Real-Aspect-Organized Generative Model) effectively improves the factual accuracy and clinical relevance of its output by retrieving the most relevant past cases from a high-quality, domain-specific report library and providing this real-world "evidence" as context to the generative model. This method has demonstrated significant potential in single-organ image report generation tasks such as chest X-rays and CT scans.

[0005] However, existing RAG frameworks are almost entirely based on the "one-to-one" assumption—that is, one image corresponds to one report on a single organ. For example... Figure 1 As shown, this paradigm falls short when dealing with typical "many-to-many" complex scenarios like fetal ultrasound. A naive "one-to-multiple" strategy forces a large amount of irrelevant information into the generation process, causing serious semantic pollution. Conversely, a "multiple-to-multiple" strategy, where most images are normal, leads to search results dominated by normal samples, masking a few but crucial abnormal findings. The resulting reports lack specificity and diagnostic value. Neither strategy effectively decouples the complex cross-organ information flow, thus failing to meet the accuracy requirements of clinical practice. It is against this research backdrop that a dedicated methodology for multimodal processing and report generation in fetal ultrasound becomes particularly necessary. This is not merely a technical optimization issue, but a crucial step in enabling AI to empower complex clinical workflows. Building an intelligent system capable of understanding multi-organ structures, accurately locating abnormalities, and generating long, structured reports that conform to clinical standards has significance far exceeding mere efficiency improvements. Summary of the Invention

[0006] Therefore, it is necessary to provide a method, system, computer equipment, storage medium, and computer program product for multimodal ultrasound data processing and report generation based on retrieval enhancement to address the aforementioned technical problems.

[0007] In a first aspect, this application provides a method for retrieval-enhanced multimodal ultrasound data processing and report generation, the method comprising:

[0008] Acquire an ultrasound image set, wherein the ultrasound image set contains at least two ultrasound images of different organs;

[0009] The ultrasound images in the ultrasound image set are analyzed and classified according to organ type;

[0010] Obtain a collection of reports for each categorized ultrasound image;

[0011] Based on the report collection of various ultrasound images, a structured ultrasound report is finally generated;

[0012] The structured ultrasound report includes ultrasound images of various organs and their corresponding reports.

[0013] In one embodiment, the analysis of the ultrasound images in the ultrasound image set and the classification of the ultrasound images according to organ type includes:

[0014] The ultrasound images in the ultrasound image set are analyzed using an organ classifier, and the ultrasound images are classified according to organ type.

[0015] For a set of ultrasound images Using a pre-trained organ classifier Label recognition is performed on each image: ;

[0016] M: Number of ultrasound images.

[0017] In one embodiment, the acquisition of a report set of each classified ultrasound image includes:

[0018] Construct an organ-specific retrieval corpus;

[0019] Searches are performed in the retrieval corpus to generate a collection of reports.

[0020] In one implementation, constructing the organ-specific retrieval corpus includes:

[0021]

[0022] : Organ-specific retrieval corpus;

[0023] The i-th reference image belonging to organ c in the corpus;

[0024] :and The corresponding organ-specific report text fragment;

[0025] : Total number of images of organ c in the retrieved corpus - report pairs.

[0026] In one embodiment, the step of performing a retrieval query in the retrieval corpus and generating a report set includes:

[0027] Each organ subset Calculate the overall embedding and retrieve:

[0028]

[0029]

[0030] The average pooled vector of all image embeddings for organ c is used as a query representation to retrieve results from the corpus.

[0031] : The number of images belonging to organ c;

[0032] The i-th image classified as organ c;

[0033] : Organ-specific search function;

[0034] From the corpus The collection of reports retrieved from the database is used to generate the final report;

[0035] E(): Embedding function.

[0036] In one implementation, the step of generating a structured ultrasound report based on the report set of various ultrasound images includes:

[0037] Dynamic routing and confidence assessment are performed on the aforementioned report set;

[0038] Select a set of reports with a confidence level greater than the confidence threshold and generate a structured ultrasound report;

[0039]

[0040] The final structured ultrasound report;

[0041] Multimodal large language model generation function;

[0042] : The original set of ultrasound images;

[0043] Prompt word templates to guide report generation format and content;

[0044] : A set of high-confidence search results filtered through dynamic routing;

[0045] Confidence threshold Control the quality of search results through filtering;

[0046] : The overall confidence score of organ c search results.

[0047] Secondly, this application also provides a retrieval-enhanced multimodal ultrasound data processing and report generation system, the system comprising:

[0048] An image acquisition module is used to acquire a set of ultrasound images, wherein the set of ultrasound images includes at least two ultrasound images of different organs.

[0049] An image classification module is used to analyze the ultrasound images in the ultrasound image group and classify the ultrasound images according to organ type;

[0050] The report acquisition module is used to acquire a collection of reports for each categorized ultrasound image.

[0051] The report generation module is used to generate a structured ultrasound report based on the report set of various ultrasound images.

[0052] The structured ultrasound report includes ultrasound images of various organs and their corresponding reports.

[0053] Thirdly, this application also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0054] Acquire an ultrasound image set, wherein the ultrasound image set contains at least two ultrasound images of different organs;

[0055] The ultrasound images in the ultrasound image set are analyzed and classified according to organ type;

[0056] Obtain a collection of reports for each categorized ultrasound image;

[0057] Based on the report collection of various ultrasound images, a structured ultrasound report is finally generated;

[0058] The structured ultrasound report includes ultrasound images of various organs and their corresponding reports.

[0059] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0060] Acquire an ultrasound image set, wherein the ultrasound image set contains at least two ultrasound images of different organs;

[0061] The ultrasound images in the ultrasound image set are analyzed and classified according to organ type;

[0062] Obtain a collection of reports for each categorized ultrasound image;

[0063] Based on the report collection of various ultrasound images, a structured ultrasound report is finally generated;

[0064] The structured ultrasound report includes ultrasound images of various organs and their corresponding reports.

[0065] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0066] Acquire an ultrasound image set, wherein the ultrasound image set contains at least two ultrasound images of different organs;

[0067] The ultrasound images in the ultrasound image set are analyzed and classified according to organ type;

[0068] Obtain a collection of reports for each categorized ultrasound image;

[0069] Based on the report collection of various ultrasound images, a structured ultrasound report is finally generated;

[0070] The structured ultrasound report includes ultrasound images of various organs and their corresponding reports.

[0071] The aforementioned method, system, computer equipment, storage medium, and computer program product for retrieval-enhanced multimodal ultrasound data processing and report generation involves: acquiring an ultrasound image set containing at least two ultrasound images of different organs; analyzing the ultrasound images in the set and classifying them according to organ type; obtaining a report set for each classified ultrasound image; and finally generating a structured ultrasound report based on the report sets of each ultrasound image. The structured ultrasound report contains ultrasound images of multiple organs and their corresponding reports. This invention addresses the problems of existing technologies that employ a naive "one-to-multiple" strategy, which forcibly injects a large amount of irrelevant information into the generation process, causing serious semantic pollution; and that if a "multiple-to-multiple" strategy is used, the search results are dominated by normal samples because most images are normal, thus masking a few but crucial anomalies and resulting in reports lacking specificity and diagnostic value. The invention achieves the technical effect of decomposing the complex multi-organ retrieval task into multiple single-organ sub-tasks through organ-level retrieval space decoupling and dynamic routing mechanisms, thereby improving the accuracy and clinical reliability of report generation. Attached Figure Description

[0072] Figure 1 Here is a diagram of the RAG framework in the prior art, where Figure 1 a is a schematic diagram of the organ-to-report correspondence. Figure 1 b is a schematic diagram of a multi-organ ultrasound report;

[0073] Figure 2 This is a schematic diagram of the architecture of the present invention;

[0074] Figure 3 A flowchart of a retrieval-enhanced multimodal ultrasound data processing and report generation method according to an embodiment of the present invention is provided.

[0075] Figure 4 This is an architecture diagram of a retrieval-enhanced multimodal ultrasound data processing and report generation system provided in one embodiment of the present invention;

[0076] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0077] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0078] Multimodal large language models (MLLMs) have shown promise in generating structured reports O=f(V,T) by processing visual input V and cues T. However, their factual accuracy remains limited due to a lack of domain-specific knowledge, making them susceptible to illusions—a particularly problematic issue in clinical applications. Retrieval-enhanced generation (RAG) addresses this problem by building generation upon external knowledge retrieved from domain-relevant corpora.

[0079] This invention addresses the "many-to-many" mapping challenge (i.e., multiple images corresponding to multiple organ descriptions) in fetal ultrasound multimodal report generation by proposing an Organ-aware Routing Mixture Retrieval-Augmented Generation (ORM-RAG) framework. This framework decomposes the complex multi-organ retrieval task into multiple single-organ sub-tasks through organ-level retrieval space decoupling and dynamic routing mechanisms, thereby improving the accuracy and clinical reliability of report generation.

[0080] ORM-RAG comprises two core modules:

[0081] Organ-aware hybrid retrieval module (O-MoR): realizes the classification of multi-organ images and the construction of organ-level retrieval corpus;

[0082] Dynamic Routing Module (DR): Based on rank-aware consistency estimation, it filters high-confidence retrieval results to suppress noise.

[0083] like Figure 2 As shown, it transforms the many-to-many retrieval problem into a mixture of multiple one-to-one retrieval problems, significantly improving the accuracy of cross-organ semantic alignment.

[0084] The implementation process of this invention will now be described in detail.

[0085] Please refer to Figure 3 This application provides a method for retrieval-enhanced multimodal ultrasound data processing and report generation, the method comprising:

[0086] S100, acquire an ultrasound image set, the ultrasound image set containing at least two ultrasound images of different organs.

[0087] The structured ultrasound report includes ultrasound images of various organs and their corresponding reports.

[0088] In this invention, the first step is to acquire a set of ultrasound images that are needed to generate an examination report. This set of ultrasound images contains at least two ultrasound images of different organs. Simply put, during fetal ultrasound screening, multiple organs of the fetus are examined, and each organ produces several ultrasound images. We first need to acquire these ultrasound images to facilitate subsequent image interpretation and report generation.

[0089] S200, Analyze the ultrasound images in the ultrasound image group and classify the ultrasound images according to organ type.

[0090] In this invention, the ultrasound images first need to be classified according to organs, such as which images belong to the heart, which to the face, which to the limbs, etc. This classification facilitates subsequent interpretation of each organ and the generation of the final report.

[0091] Specifically, in this invention, an organ classifier is used to analyze ultrasound images in an ultrasound image set and classify the ultrasound images according to organ type.

[0092] For a set of ultrasound images Using a pre-trained organ classifier Label recognition is performed on each image: M: The number of ultrasound images. The organ classifier can use a CNN-based classification model, such as RESNET and its variants, YOLO series models, etc., which are not limited here.

[0093] S300: Obtain the report set of each ultrasound image after classification.

[0094] After classifying the ultrasound images, we need to obtain a report set of ultrasound images corresponding to each organ. Unlike existing technologies, which typically generate reports in a one-to-many or many-to-many manner—that is, retrieving ultrasound images of one organ from all complete ultrasound reports of all organs (one-to-many), or retrieving ultrasound images of multiple organs from all complete ultrasound reports of all organs (many-to-many)—this approach results in reports with low accuracy and reliability, and is inefficient and wasteful of resources. In this invention, we first classify the ultrasound images by organ, and then retrieve the ultrasound images of each organ from the specific ultrasound reports for that organ, significantly improving both accuracy and efficiency.

[0095] Specifically, step S300 includes the following two steps:

[0096] S310, Construct an organ-specific retrieval corpus.

[0097] In this invention, we need to perform searches within organ-specific ultrasound reports. This requires us to first construct an organ-specific search corpus, which contains report statements corresponding to specific organs. Because RAG technology requires a corpus to query the problem, once the corpus is built, it is used to match ultrasound images with organ-level reports, directly obtaining the report for the organ corresponding to the ultrasound image. Existing corpora contain complete reports, which does not conform to the fetal reporting process. Fetal ultrasound scans multiple organs at once to generate a report, and some sentences in the report correspond to a specific organ section, unlike existing technologies where a single organ corresponds to a complete report.

[0098] Specifically, methods for constructing organ-specific retrieval corpora include:

[0099]

[0100] : Organ-specific retrieval corpus;

[0101] The i-th reference image belonging to organ c in the corpus;

[0102] :and The corresponding organ-specific report text fragment;

[0103] : Total number of image-report pairs of organ c in the retrieved corpus;

[0104] The superscript c indicates an organ-specific subspace, enabling decoupling of many-to-many problems.

[0105] S320, Perform a search query in the retrieval corpus and generate a report set.

[0106] After constructing an organ-specific retrieval corpus, we input the corresponding ultrasound images of the organs into the retrieval corpus for searching and formed a report set of corresponding ultrasound images of the organs.

[0107] Specifically, each organ subset Calculate the overall embedding and retrieve:

[0108]

[0109]

[0110] The average pooled vector of all image embeddings for organ c is used as a query representation to retrieve results from the corpus.

[0111] : The number of images belonging to organ c;

[0112] The i-th image classified as organ c;

[0113] : Organ-specific search function;

[0114] From the corpus The collection of reports retrieved from the database is used to generate the final report;

[0115] E() is the embedding function.

[0116] S400 generates a structured ultrasound report based on the report set of various ultrasound images.

[0117] In the preceding steps, we obtained a report set of ultrasound images corresponding to each organ. In this step, we generate a structured ultrasound report based on the aforementioned report set.

[0118] Specifically, step S400 includes the following steps:

[0119] S410, Perform dynamic routing and confidence assessment on the report set.

[0120] When generating fetal ultrasound reports, directly inputting all retrieved examples into the LLM (Limited Library Model) often introduces redundancy, noise, and even errors. To mitigate this, we propose a dynamic routing mechanism that uses a novel rank-aware consistency estimation module to filter search results. This mechanism selectively routes only high-confidence searches to the LLM for report generation. In simpler terms, we filter the information retrieved from the organ-specific corpus—the report set—preventing the use of all results. Specifically, this is achieved by ranking each search result, calculating the confidence level of each result, and selecting those with confidence levels exceeding a threshold. This is one of the innovative aspects of this invention.

[0121] Specifically, this invention also provides specific methods for dynamically routing and evaluating the confidence level of search results.

[0122] Define the probability of a globally consistent match:

[0123]

[0124] Smoothly weighted label matching probability;

[0125] Binary match indicator ;

[0126] The diagnostic label of the top-ranked search result (0=normal, 1=abnormal);

[0127] : The diagnostic label of the i-th search result;

[0128] Exponentially decaying weights ;

[0129] : Decay coefficient, which controls the rate at which the weight decreases as the ranking decreases;

[0130] γ: Laplace smoothing constant, used to prevent extreme values ​​in probability estimates;

[0131] It is not used directly here. This is because it ignores the smoothness of the order between the retrieved tags.

[0132] Calculate the overall confidence level:

[0133]

[0134] The overall confidence score of organ c search results;

[0135] Global uncertainty entropy based on label consistency;

[0136] Label trend continuity score, which measures the smoothness of the diagnostic label sequence;

[0137] : Control parameters of the Sigmoid function;

[0138] : The balance coefficient of the influence of trend continuity.

[0139] Based on the above method, we can obtain the confidence score of the report corresponding to the organ, which makes it easier to select the appropriate report to generate a structured report based on the confidence score.

[0140] Optionally, in fetal ultrasound, the differences between normal and abnormal images may be visually subtle, making it difficult for standard contrastive learning to accurately distinguish between them. To address this issue, this invention enhances standard contrastive learning using supervised learning of abnormality perception. Specifically, given from... Embedding, i.e. ,have:

[0141]

[0142] in Indicates the first Zhang Image It is a prediction The classification header, This is a label indicating whether an existing organ is normal (0) or abnormal (1). Finally, the optimization objective is defined as follows:

[0143]

[0144] S420: Select a set of reports with a confidence level greater than the confidence threshold and generate a structured ultrasound report.

[0145] In this invention, a confidence threshold can be set according to actual needs, and a set of reports that meet the confidence threshold can be selected to generate a structured report.

[0146] Specifically, reports are generated in the following ways:

[0147]

[0148] The final structured ultrasound report;

[0149] Multimodal large language model generation function;

[0150] : The original set of ultrasound images;

[0151] Prompt word templates to guide report generation format and content;

[0152] : A set of high-confidence search results filtered through dynamic routing;

[0153] Confidence threshold Control the quality of search results through filtering;

[0154] : The overall confidence score of organ c search results.

[0155] To verify the effectiveness of the retrieval-enhanced multimodal ultrasound data processing and report generation method of this invention, a comprehensive dataset called FetusR was collected from Shenzhen Maternity & Child Health Hospital. FetusR contains 15,594 confirmed prenatal cases, totaling 172,851 images, spanning from January 2014 to March 2024. This study has been approved by the local hospital's ethics committee (Approval No.: LLYJ2024-202-089). For details regarding the FetusR dataset, please refer to Appendix B.

[0156] Experimental details

[0157] To obtain organ-specific sentences, this invention uses a rule-based method to extract keywords and refines the results using LLM. For the visual retrieval system, this invention uses ViT-Base as the backbone network and pre-trains it on ImageNet-21K. During training, only the last two Transformer blocks are fine-tuned. This invention uses the Adam optimizer with a learning rate of [missing information]. The batch size is 8, and training lasts for 30 epochs. This invention indexes images in FAISS (a statistical library for similarity search) using 768-dimensional features of the average pooled patch token obtained from the last layer of ViT. Dynamic routing is then applied to retain only confident, informative samples from the retrieval set. For the multimodal base model, this invention uses InternVL as the backbone network. All experiments were implemented using PyTorch and performed on four NVIDIA A100 GPUs. More training details can be found in Appendix D.

[0158] Experimental results

[0159] To thoroughly evaluate the effectiveness of the proposed method, Table 1 presents a performance comparison with several state-of-the-art medical report generation methods, including RAG-based methods (CXR-RePaiR, X-REM, and BiomedCLIP) and non-RAG methods (LLaVA, LLaVA-Med, R2GenGPT, and MicarVLMOE). All models here were fine-tuned on the FetusR dataset proposed in this invention to ensure fair comparisons under the same training settings. The evaluation of this invention considers two main dimensions: textual similarity and diagnostic accuracy.

[0160] Table 1. Model Comparison Results

[0161]

[0162] Table 1 shows that the present invention achieved optimal performance across all metrics. The high degree of consistency with the gold standard report indicates that the proposed method can effectively generate accurate report descriptions, reduce physician workload, and has great potential for clinical application.

[0163] The attributes in Table 1 have the following meanings: Left side: BLEU (B-*), Meteor (M), Rouge-L (RL), CIDEr (CID) indicators; Right side: Abnormal and normal (NM) indicators for the heart (CA), placenta (PA), face (FA), limbs (LA), head (HA), and abdomen (AA). Diagnostic accuracy: Considering that high text similarity does not necessarily guarantee clinical correctness or accurate diagnosis, this invention also introduces diagnostic accuracy to evaluate the model's ability to detect and classify medical abnormalities across multiple clinical categories. Table 1 shows that the SOTA model achieves strong text similarity but poor diagnostic accuracy. This stems from the high visual similarity between multiple ultrasound images, making it difficult to distinguish between normal and abnormal, and data imbalance, particularly the scarcity of abnormal samples for FA, LA, and AA, resulting in near-zero accuracy for these categories. This invention overcomes these challenges by decoupling images by organ and applying organ-level retrieval, thereby improving accuracy for all abnormality types.

[0164] This invention analyzes the performance of different modules in the proposed method. Here, InternVL2-1B is used as the baseline model. Performance is evaluated in three aspects: retrieval accuracy—Rank 1 (R@1) and Rank 5 (R@5), text similarity—BLEU-4 (B-4) and CIDER (CID), and diagnostic accuracy—abnormal accuracy (AB), normal accuracy (NM), and average accuracy (AVG).

[0165] This invention consists of two parts: Hybrid Retrieval (MoR) and Dynamic Routing (DR). Table 2 shows that MoR significantly improves text similarity and diagnostic accuracy, demonstrating the effectiveness of organ-level retrieval. Furthermore, using DR to remove noise or irrelevant reports can improve average diagnostic accuracy by 3.3%.

[0166] Table 2 Ablation Experiment

[0167]

[0168] in conclusion

[0169] This invention pioneers the automated generation of fetal ultrasound reports, supporting multi-organ, multi-view analysis and filling a long-standing gap in the field of medical imaging—previous research focused solely on single-organ imaging techniques such as X-rays and CT scans. This invention constructs an innovative retrieval enhancement framework, ORM-RAG. This framework successfully solves the challenge of accurate matching between multi-organ images and text by decomposing the retrieval task into organ-specific operations and combining them with a dynamic routing mechanism for high-confidence reports. Through integration with a multimodal large language model (MLLM), this invention generates detailed organ reports and achieves state-of-the-art performance, surpassing existing MLLM-based methods. This invention enables scalable, structured reports for complex prenatal ultrasound examinations, possessing significant clinical application value.

[0170] In summary, this application obtains an ultrasound image set containing at least two ultrasound images of different organs; analyzes the ultrasound images in the set and classifies them according to organ type; obtains a report set for each classified ultrasound image; and finally generates a structured ultrasound report based on the report sets of each ultrasound image. The structured ultrasound report contains ultrasound images of multiple organs and their corresponding reports. This addresses the problems of existing technologies that use a naive "one-to-multiple" strategy, which forcibly injects a large amount of irrelevant information into the generation process, causing serious semantic pollution; and that using a "multiple-to-multiple" strategy, where most images are normal, leading to search results dominated by normal samples and masking a few but crucial anomalies, resulting in reports lacking specificity and diagnostic value. This application achieves the technical effect of decomposing the complex multi-organ retrieval task into multiple single-organ sub-tasks through organ-level retrieval space decoupling and dynamic routing mechanisms, thereby improving the accuracy and clinical reliability of report generation.

[0171] At least some steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0172] Based on the same inventive concept, this application also provides a system for implementing the aforementioned retrieval-enhanced multimodal ultrasound data processing and report generation system. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the retrieval-enhanced multimodal ultrasound data processing and report generation system provided below can be found in the limitations of the retrieval-enhanced multimodal ultrasound data processing and report generation method described above, and will not be repeated here.

[0173] In one embodiment, such as Figure 4 As shown, a retrieval-enhanced multimodal ultrasound data processing and report generation system is provided, including:

[0174] The image acquisition module 100 is used to acquire a set of ultrasound images, which includes at least two ultrasound images of different organs.

[0175] The image classification module 200 is used to analyze the ultrasound images in the ultrasound image group and classify the ultrasound images according to organ type.

[0176] The report acquisition module 300 is used to acquire a collection of reports for each classified ultrasound image.

[0177] The report generation module 400 is used to generate a structured ultrasound report based on the report set of various ultrasound images. This structured ultrasound report contains ultrasound images of multiple organs and their corresponding reports.

[0178] In one embodiment, the image classification module 200 is further configured to: analyze the ultrasound images in the ultrasound image group using an organ classifier, and classify the ultrasound images according to organ type;

[0179] For a set of ultrasound images Using a pre-trained organ classifier Label recognition is performed on each image: M: Number of ultrasound images.

[0180] In one implementation, the report acquisition module 300 is further configured to: construct an organ-specific retrieval corpus; perform retrieval queries in the retrieval corpus, and generate a report set. Constructing the organ-specific retrieval corpus includes:

[0181]

[0182] : Organ-specific retrieval corpus;

[0183] The i-th reference image belonging to organ c in the corpus;

[0184] :and The corresponding organ-specific report text fragment;

[0185] : Total number of images of organ c in the retrieved corpus - report pairs.

[0186] In one implementation, the report generation module 300 is further configured to:

[0187] Each organ subset Calculate the overall embedding and retrieve:

[0188]

[0189]

[0190] The average pooled vector of all image embeddings for organ c is used as a query representation to retrieve results from the corpus.

[0191] : The number of images belonging to organ c;

[0192] The i-th image classified as organ c;

[0193] : Organ-specific search function;

[0194] From the corpus The collection of reports retrieved from the database is used to generate the final report;

[0195] E() is the embedding function.

[0196] In one implementation, the report generation module 400 is further configured to: perform dynamic routing and confidence assessment on the report set; select a report set with a confidence level greater than a confidence threshold, and generate a structured ultrasound report.

[0197]

[0198] The final structured ultrasound report;

[0199] Multimodal large language model generation function;

[0200] : The original set of ultrasound images;

[0201] Prompt word templates to guide report generation format and content;

[0202] : A set of high-confidence search results filtered through dynamic routing;

[0203] Confidence threshold Control the quality of search results through filtering;

[0204] : The overall confidence score of organ c search results.

[0205] The modules in the aforementioned retrieval-enhanced multimodal ultrasound data processing and report generation system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.

[0206] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores preset data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a retrieval-enhanced multimodal ultrasound data processing and report generation method.

[0207] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0208] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described retrieval-enhanced multimodal ultrasound data processing and report generation method.

[0209] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described retrieval-enhanced multimodal ultrasound data processing and report generation method.

[0210] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described retrieval-enhanced multimodal ultrasound data processing and report generation method.

[0211] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0212] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0213] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for multimodal ultrasound data processing and report generation based on retrieval enhancement, characterized in that, The method includes: Acquire an ultrasound image set, wherein the ultrasound image set contains at least two ultrasound images of different organs; The ultrasound images in the ultrasound image set are analyzed and classified according to organ type; Obtain a collection of reports for each categorized ultrasound image; Based on the report collection of various ultrasound images, a structured ultrasound report is finally generated; The structured ultrasound report includes ultrasound images of various organs and their corresponding reports. The step of analyzing the ultrasound images in the ultrasound image set and classifying the ultrasound images according to organ type includes: The ultrasound images in the ultrasound image set are analyzed using an organ classifier, and the ultrasound images are classified according to organ type. For a set of ultrasound images Using a pre-trained organ classifier Label recognition is performed on each image: ; M: Number of ultrasound images; The report set of each classified ultrasound image includes: Construct a corpus for organ retrieval; Perform retrieval queries in the retrieval corpus to generate a collection of reports; The constructed organ retrieval corpus includes: ; : Corpus for retrieving organ c; The i-th reference image belonging to organ c in the corpus; :and The corresponding organ-specific report text fragment; : Total number of image-report pairs of organ c in the retrieved corpus; The process of performing a search query in the retrieval corpus and generating a report set includes: Each organ subset Calculate the overall embedding and retrieve: , ; The average pooled vector of all image embeddings for organ c is used as a query representation to retrieve results from the corpus. : The number of images belonging to organ c; The i-th image classified as organ c; : The retrieval function for organ c; From the corpus The collection of reports retrieved from the database is used to generate the final report; E(): Embedding function.

2. The method according to claim 1, characterized in that, Based on the report set of various ultrasound images, a structured ultrasound report is finally generated, including: Dynamic routing and confidence assessment are performed on the aforementioned report set; Select a set of reports with a confidence level greater than the confidence threshold and generate a structured ultrasound report; ; The final structured ultrasound report; Multimodal large language model generation function; : The original set of ultrasound images; Prompt word templates to guide report generation format and content; : A set of high-confidence search results filtered through dynamic routing; : Confidence threshold, ϵ∈[0,1], controls the quality filtering of search results; : The overall confidence score of organ c search results.

3. A retrieval-enhanced multimodal ultrasound data processing and report generation system, characterized in that, The system is used to perform the method according to any one of claims 1-2, the system comprising: An image acquisition module is used to acquire a set of ultrasound images, wherein the set of ultrasound images includes at least two ultrasound images of different organs. An image classification module is used to analyze the ultrasound images in the ultrasound image group and classify the ultrasound images according to organ type; The report acquisition module is used to acquire a collection of reports for each categorized ultrasound image. The report generation module is used to generate a structured ultrasound report based on the report set of various ultrasound images. The structured ultrasound report includes ultrasound images of various organs and their corresponding reports.

4. A computer device, the computer device comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 2.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 2.

6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Large-model medical diagnosis and treatment scheme generation method based on memory and retrieval enhancement

    CN118692663A

  • Medical image report text generation model training method and device

    CN120432068A