Medical report automatic generation method and system based on category guidance

By employing a category-guided approach, utilizing diagnostic category information and a structured clinical knowledge base, local lesion perception and global diagnostic semantic prompts are generated. This addresses the shortcomings in the accuracy and interpretability of report generation in existing technologies, enabling more accurate and transparent medical report generation.

CN121662260APending Publication Date: 2026-03-13JINAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Current medical image report generation technologies lack precise guidance on disease-specific information, fail to effectively integrate local lesion information and clinical context, resulting in a disconnect between report content and actual condition, and lack interpretability and knowledge enhancement mechanisms.

Method used

A category-guided approach is adopted to generate diagnostic category information through a pre-trained disease classification model, which in turn derives local lesion perception and global diagnostic semantic prompts. Combined with a structured clinical knowledge base, this enhances the interpretability and accuracy of the report generation model.

Benefits of technology

It improves the accuracy and reliability of reports, reduces missed diagnoses and misdiagnoses, enhances the clinical relevance and interpretability of reports, adapts to multiple disease entities and complex scenarios, and improves the professionalism and dynamism of report generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121662260A_ABST
    Figure CN121662260A_ABST
Patent Text Reader

Abstract

The invention discloses a medical report automatic generation method and system based on category guidance. The method comprises the following steps: acquiring a medical image; the medical image is analyzed through a pre-trained disease classification model, diagnosis category information corresponding to at least one disease entity is generated, and the diagnosis category information is used for representing the existence state of the disease entity in the image; based on the diagnosis category information, guiding prompt information is derived; and controlling a report generation model to convert the medical image into a structured diagnosis report text by utilizing the guide prompt information. Diagnosis category information is introduced to serve as a guiding mechanism, the model can more accurately pay attention to image features related to diseases, and the possibility of missed diagnosis or misdiagnosis is reduced; in the knowledge enhancement step, the latest medical evidence is fused into the generation process by retrieving and fusing structured clinical knowledge; according to the method, the professional, dynamic and interpretable report generation is realized through a multi-level guidance and fusion mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image analysis technology, specifically to a category-guided method and system for automatically generating medical reports. Background Technology

[0002] Medical image report generation is a crucial step in clinical diagnosis. Traditional methods primarily rely on radiologists manually writing reports, which is not only time-consuming and labor-intensive but also susceptible to subjective factors, leading to inconsistent report quality. With the development of artificial intelligence, existing technologies are beginning to automate report generation using deep learning models, such as encoder-decoder architecture models, which directly extract features from medical images and generate text descriptions. However, these methods have significant drawbacks. First, existing models often lack precise guidance on disease-specific information, potentially causing reports to omit key lesion details or contain fabricated content. Second, most methods rely solely on global image features, failing to effectively integrate local lesion information and clinical context, resulting in reports that are disconnected from the actual patient's condition. Furthermore, existing technologies often cannot dynamically adapt to multiple disease entities or complex clinical scenarios. For example, when multiple positive or indeterminate disease states exist in an image, the model struggles to accurately prioritize and describe these findings.

[0003] Another prominent issue is the lack of interpretability and knowledge enhancement mechanisms in existing report generation systems. The process of model-generated reports is often a black box, making it impossible for physicians to understand the rationale behind their decisions, thus reducing the trustworthiness of the reports. Furthermore, these systems rarely incorporate structured clinical knowledge, such as historical diagnostic reports or medical guidelines, resulting in a lack of clinical depth and consistency in the generated content.

[0004] In summary, existing technologies have shortcomings in terms of guidance mechanisms, multimodal fusion, and interpretability, and there is an urgent need for a more intelligent and transparent automated report generation solution to address these issues. Summary of the Invention

[0005] In order to overcome the shortcomings of the prior art, the purpose of this invention is to provide a method and system for automatically generating medical reports based on category guidance, so as to solve the core problems such as insufficient guidance mechanism and insufficient intelligence in the prior art.

[0006] To solve the above problems, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a category-guided method for automatically generating medical reports, comprising the following steps: Acquiring medical images; The medical image is analyzed by a pre-trained disease classification model to generate diagnostic category information corresponding to at least one disease entity. The diagnostic category information is used to characterize the existence status of the disease entity in the image. Based on the diagnostic category information, guidance and prompts are generated; Using the aforementioned guidance and prompts, a report generation model is controlled to convert the medical images into structured diagnostic report text.

[0007] In some embodiments, the step of deriving guidance prompts based on the diagnostic category information includes: The system generates local lesion perception prompts, which are used to guide the image encoder of the report generation model to focus on local image features related to disease diagnosis.

[0008] In some embodiments, the method for generating the local lesion perception prompt information includes: The disease classification model is subjected to interpretability analysis to generate a category activation map corresponding to the diagnostic category information in order to locate the region of interest in the image; The visual features of the region of interest are extracted, and the aggregated visual features are converted into local lesion perception prompts.

[0009] In some embodiments, when the diagnostic category information indicates that multiple disease entities are in a positive or uncertain state, the step of generating the category activation map includes: The activation regions corresponding to each disease entity are aggregated to form a comprehensive region of interest.

[0010] In some embodiments, the step of deriving guidance prompts based on the diagnostic category information includes: Generate global diagnostic semantic prompts, which are constructed based on predefined semantic tags corresponding to the diagnostic category information and are used as prefixes to be input into the text decoder of the report generation model.

[0011] In some embodiments, the method further includes: Based on the different diagnostic category information, a differentiated text generation strategy is invoked; The differentiated text generation strategy includes at least one of the following: When the category is negative, the strategy tends to exclude abnormal descriptions of the corresponding anatomical structures from the report; When the category is uncertain, the policy triggers the generation of text content containing review suggestions.

[0012] In some embodiments, the method further includes a knowledge enhancement step: In response to the diagnostic category information indicating that a specific disease entity is in a positive state, relevant textual knowledge is retrieved from a structured clinical knowledge base; The retrieved textual knowledge is fused with the visual features of the medical images to enhance the diagnostic basis of the report generation model.

[0013] In some embodiments, the structured clinical knowledge base is constructed in the following manner: From historical diagnostic reports, sentences are selected based on the rule of "disease entity × positive status" and then standardized and deredundant to form a sentence-level knowledge set.

[0014] In some embodiments, the retrieval step includes: mapping positive disease entities to a high-dimensional semantic space, and retrieving the most relevant text knowledge by calculating their semantic similarity to sentences in the knowledge base; The disease classification model is trained using the cross-entropy loss function and with predefined disease labels as supervision signals.

[0015] Secondly, the present invention provides a medical report generation system for implementing the method described above, comprising: Image acquisition module, configured for acquiring medical images; The category analysis module contains the disease classification model and is configured to generate diagnostic category information; The prompt generation module is configured to derive guiding prompt information based on the diagnostic category information, the guiding prompt information including local lesion perception prompt information and global diagnostic semantic prompt information; The report generation module contains the report generation model and is configured to receive the prompt information and the medical image, and output a structured diagnostic report text.

[0016] Compared with the prior art, the present invention has at least the following beneficial effects: 1. First, by introducing diagnostic category information as a guiding mechanism, the model can more accurately focus on disease-related imaging features, reducing the possibility of missed or misdiagnosed diagnoses. Local lesion perception prompts generate category activation maps through interpretability analysis, ensuring the model focuses on key lesion areas and enhancing its ability to identify subtle lesions. Global diagnostic semantic prompts, constructed through semantic tags, provide rich clinical context for report generation, making the report content more aligned with actual diagnostic needs.

[0017] 2. Secondly, the knowledge enhancement step incorporates the latest medical evidence into the generation process by retrieving and fusing structured clinical knowledge, thereby improving the academic rigor and practicality of the report. The adaptive learning mechanism enables the model to continuously evolve from new data and feedback, adapting to the ever-changing clinical environment. Meanwhile, interpretability integration, through saliency plots and confidence scores, makes the generation process transparent, enhancing physicians' trust in the report. Overall, this invention, through multi-level guidance and fusion mechanisms, achieves professionalism, dynamism, and interpretability in report generation, providing reliable support for clinical decision-making.

[0018] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0019] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a category-guided automatic medical report generation method as one embodiment.

[0021] Figure 2 This is a schematic diagram of the framework of a category-guided automatic medical report generation system provided in one embodiment.

[0022] Figure 3 This is a schematic diagram of the complete architecture of the model in one embodiment.

[0023] Figure 4 This is a schematic diagram illustrating the local lesion perception and prompting in one embodiment. Detailed Implementation

[0024] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0026] In the description of this invention, when a specific device is described as being located between a first device and a second device, an intermediary device may or may not be present between the specific device and the first or second device. When a specific device is described as being connected to other devices, the specific device may be directly connected to the other devices without an intermediary device, or it may not be directly connected to the other devices but may have an intermediary device.

[0027] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0028] The applicant discovered: In recent years, artificial intelligence (AI) technology has made significant progress in the field of medical image analysis. For example, AI-based image fusion and image segmentation technologies have achieved pixel-level analysis, effectively improving radiologists' sensitivity in daily diagnosis and significantly shortening image interpretation time. However, despite the excellent performance of these AI algorithms in image analysis, the time-consuming and critical task of writing diagnostic reports still mainly relies on radiologists, becoming a major bottleneck in the clinical workflow. Medical report generation technology has emerged to address this need. It aims to automatically convert visual information from images into structured text reports, covering image presentation, diagnostic impressions, and follow-up recommendations, thereby reducing descriptive omissions, minimizing subjective bias, and improving the standardization of report terminology. This represents an inevitable direction for the development of automated chest X-ray analysis.

[0029] However, in existing technologies, models struggle to establish strong correlations between disease image features and textual descriptions, and their diagnostic reliability is insufficient for long-tailed diseases. Specifically, this includes at least the following two levels of problems: 1. Weak Feature-Description Correlation: Existing methods often result in reports where image findings are disconnected from textual descriptions, or where key lesion features are vaguely described or omitted. This invention aims to deeply embed structured diagnostic category information into the entire process of model encoding (understanding images) and decoding (generating text) through a dual-guided prompting strategy. This forces the model to align global semantics with local lesion information during the learning process, enhancing the clinical relevance between report content and image evidence.

[0030] 2. Poor Diagnostic Performance for Long-Tail Diseases: To address the problem of weak model discrimination and inaccurate description caused by the scarcity of rare disease samples in the training data, this invention designs a cross-modal knowledge enhancement module. This module can dynamically retrieve and integrate relevant knowledge from external clinical knowledge bases, providing the model with additional diagnostic evidence, especially for rare or atypical cases. This effectively improves the model's ability to identify and describe long-tail diseases, enhancing the robustness and generalization of the overall system.

[0031] In view of this, refer to Figure 1 In a first aspect, embodiments of this application propose a category-guided method for automatically generating medical reports, comprising the following steps: Acquire medical images of the target patient; The medical image is analyzed by a pre-trained disease classification model to generate diagnostic category information corresponding to at least one disease entity. The diagnostic category information is used to characterize the existence status of the disease entity in the image. Based on the diagnostic category information, guidance and prompting information is derived through interpretability analysis. The guidance and prompting information includes local lesion perception prompts and global diagnostic semantic prompts. Using the aforementioned guidance and prompts, a report generation model is controlled to convert the medical images into structured diagnostic report text.

[0032] The steps for generating the guidance prompt information include: generating a category activation map based on the diagnostic category information to locate the region of interest in the image, and extracting visual features to convert them into local lesion perception prompts; at the same time, constructing a global diagnostic semantic prompt based on predefined semantic labels as a prefix input to the report generation model.

[0033] It should be noted that the diagnostic category information is the output generated after analyzing medical images using a pre-trained disease classification model. This information is used to characterize the presence status of specific disease entities in the images. Specifically, for each disease entity, the model outputs a probability distribution or classification label indicating whether the entity is blank, positive (lesion present), negative (no lesion), or uncertain. This representation is based on features extracted from the images by a deep learning model, implemented through a softmax layer or a sigmoid activation function, thereby quantifying the probability of disease presence. For example, for the disease entity of a lung nodule, the diagnostic category information may represent the presence of the nodule with a high probability value and reflect its certainty through a confidence score, thus providing structured and interpretable input for subsequent steps.

[0034] The core innovation of this embodiment lies in deriving guiding prompts through interpretability analysis. Preferably, the interpretability analysis employs Gradient Weighted Class Activation Mapping (Grad-CAM++) technology to backpropagate the disease classification model, generating category activation maps corresponding to the diagnostic category information. These activation maps, by calculating the weighted sum of feature map gradients, highlight regions of interest in the image related to specific disease entities, thereby locating lesions. The activation maps for each disease are weighted and fused according to their confidence levels to obtain a fused activation map for each sample. Based on the fused activation map, the system uses a lesion region localization method based on local maxima detection to locate the regions of interest in the disease classification model and extract visual features from these regions, such as through weighted feature pooling of feature maps from a convolutional neural network, and aggregates these features to convert them into local lesion perception prompts.

[0035] Meanwhile, global diagnostic semantic prompts are constructed using a predefined semantic tag library, mapping diagnostic category information to a high-dimensional semantic space and generating semantic vectors as prefix inputs. This process ensures that the prompt information originates from imaging data while incorporating clinical semantics, enhancing the transparency and controllability of the generation process.

[0036] More specifically, the primary role of local lesion perception cues is to guide the image encoder of the report generation model to focus on posterior image features relevant to disease diagnosis. After applying spatially weighted pooling to the convolutional feature map to obtain region features, each region feature is projected onto the latent embedding space through a linear transformation, generating lesion perception cues that better align with the encoder's semantics. These lesion perception cues, as prefixes to the encoder, are appended to the encoder's input sequence to form an enhanced visual embedding sequence, guiding the encoder to focus on clinically salient regions during the visual-text representation learning process.

[0037] Global diagnostic semantic cues provide clinical context, guiding the report generation model's text decoder to produce report content that conforms to medical standards. In the decoder, these cues serve as initial tags or prefix sequences, influencing subsequent word predictions through an autoregressive generation process. Preferably, in a Transformer-based decoder, the global cues interact with the encoder output via a cross-attention mechanism, ensuring that the generated text is semantically consistent with the diagnostic category. Its working principle involves transforming disease states into understandable textual clues through semantic embedding, such as mapping a "positive" state to a description like "suspected malignant tumor," thereby reducing the risk of generating irrelevant or erroneous content and improving the report's structure.

[0038] The process of generating reports using guided prompts involves multimodal fusion and control mechanisms. The prompts act as a bridge connecting the image encoding and text generation stages. In the encoding stage, local lesion perception prompts optimize image feature extraction, making the encoder output more lesion-specific feature representations. In the decoding stage, global diagnostic semantic prompts initialize the generation sequence, and conditional generation mechanisms constrain the text output. The model learns how to respond to these prompts through end-to-end training, preferably using cross-entropy loss and negative log-likelihood loss to ensure semantic and structural consistency between the report text and the prompts. This guidance mechanism enables the model to dynamically adapt to different disease states; for example, when the diagnostic category indicates an uncertain state, the prompts trigger the generation of text containing follow-up recommendations, thereby improving the clinical usability and accuracy of the report.

[0039] Furthermore, the steps for generating guidance and prompts include: First, based on diagnostic category information, Grad-CAM++ is used to generate category activation maps. A heatmap is generated by calculating the product of the feature map gradient and weights to locate regions of interest. Second, the activation maps for each disease are weighted and fused according to their confidence levels to obtain a fused activation map for each sample. Next, visual features of the regions of interest are extracted through weighted feature pooling and aggregated into local prompt vectors using fully connected layers. Finally, global prompts are generated by querying a predefined semantic tag library and using a word embedding model such as SentenceBERT to generate semantic vectors, which are then input into the decoder as prefix sequences. This processing ensures that the prompts accurately capture image details and are clinically relevant, thus providing reliable guidance for report generation.

[0040] As one implementation method, the step of deriving guidance prompts based on the diagnostic category information includes: The system generates local lesion perception prompts, which are used to guide the image encoder of the report generation model to focus on local image features related to disease diagnosis.

[0041] More specifically, local lesion perception cues guide the image encoder of the report generation model to focus on local image features relevant to disease diagnosis by providing spatial attention weights. Specifically, these cues originate from diagnostic category information, are generated through interpretability analysis, and represent the salience of key regions in the image as vectors or heatmaps. In the image encoder, typically based on a convolutional neural network (CNN) or Transformer architecture, the local cues are integrated into the attention mechanism. Optionally, before the image features extracted by the encoder are injected into the decoder through a cross-attention mechanism, the cues serve as an additional prefix, dynamically adjusting the weight distribution of the feature extraction process, causing the encoder to prioritize highly salient regions, such as lesion edges, texture anomalies, or areas of density variation. This guidance mechanism strengthens the representation of relevant features while suppressing background or irrelevant regions, thereby improving the discriminativeness of the encoder's output feature map and ensuring that subsequent report generation is based on accurate visual evidence.

[0042] Optionally, the method for generating the local lesion perception prompt information includes: The disease classification model is subjected to interpretability analysis to generate a category activation map corresponding to the diagnostic category information in order to locate the region of interest in the image; The visual features of the region of interest are extracted, and the aggregated visual features are converted into local lesion perception prompts.

[0043] The interpretability analysis employs Gradient Weighted Class Activation Mapping (Grad-CAM++) to backpropagate a pre-trained disease classification model, generating class activation maps corresponding to diagnostic category information. This process first calculates the feature maps of the target category (e.g., specific disease entities) through the convolutional layers, obtains weights through global average pooling, and generates a heatmap by weighted summation, highlighting disease-related regions of interest in the image. These activation maps are directly associated with the regions of interest because they visualize the model's decision-making process and locate lesions, such as lung nodules or hemorrhage areas. Subsequently, the system extracts visual features of the regions of interest through weighted feature pooling, for example, cropping and pooling specific regions from the CNN feature maps to obtain fixed-dimensional feature vectors. When aggregating these features, a fully connected layer is used for weighted fusion, transforming multi-region features into a compact vector of local lesion-aware cues. This vector encodes the spatial and semantic information of the lesions and serves as a guiding signal input to the report generation model, ensuring that the cues retain both local details and generalization.

[0044] Optionally, when the diagnostic category information indicates that multiple disease entities are in a positive or uncertain state, the step of generating the category activation map includes: The activation regions corresponding to each disease entity are aggregated to form a comprehensive region of interest.

[0045] It should be noted that when diagnostic category information indicates multiple disease entities as positive or indeterminate, forming a comprehensive area of ​​interest is to avoid a dispersion of focus and ensure that the report generation model comprehensively covers all relevant lesions without overlooking potential lesions. A positive status indicates the presence of disease, while an indeterminate status suggests the need for further investigation; both should be prioritized.

[0046] The aggregation method involves fusing the category activation maps corresponding to each disease entity: First, weighted fusion is performed based on classifier prediction confidence, merging multiple activation maps to generate a comprehensive heatmap that identifies all potential regions of interest; second, local maxima detection removes redundant regions, highlighting key areas; finally, aggregated visual features are extracted based on the comprehensive heatmap, such as using weighted feature pooling to capture features of lesions of different sizes, and further compressed into a unified cue vector through linear transformation. This aggregation ensures the robustness and completeness of the cue information across multiple disease scenarios, guiding the model to generate more comprehensive and accurate report content.

[0047] As one implementation method, the step of deriving guidance prompts based on the diagnostic category information includes: Generate global diagnostic semantic prompts, which are constructed based on predefined semantic tags corresponding to the diagnostic category information and are used as prefixes to be input into the text decoder of the report generation model.

[0048] The generation of global diagnostic semantic prompts relies on a predefined semantic tag library, which is associated with specific disease entities, such as "positive pulmonary nodules" or "uncertain pleural effusion". Diagnostic category information (such as disease status) is converted into corresponding semantic tags through mapping functions, such as using lookup tables or embedding models to map numerical categories to natural language descriptions.

[0049] When constructing global diagnostic semantic prompts, the system employs word embedding technology to transform semantic tags into high-dimensional vector representations. These vectors capture the semantic relationships between tags, and optionally, cosine similarity is used to measure the correlation between different disease descriptions. The generation process includes: first, selecting relevant semantic tags based on diagnostic category information; second, serializing these tags into a series of word-level text prefixes, such as "[POS],[UNC],...,[POS]"; and finally, encoding these text prefixes into fixed-dimensional vectors as global prompts.

[0050] When used as a prefix input to the text decoder of the report generation model, global diagnostic semantic hints guide the initialization of the text sequence through a conditional generation mechanism. In the Transformer-based decoder, the prefix vector is pre-loaded into the decoder's input sequence as an initial label. During autoregressive generation, the decoder predicts subsequent words based on this prefix, propagating the prefix semantics to the entire output sequence through attention weights. For example, the prefix "positive lung nodule" will cause the decoder to prioritize generating terms related to the nodule description, such as "rough edges" or "follow-up recommended," thus ensuring that the report content is consistent with the diagnostic status. This working principle leverages the contextual learning capabilities of language models, reducing the probability of generating irrelevant content through prefix constraints, thereby improving the relevance and reliability of the report.

[0051] The integration of global diagnostic semantic cues enhances the multimodal alignment capabilities of the report generation model. By combining image-derived information with textual semantics, the model can better understand clinical intent and reduce hallucination generation. Simultaneously, the cues, acting as prefixes, serve as soft constraints, allowing the decoder to adhere to medical logic while maintaining flexibility.

[0052] Furthermore, the method also includes: Based on the different diagnostic category information, a differentiated text generation strategy is invoked; The diagnostic category information includes blank, positive (lesion present), negative (no lesion), or uncertain states. Each state serves as an input signal, influencing the text decoder's word selection, sentence structure, and content prioritization. The strategy is implemented through predefined rules or machine learning models (such as attention-weighted conditional generation), where the diagnostic category information is encoded as a control vector and integrated into the intersecting attention layer or prefix sequence of the Transformer decoder.

[0053] The differentiated text generation strategy includes at least one of the following: When the category is blank, the diagnostic category information indicates that no disease entity or related findings were detected. The strategy tends to generate neutral or standardized report content, avoiding the introduction of any abnormal descriptions. The report generation model will prioritize outputting general observation statements, such as "No obvious abnormalities were found on imaging examination," or directly omit detailed discussions of specific anatomical structures to maintain the conciseness and objectivity of the report. The working principle is to train the model on a large amount of training data to understand additional diagnostic category information ([POS], [NEG], [UNC], [BLA]), thereby ensuring that the output only reflects normal or irrelevant findings.

[0054] When the category is positive, the diagnostic category information confirms the presence of a lesion. The strategy focuses on generating a detailed and structured description of the abnormality, including the lesion's location, shape, size, and characteristics. The report generation model activates a lexicon related to the lesion, enhancing attention weights to key areas, and outputting information such as "A solid nodule with irregular margins and a diameter of approximately 1.2 cm is seen in the upper lobe of the right lung." The working principle utilizes diagnostic category information as a reinforcement signal, prioritizing terms consistent with the positive status during decoding, and ensuring the description conforms to medical standards through contextual learning, thereby providing a highly informative diagnostic opinion.

[0055] When the category is negative, the diagnostic category information indicates no lesion. The strategy aims to exclude any abnormal descriptions of the corresponding anatomical structures to avoid misleading content. The report generation model actively filters out lesion-related words, for example, by using a masking mechanism to skip abnormal terms in the generated sequence and instead output confirmatory statements such as "no nodules or exudates were found in the lungs." The working principle is based on negative sample learning. During training, the model is optimized to identify negative states, and during inference, conditional probabilities are adjusted to ensure that the output only contains normal or negative findings, enhancing the reliability of the report.

[0056] When the category is uncertain, the policy triggers the generation of text content containing review suggestions.

[0057] As one implementation, the method further includes a knowledge enhancement step: In response to the diagnostic category information indicating that a specific disease entity is in a positive state, relevant textual knowledge is retrieved from a structured clinical knowledge base; The retrieved textual knowledge is fused with the visual features of the medical images to enhance the diagnostic basis of the report generation model. This addresses the problem of reports lacking medical evidence support, improving the academic depth and diagnostic basis of the reports through knowledge retrieval and fusion.

[0058] Preferably, the structured clinical knowledge base is constructed in the following manner: From historical diagnostic reports, sentences are selected based on the rule of "disease entity × positive status," and then standardized and deredundanted to form a sentence-level knowledge set. The structured clinical knowledge base is constructed by selecting and standardizing sentences from historical diagnostic reports, for example, based on the "disease entity × positive status" rule. This solves the problem of low efficiency in knowledge base construction, ensuring high quality and usability through automated selection and processing.

[0059] Preferably, the retrieval step includes: mapping positive disease entities to a high-dimensional semantic space, and retrieving the most relevant text knowledge by calculating their semantic similarity with sentences in the knowledge base; The disease classification model is trained using a cross-entropy loss function and predefined disease labels as supervision signals. This addresses the problem of inaccurate knowledge retrieval by ensuring that the retrieved content is highly relevant to the current case through semantic matching, thus enhancing the fusion effect.

[0060] It should be noted that the knowledge enhancement step improves the diagnostic basis and academic depth of the medical report generation model by integrating a structured clinical knowledge base. The core of this step lies in responding to positive states in diagnostic category information, automatically retrieving relevant textual knowledge, and performing multimodal fusion with image visual features, thereby enhancing the accuracy and clinical relevance of the report content.

[0061] Specifically, the fusion of retrieved textual knowledge and medical image visual features employs a cross-modal attention mechanism. First, textual knowledge is converted into high-dimensional semantic vectors through a pre-trained word embedding model, capturing its semantic information. Simultaneously, visual features are extracted from the image encoder output and represented as spatial feature maps. The fusion process is performed through a cross-attention layer: visual features serve as the query, and text vectors as the key and value. Attention weights are calculated, allowing the model to dynamically focus on the correlation between the textual description and the image region. For example, for the positive disease entity "pulmonary nodule," retrieved textual knowledge such as "rough nodule edges" will enhance the representation of the corresponding edge region in the visual features through attention weights. The fused features are then integrated through fully connected layers or gating mechanisms to generate an enhanced multimodal representation, which is input into the report generation model. This mechanism ensures that textual knowledge complements and corrects visual evidence, reduces model illusions, and improves the logical coherence of the report.

[0062] Furthermore, the establishment of the "disease entity × positive status" rule is based on the structured analysis of historical diagnostic reports, aiming to automatically filter high-quality sentences. Rule construction first defines a disease entity database (e.g., "pulmonary nodule," "pleural effusion") and status labels (positive status), then identifies report sentences using pattern matching or keyword extraction techniques. Specifically, natural language processing tools are used to scan historical reports, filtering sentences that simultaneously contain disease entity names and positive indicator words (e.g., "visible," "existing," or "confirmed"). For example, the sentence "A solid nodule is visible in the upper lobe of the right lung" is captured because it matches the "pulmonary nodule × positive status" pattern. After filtering, the sentences undergo standardization (e.g., removing stop words, standardizing terminology) and redundancy removal (e.g., clustering based on semantic similarity) to form a sentence-level knowledge set. This rule-based filtering ensures the relevance and consistency of the knowledge base, solving the problem of low efficiency in manual construction.

[0063] The disease classification model is trained using a cross-entropy loss function with predefined disease labels as supervision signals. The cross-entropy loss function calculates the difference between the model's predicted probability distribution and the true labels. During training, the model optimizes the loss through backpropagation and gradient descent to make the predicted values ​​approximate the true distribution. The supervision signals come from predefined disease labels, which are based on clinical annotations (such as disease states confirmed by radiologists) and serve as the gold standard in the training data. For example, for an image sample, the label might be "positive lung nodule," and the model learns to distinguish between positive and negative states by minimizing the cross-entropy loss. This training mechanism ensures the robustness and generalization ability of the disease classification model, providing reliable diagnostic category information for knowledge augmentation steps.

[0064] In one implementation, the method further includes an interpretability and confidence integration step: adding an interpretability module to generate a saliency map to highlight key areas in the image that affect diagnostic decisions, and embedding the saliency map into the diagnostic report as visual evidence; outputting a confidence score for each diagnostic category, calculating uncertainty based on a probability distribution, and providing a note in the report; wherein the guidance and prompting information includes explanatory prompts, provides decision-making logic, and extracts supporting literature fragments from the knowledge base through an evidence retrieval function and cites them in the report.

[0065] By adding an interpretability module to the current system, such as integrating a saliency map generator after the disease classification model, key regions in the image that influence diagnostic decisions are highlighted, and these maps are embedded in the report as visual evidence. Simultaneously, the report generation model can output confidence scores; for each diagnostic category, the model calculates the probability distribution and indicates uncertainty in the report as a footnote; for example, a low-confidence category can trigger the generation of text "further investigation recommended." The prompt generation module can be expanded into a multi-level prompt system, where global prompts control the report structure, local prompts refine the description, and explanatory prompts provide the decision-making logic. The knowledge enhancement module can integrate evidence retrieval functionality, extracting literature snippets supporting the diagnosis from the knowledge base and directly citing them in the report. The entire solution achieves interpretability through end-to-end training, uses an attention mechanism to associate output and input features, and introduces consistency checks to ensure that the interpretation matches the report content. This not only solves the credibility problem of black-box models but also makes this technology more reliable in clinical deployment.

[0066] Reference Figure 2 Secondly, the present invention provides a medical report generation system for implementing the method as described in the above embodiments, comprising: Image acquisition module, configured for acquiring medical images; The category analysis module contains the disease classification model and is configured to generate diagnostic category information; The prompt generation module is configured to derive guiding prompt information based on the diagnostic category information, the guiding prompt information including local lesion perception prompt information and global diagnostic semantic prompt information; The report generation module contains the report generation model and is configured to receive the prompt information and the medical image, and output a structured diagnostic report text.

[0067] By coordinating the image acquisition module, category analysis module, prompt generation module, and report generation module, the method described in the above embodiments is implemented, which solves the problem of low system integration. Modular design improves processing smoothness and enhances overall performance.

[0068] The following example illustrates the specific technical aspects: Reference Figure 3 , Figure 3 To provide the complete model architecture of the medical report generation system presented in this embodiment, we propose a dual-guidance prompting strategy to capture the similarity of reports within the same diagnostic category. Specifically, our framework includes a disease classification model that categorizes each disease into four classes (Blank, Positive, Negative, Uncertain). These classification labels are predefined disease labels from the reports. Next, we train the disease classification model based on these disease labels using a cross-entropy loss function.

[0069] The global awareness cue output by the disease classifier serves as a prefix for the decoder, namely four different token cuees: [BLA], [POS], [NEG], and [UNC]. These cuees are added to the decoder's vocabulary, providing explicit global disease semantics and uncertainty priors to ensure that the reported content is consistent with the diagnostic conclusion.

[0070] Meanwhile, to enhance the alignment between visual representation and disease semantics, we introduced local lesion perception cues on the encoder side, such as... Figure 4 Specifically, the disease classifier uses gradcam++ to generate class activation maps that indicate the regions most relevant to each disease prediction.

[0071] For samples with multiple positive or uncertain predictions, multi-disease aggregation is further performed to retain all potentially abnormal areas. Next, the regions with the highest activation in the multi-disease aggregated heatmap are extracted and located based on the lesion region. Subsequently, the visual features corresponding to these regions are aggregated and transformed into a set of learnable lesion perception cues through a lightweight fully connected layer. These cues are then injected as additional input into the image encoder, guiding the encoding process to focus on clinically significant local areas.

[0072] By introducing these local cues, the encoder is able to more sensitively capture subtle disease clues, achieving finer-grained visual-text alignment. Through a complementary global-local design, the dual guidance and cues strategy effectively alleviates the challenges of lack of global consistency and neglect of key details in medical report generation.

[0073] In terms of clinical rule embedding, each of the four states plays an important role.

[0074] Blank states tend to mask irrelevant anatomical descriptions and reduce redundant information; A standardized medical terminology database for positive state triggers ensures the professionalism and accuracy of lesion descriptions; A negative result tends to exclude suspicious areas, reducing the risk of overdiagnosis; For uncertain cases, reasonable suggestions for follow-up examinations are given based on the degree of image blurriness, providing direction for subsequent diagnosis.

[0075] The four-state label is embedded as a semantic prefix in the decoding process, which transforms the diagnostic thinking of radiologists into computable rules and realizes the explicit expression of diagnostic logic.

[0076] Our method establishes a robust dual-guidance prompting strategy. Through the strong association between state and description, it enables precise control of the text generation range based on the diagnostic status of each entity, ensuring a high degree of consistency between the report content and the images, thereby improving the correlation between the disease and the text.

[0077] Furthermore, addressing the issue of poor model performance in handling rare diseases within the long-tail disease problem scenario, this study proposes cross-modal knowledge enhancement centered on diagnostic needs to improve the ability to acquire and apply medical knowledge. First, a structured knowledge base is constructed based on positive diseases. From the training data, key reports are selected according to the rule of "disease entity × positive state," and the reports are segmented into sentences. High-frequency sentences are downsampled, and the maximum frequency of each sentence is controlled to form a standardized sentence-level knowledge base. Next, in the retrieval mechanism, the retrieval of the corresponding knowledge base partition is activated only when the diagnostic classifier outputs a disease entity as a positive state.

[0078] Finally, we use cosine similarity to map positive disease entities to a high-dimensional semantic space, thereby identifying equivalent expressions of different terms. In this way, we can retrieve report sentences with similar content and provide them as reference information to the decoder for generating the final diagnostic report.

[0079] The system was specifically designed to address the distribution issues of medical image data. When dealing with rare diseases or complex symptoms, the positive state of entities is used to increase the likelihood of retrieving the corresponding knowledge sub-base, reducing the global similarity ranking's bias towards high-frequency diseases and negative samples, thus improving the model's performance in scenarios with imbalanced disease distribution.

[0080] Therefore, the key technical point of this embodiment is: 1. A report generation mechanism employing a dual-guidance prompting strategy: This mechanism creatively uses structured diagnostic information as a guiding signal, simultaneously injecting it into both the image encoding and text decoding stages. This not only ensures a high degree of relevance between the generated report and the image content, but more importantly, it shifts the model's focus from "generating fluent text" to "making an accurate diagnosis," thereby solving the core problem of weak feature-description correlation. Its key lies in achieving a synergistic alignment between global diagnostic intent and local lesion features.

[0081] 2. Cross-modal knowledge enhancement method for long-tail diseases: A dynamic cross-modal knowledge enhancement module was designed. When the model encounters rare or uncertain signs, it can retrieve relevant knowledge from a structured clinical knowledge base and represent it as a fused disease clue enhancement feature. This is equivalent to equipping the model with a "clinical decision support system," effectively compensating for the lack of prior knowledge in data-driven models for long-tail diseases and improving the reliability of diagnosis.

[0082] In summary, compared with the prior art, the above embodiments have at least the following technical advantages: 1. First, by introducing diagnostic category information as a guiding mechanism, the model can more accurately focus on disease-related imaging features, reducing the possibility of missed or misdiagnosed diagnoses. Local lesion perception prompts generate category activation maps through interpretability analysis, ensuring the model focuses on key lesion areas and enhancing its ability to identify subtle lesions. Global diagnostic semantic prompts, constructed through semantic tags, provide rich clinical context for report generation, making the report content more aligned with actual diagnostic needs.

[0083] 2. Secondly, the knowledge enhancement step incorporates the latest medical evidence into the generation process by retrieving and fusing structured clinical knowledge, thereby improving the academic rigor and practicality of the report. The adaptive learning mechanism enables the model to continuously evolve from new data and feedback, adapting to the ever-changing clinical environment. Meanwhile, interpretability integration, through saliency plots and confidence scores, makes the generation process transparent, enhancing physicians' trust in the report. Overall, this invention, through multi-level guidance and fusion mechanisms, achieves professionalism, dynamism, and interpretability in report generation, providing reliable support for clinical decision-making.

[0084] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.

Claims

1. A method for automatically generating medical reports based on category guidance, characterized in that, Includes the following steps: Acquiring medical images; The medical image is analyzed by a pre-trained disease classification model to generate diagnostic category information corresponding to at least one disease entity. The diagnostic category information is used to characterize the existence status of the disease entity in the image. Based on the diagnostic category information, guiding prompts are generated; Using the guidance and prompting information, a report generation model is controlled to convert the medical image into a structured diagnostic report text.

2. The method according to claim 1, characterized in that, The step of deriving guidance prompts based on the diagnostic category information includes: The system generates local lesion perception prompts, which are used to guide the image encoder of the report generation model to focus on local image features related to disease diagnosis.

3. The method according to claim 2, characterized in that, The method for generating the local lesion perception prompt information includes: The disease classification model is subjected to interpretability analysis to generate a category activation map corresponding to the diagnostic category information in order to locate the region of interest in the image; Visual features of the region of interest are extracted, and the aggregated visual features are converted into local lesion perception prompts.

4. The method according to claim 3, characterized in that, When the diagnostic category information indicates that multiple disease entities are in a positive or uncertain state, the step of generating the category activation map includes: The activation regions corresponding to each disease entity are aggregated to form a comprehensive region of interest.

5. The method according to any one of claims 1 to 4, characterized in that, The step of deriving guidance prompts based on the diagnostic category information includes: Generate global diagnostic semantic prompts, which are constructed based on predefined semantic tags corresponding to the diagnostic category information and are used as prefixes to be input into the text decoder of the report generation model.

6. The method according to claim 5, characterized in that, The method further includes: Based on the different diagnostic category information, a differentiated text generation strategy is invoked; The differentiated text generation strategy includes at least one of the following: When the category is negative, the strategy tends to exclude abnormal descriptions of the corresponding anatomical structures from the report; When the category is uncertain, the policy triggers the generation of text content containing review suggestions.

7. The method according to claim 6, characterized in that, The method also includes a knowledge enhancement step: In response to the diagnostic category information indicating that a specific disease entity is in a positive state, relevant textual knowledge is retrieved from a structured clinical knowledge base; The retrieved textual knowledge is fused with the visual features of the medical images to enhance the diagnostic basis of the report generation model.

8. The method according to claim 7, characterized in that, The structured clinical knowledge base is constructed in the following ways: From historical diagnostic reports, sentences are selected based on the rule of "disease entity × positive status" and then standardized and deredundant are processed to form a sentence-level knowledge set.

9. The method according to claim 7, characterized in that, The retrieval steps include: mapping positive disease entities to a high-dimensional semantic space, and retrieving the most relevant text knowledge by calculating their semantic similarity to sentences in the knowledge base; The disease classification model is trained using the cross-entropy loss function and with predefined disease labels as supervision signals.

10. A medical report generation system for implementing the method according to any one of claims 1-9, characterized in that, include: Image acquisition module, configured for acquiring medical images; The category analysis module contains the disease classification model and is configured to generate diagnostic category information; The prompt generation module is configured to derive guiding prompt information based on the diagnostic category information, the guiding prompt information including local lesion perception prompt information and global diagnostic semantic prompt information; The report generation module contains the report generation model and is configured to receive the prompt information and the medical image, and output a structured diagnostic report text.