Radiology report automatic generation method, system and equipment based on patient-specific priori knowledge and medium

By integrating patient-specific prior knowledge through spatiotemporal fusion networks and progressive fusion networks based on prior perception, this approach addresses the issues of insufficient utilization of patient-specific information and inadequate disease evolution modeling in existing technologies. This improves the personalization and accuracy of radiology reports, reduces the risk of generating hallucinations, and enhances the clinical credibility and practicality of radiology reports.

CN120954609APending Publication Date: 2025-11-14XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511078531.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing methods for generating radiology reports neglect patient-specific prior knowledge, resulting in reports lacking personalized content, failing to accurately reflect the patient's specific condition, and having insufficient modeling of disease evolution, which can easily lead to generated hallucinations and affect the accuracy and reliability of diagnosis.

Method used

By integrating patients' past medical images through a spatiotemporal fusion network, a disease evolution trajectory is constructed. Combined with a progressive fusion network based on prior perception, patient-specific prior knowledge and hierarchical visual features are gradually integrated. A two-stage training strategy is adopted to improve the accuracy of medical image and text retrieval and the quality of radiology report generation.

Benefits of technology

By effectively integrating patient-specific clinical context information, radiology reports can be personalized and accurate, the risk of generating hallucinations can be reduced, the model's ability to understand disease progression can be enhanced, and the clinical credibility and usability of radiology reports can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954609A_ABST
    Figure CN120954609A_ABST
Patent Text Reader

Abstract

A radiology report automatic generation method, system, device and medium based on patient specific priori knowledge uses a special token to represent missing patient clinical context information, so that a text encoder can process complete and incomplete clinical context input in a unified manner, and then robust clinical context features are extracted; the method comprises the following steps: constructing a space-time fusion network STF, integrating previous medical images of a patient, establishing a difference mapping relation between a current image and a historical image for modeling an evolutionary process of a disease, and extracting space-time visual features with time dependence; establishing an attention-enhanced hierarchical fusion network for fusing multiple layers of hidden states in a visual encoder so as to extract multiple layers of hierarchical visual features with rich semantics; a prior perception progressive fusion network is introduced, and patient specific prior knowledge and hierarchical visual features are gradually fused in a coarse-to-fine mode to generate multi-modal features facing radiology report generation; the method comprises the following steps: designing a two-stage training strategy: in the first stage, aiming at image-text alignment, improving the accuracy of medical image-text retrieval; and in the second stage, the generation of the radiology report is taken as a target, a text decoder is optimized, and the performance of the generated radiology report in the aspects of clinical semantic accuracy and language expression quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic text generation technology, and in particular to a method, system, device and medium for automatically generating radiology reports based on patient-specific prior knowledge. Background Technology

[0002] Chest X-ray (CXR) is one of the most commonly used medical imaging examinations, widely used in the screening and diagnosis of diseases such as lung infection, pleural effusion, pneumothorax, cardiomegaly, and fractures. Radiology reports, as an important output form of image interpretation, play a key role in clinical diagnosis and treatment decisions. Traditional manual report writing methods have two main problems: (1) the radiology report generation process is time-consuming and cumbersome, which can easily cause diagnostic delays, especially in high-intensity work scenarios; (2) different radiologists have different writing styles, and radiology reports lack standardization and consistency. Therefore, radiology report generation (RRG) based on artificial intelligence (AI) has become a research hotspot. This technology automatically interprets images and generates a draft radiology report through a model, which is expected to improve the efficiency and consistency of radiology report writing. However, most existing technologies adopt an end-to-end image-to-text modeling approach, generating radiology reports based only on the current images, ignoring the patient's clinical context and previous images, making it difficult for the model to achieve personalized diagnosis, and the quality of radiology reports still cannot meet the actual clinical needs.

[0003] To improve the clinical accuracy of generated radiology reports, researchers have proposed various technical approaches, including knowledge graphs, retrieval enhancement mechanisms, reinforcement learning, large language model frameworks, and disease tags. In the knowledge graph-based RRG method, M2KT constructs a knowledge base from radiology reports and automatically distills and preserves medical knowledge, achieving semantic alignment and disease classification, thus improving the linguistic quality and clinical accuracy of radiology reports. GSKET addresses the data bias problem in medical imaging by introducing prior knowledge, injecting radiology reports with similar disease tags as specific knowledge into the model, enhancing its generality and specificity, thereby significantly improving performance. DCL further constructs a dynamic graph to enrich knowledge representation and enhance the expressive power of image features. In retrieval enhancement methods, FSE proposes a fact sequence enhancement strategy, enabling the model to focus on cross-modal alignment between fact sequences composed of images and clinical keywords, while simultaneously retrieving similar historical cases to guide the generation of higher-quality radiology reports. In reinforcement learning, Delbrouck et al. designed a semantic reward-based mechanism to improve the clinical accuracy of radiology reports and used Self-critical Sequence Training (SCST) to transform non-differentiable semantic indicators into differentiable functions, optimizing the training process. Within the framework of large language models, Yan et al. first extracted clinical keyword sequences and then combined them with large language models (such as ChatGPT and LLaMA) for scene learning to generate the final radiology report. Liu et al. achieved cross-modal alignment and generation between medical images and radiology reports by calling MiniGPT-4 twice. In disease-label-based methods, FMVP used a CheXbert pre-trained model to label images with 14 common diseases (such as cardiomegaly, edema, and lung infection), and combined physician knowledge to construct medical concepts for each disease, assisting the model in cross-modal alignment and extracting more representative visual features. Although the above methods have made significant progress in language quality and clinical accuracy, they still generally rely on current medical images and lack effective integration of patient-specific prior knowledge, limiting the model's ability to understand patient contextual information and track disease evolution.

[0004] Existing automated radiology report generation methods primarily aim to generate high-quality preliminary radiology reports to alleviate the workload of radiologists. However, in actual clinical scenarios, radiologists often rely on patient-specific prior knowledge (including clinical context information and previous images, where the clinical context information includes patient indications and medical history) for personalized diagnosis and to write accurate radiology reports accordingly. Current mainstream technologies have the following two core technical problems: (1) Insufficient utilization of clinical context information. Most existing methods generate radiology reports based solely on current images, ignoring patient context information such as indications, medical history, and symptoms, resulting in radiology reports lacking individualized content and failing to accurately reflect the patient's specific condition. (2) Limited disease evolution modeling capabilities. Some technologies fail to effectively model the dynamic differences between current and previous images, making it difficult to systematically track the disease progression process. This limitation easily leads to the "generation illusion" problem, where the content generated by the model deviates from the actual clinical evolution path, affecting the accuracy and reliability of the diagnosis.

[0005] To address the aforementioned issues, existing research has attempted to incorporate patient contextual information to improve the personalization and accuracy of radiology report generation. Huang et al. used BiLSTM to encode patient indication information, generating more targeted radiology reports. Nguyen et al. combined indication information with an LLaMA model to predict positive diseases to guide radiology report generation; however, their efforts in deeply mining indication information remain insufficient. Considering the potential noise in indication text, SEI improves data quality by cleaning illegal characters and invalid words and standardizing gender expressions. Simultaneously, it designs a cross-modal fusion network to integrate indication information, thereby enhancing the model's ability to understand patient contextual information. On the other hand, to alleviate the illusion problem in describing disease evolution, researchers have attempted to model the differences between current and previous images to achieve effective tracking of disease progression. RECAP utilizes previous radiology reports to model the disease evolution status of specific regions in images, achieving significant results in improving the coherence and accuracy of radiology reports; HERGen models the disease evolution process using a population causal Transformer. However, these methods still fail to fully integrate patient contextual information, resulting in an incomplete understanding of the examination context, which in turn affects the overall quality of radiological reports. Summary of the Invention

[0006] To address the problems existing in the prior art, the present invention discloses a method, system, device, and medium for automatically generating radiology reports based on patient-specific prior knowledge. Through a spatiotemporal fusion network, it flexibly integrates existing patient medical images to construct a difference mapping relationship between current and past images, used to model the disease evolution trajectory and extract spatiotemporal visual features. Through a progressive fusion network based on prior perception, it gradually integrates patient-specific prior knowledge and hierarchical visual features in a coarse-to-fine manner to generate personalized radiology reports. Through two-stage training, it improves the accuracy of medical image retrieval and the quality of radiology report generation. The present invention can effectively integrate patient-specific clinical context information and model the disease evolution process, improving the overall performance of the model in terms of personalization, accuracy, and clinical applicability.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A method for automatically generating radiology reports based on patient-specific prior knowledge, specifically including the following steps:

[0009] Step 1: Establish a special token missing encoding strategy. By introducing a special token, missing patient clinical context information is explicitly encoded. This clinical context information includes the patient's indications and medical history, enabling the text encoder to uniformly process complete and incomplete clinical context inputs, thereby extracting robust contextual features.

[0010] Step 2: Establish a spatiotemporal fusion network (STF). Input current medical images and historical patient images into the network, establish a difference mapping relationship between them, and use this to model the disease evolution process and extract time-dependent spatiotemporal visual features.

[0011] Step 3: Establish an attention-enhanced hierarchical fusion network to effectively integrate the multi-layered hidden states of the current medical image extracted by the visual encoder, thereby extracting multi-layered, semantically rich hierarchical visual features.

[0012] Step 4: Establish a progressive fusion network of prior perceptions, gradually integrating patient-specific prior knowledge and hierarchical visual features in a coarse-to-fine manner. Specifically, this includes fusing the clinical context features T extracted in Step 1. c The spatiotemporal visual features extracted in step 2 and the hierarchical visual features extracted in step 3 This generates multimodal features for generating radiology reports;

[0013] Step 5: Construct a two-stage training strategy based on prior knowledge to improve the accuracy of medical image and text retrieval and the quality of radiology report generation. This includes the following two stages: Stage 1: Perform prior-guided image-text alignment and comparison pre-training. By introducing patient-specific prior knowledge, a self-supervised contrastive learning approach is used to pre-train the current medical image and radiology report, optimizing their alignment capabilities in the shared semantic space, thereby improving the accuracy of medical image and text retrieval. Stage 2: Perform prior-aware radiology report generation training. Based on the model parameters from Stage 1, the multimodal fusion features generated in Step 4 for radiology report generation are further input into the text generator for supervised training, thereby improving the accuracy of the final generated radiology report in clinical semantic expression and the fluency of language organization.

[0014] The specific method for step 1 is as follows:

[0015] The patient's indications {ind-data} and patient's medical history {hx-data} are encoded in the following format:

[0016] t i ="[INDICATION]{ind-data i}"+"[SEP]"+"[HISTORY]{hx-data i}" (1)

[0017] Here, "[INDICATION]", "[SEP]", and "[HISTORY]" all represent special tokens used to clearly distinguish information originating from patient indications and patient medical history. Multiple input information types are uniformly encoded into standardized strings, simplifying the feature extraction process for patient clinical context information. Subsequently, the standardized strings are input into a text encoder to extract clinical context features. Where B is the batch size, s represents the number of tokens, and d is the dimension of each token.

[0018] The specific method for step 2 is as follows:

[0019] First, the current medical image and previous medical images are input into the visual encoder to extract the corresponding visual feature representations, denoted as V respectively. c and V p V c V represents the visual characteristics of current medical imaging. p Indicates the visual characteristics of previous medical images;

[0020] Subsequently, the visual feature V c and V pAdd view position code E respectively view (·), Projector head P proj (·) and time location code E time (·), which aims to enhance the visual encoder's ability to perceive and model imaging perspective differences and time-series information, is represented as follows:

[0021]

[0022] Next, the enhanced visual feature V mentioned above... cur and V pri The data is input into a Spatiotemporal Fusion Network (STF) to model the semantic differences between current and previous images and to characterize the evolution of the disease over time, thereby outputting spatiotemporal visual features. The calculation process of the spatiotemporal fusion network STF is expressed as follows:

[0023]

[0024] Where LN(·) is the layer normalization operation, FFN(·) represents the feedforward neural network, ATTN(Q,K,V) represents the cross-attention mechanism for spatiotemporal feature fusion, and d represents the dimension of the query vector and key vector in the attention mechanism; for samples without prior medical images, the enhanced visual feature V cur Considered as spatiotemporal visual features Where B is the batch size and p represents the number of visual patches.

[0025] The specific method for step 3 is as follows:

[0026] First, the current medical image is input into the visual encoder, where it undergoes layer-by-layer feature extraction. Assuming the visual encoder contains an L-layer neural network structure, the output multi-layer hidden states are denoted as...

[0027] Subsequently, the multi-layered hidden state H is input into the channel attention CAM(·) and spatial attention SAM(·) modules of CBAM to achieve dynamic highlighting of important semantics, and the output feature is denoted as:

[0028]

[0029] Next, the above feature H cbam The input is fed into the Channel Fusion Projector (CFP) module for feature compression and transformation. The output features are denoted as:

[0030] H out=CFP(H) cbam )=Conv2D(BN(ReLU(Conv2D(H cbam (5)

[0031] Where Conv2D(·) represents a 2D convolution operation, BN(·) represents batch normalization, and ReLU(·) represents the activation function; the output is a feature fused from multiple hidden states. This achieves feature simplification and unification while retaining multiple layers of information;

[0032] Finally, the feature H that integrates multiple hidden states. out After projection head transformation and the addition of temporal position encoding, followed by layer normalization, a hierarchical visual representation is obtained that simultaneously integrates fine-grained local details and high-level semantic features.

[0033] The prior-aware progressive fusion network described in step 4 is built on the Perceiver architecture. Perceiver(P,Q) is a general architecture that iteratively compresses the input Q into a set of compact, learnable latent representations P through a cross-attention mechanism to achieve unified fusion of multimodal information. Specifically, the computation process of the prior-aware progressive fusion network is as follows:

[0034]

[0035]

[0036]

[0037] in, Let T be the initial implicit representation, N be the number of implicit representations, and T be the number of implicit representations. c This represents the clinical context features obtained in step 1. It is a compressed representation of the clinical context. Represents the spatiotemporal features of integrated clinical semantics. To achieve the final multimodal features that integrate patient-specific prior knowledge with visual semantics.

[0038] The specific method for step 5 is as follows:

[0039] a. First stage: Prior-guided comparative pre-training, the process is as follows:

[0040] First, the compressed clinical context representation is obtained using formulas (6) and (7). And spatiotemporal features that integrate clinical semantics This allows for the introduction of patient clinical context information into the visual encoding process, enabling prior knowledge to guide the extraction of spatiotemporal visual features.

[0041] Then, splicing and After global average pooling and L2 normalization, global visual features are obtained.

[0042] Next, the similarity (logits) between medical images and radiology reports is calculated. The calculation process can be expressed as follows:

[0043]

[0044] in, The global textual features of the radiology report are defined, and τ is the temperature coefficient. Similarly, the similarity between the radiology report and the medical image is calculated. The calculation process can be expressed as follows:

[0045]

[0046] Considering that a patient may have multiple medical images during a single visit, all medical images and radiology reports associated with the same visit are considered as positive sample pairs, thus generating multiple positive sample pairs; therefore, a true matching label matrix is ​​defined. The calculation process for indicating the matching relationship between medical image and radiology report pairs can be expressed as follows:

[0047]

[0048] in, It is an indicator function, when y i =y j The value is 1 if it is true, and 0 otherwise.

[0049] Finally, the loss function for prior-guided contrastive pre-training is the cross-entropy loss between q and p, and its calculation process can be expressed as follows:

[0050]

[0051] b. Second stage: Generation of prior-perceived radiological reports, the process is as follows:

[0052] First, the compressed clinical context representation obtained in step 5 is... Spatiotemporal representation incorporating clinical semantics And the final multimodal features that integrate patient-specific prior knowledge and visual semantics. The sequences are concatenated to form the input sequence for the text generator; this input sequence helps to provide the text generator with rich and diverse semantic context, thereby guiding the text generator to output more personalized and clinically relevant radiology reports.

[0053] Then, the radiology report generated by the text generator is minimized. Radiology reports written by actual radiologists i The cross-entropy loss function between the two is used to supervise the learning of the text generator, and its optimization objective can be expressed as:

[0054]

[0055] Where K is the maximum length of the text generator output. B represents all tokens generated up to the k-th token, and B is the batch size.

[0056] An automated report generation system based on a patient-specific prior knowledge-based method for automatically generating radiology reports includes:

[0057] The token missing encoding module is used in step 1 to represent missing patient clinical context information using a special token, so that the text encoder can uniformly process complete and incomplete input cases, in order to extract personalized text features.

[0058] The spatiotemporal fusion network module is used to integrate the patient's previous medical images in step 2, construct a difference mapping between the current medical images and previous medical images, model the evolution trajectory of the disease, and thus extract spatiotemporal visual features.

[0059] The attention-enhanced hierarchical fusion network module is used to efficiently integrate the multi-layer hidden states in the visual encoder in step 3, thereby facilitating the extraction of hierarchical visual features.

[0060] A priori-aware progressive fusion network module is used to progressively integrate the personalized text features extracted in step 1 in step 4. Spatiotemporal visual features extracted in step 2 and the hierarchical visual features generated in step 3

[0061] A two-stage training module is used in step 5 to optimize the consistency of multimodal representation between medical images and radiology reports, as well as the quality of radiology report generation.

[0062] An automated report generation device for a patient-specific prior knowledge-based method for automatically generating radiology reports, comprising:

[0063] Memory, used to store computer programs;

[0064] A processor for executing the computer program to implement the method for automatically generating radiological reports based on patient-specific prior knowledge.

[0065] A computer-readable storage medium storing a computer program, characterized in that, when executed by a processor, the computer program is capable of automatically generating a radiology report according to the patient-specific prior knowledge-based automatic radiology report generation method.

[0066] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0067] 1. A special token missing encoding strategy enhances robust modeling capability for clinical context. In step 1, this invention designs a special token missing encoding strategy. By explicitly introducing special markers to represent missing information, the text encoder can uniformly process complete and incomplete clinical context inputs, effectively reducing representation bias caused by missing clinical context. This improves the model's robust modeling capability for patient information in real clinical scenarios, ensuring better adaptability in real-world clinical situations.

[0068] 2. Spatiotemporal fusion modeling to capture disease evolution patterns and alleviate hallucination problems. In step 2, this invention introduces a spatiotemporal fusion network (STF). By jointly modeling the differences between current medical images and the patient's historical images, and introducing perspective encoding and temporal encoding mechanisms, it constructs the temporal dependencies in the disease development process. This extracts more semantically continuous spatiotemporal visual features, improves the system's ability to understand pathological evolution trends, effectively reduces the risk of false, redundant, or inconsistent information in generated radiology reports, and enhances the clinical credibility and practicality of radiology reports.

[0069] 3. Multi-level visual modeling and prior-driven multimodal fusion mechanism to comprehensively enhance semantic expression capabilities. This invention proposes a comprehensive optimization mechanism integrating "semantic enhancement, prior guidance, and progressive fusion." First, in terms of visual representation, by designing an attention-enhanced hierarchical fusion network, the hidden states at different levels in the visual encoder are fully integrated to extract multi-scale, semantically deep, and hierarchical visual features, thereby improving the subtlety and structural integrity of visual semantic expression. Second, in terms of feature fusion, a prior-aware progressive multimodal fusion network is constructed, fusing patient-specific clinical context information, spatiotemporal visual features, and hierarchical visual features layer by layer in a coarse-to-fine order. This not only preserves the hierarchical structure of each modality feature but also achieves cross-modal semantic collaborative optimization, generating multimodal features that are more consistent with medical diagnostic logic. Finally, in terms of the training mechanism, a two-stage training strategy based on prior knowledge is innovatively introduced. Image-text semantic alignment is optimized through comparative learning, and then the text generator is optimized based on multimodal features, thereby improving the accuracy of image-text retrieval and the quality of radiology report generation, respectively.

[0070] In summary, this invention systematically addresses the problems of insufficient utilization of clinical context information, inadequate disease evolution modeling, and incomplete visual semantic representation in existing technologies by constructing a multimodal modeling framework that integrates patient clinical context information, spatiotemporal visual features, and multi-level visual features. Simultaneously, the proposed progressive fusion mechanism based on prior perception enables synergistic optimization of the structural and semantic levels among multimodal features. Combined with a two-stage training strategy guided by prior knowledge, it first optimizes the image-text semantic alignment capability through comparative learning, and then performs specialized training for the radiology report generation task, effectively improving the accuracy of image-text retrieval and the quality of generated radiology reports. This method, while ensuring the consistency of clinical semantics and the accuracy of language expression in generated radiology reports, enhances the adaptability of this invention to complex real-world clinical data, possessing comprehensive technical effects and practical application value, including improving the clinical practicality of radiology reports, assisting physicians in efficient decision-making, and significantly reducing the workload of radiologists. Attached Figure Description

[0071] Figure 1 This is a schematic diagram of the spatiotemporal fusion network and the attention-enhanced layered fusion network algorithm framework; where (A) shows the spatiotemporal fusion network and (B) shows the attention-enhanced layered fusion network.

[0072] Figure 2 This is a framework for an algorithm to automatically generate radiology reports based on patient-specific prior knowledge.

[0073] Figure 3 This is the consensus assessment result of expert preferences on the MIMIC-CXR dataset.

[0074] Figure 4 The results show a performance comparison of medical image and text retrieval on the MIMIC-5x200 dataset, where the method of this invention is compared and validated with existing mainstream medical image and text retrieval methods. "CC" and "PI" represent clinical context information and previous images, respectively; where, Figure 4 (a) shows the Study-Precision@K evaluation index, which measures whether the search results are derived from radiological studies with the same queried images, and mainly reflects the semantic alignment effect at the patient level. Figure 4 (b) shows the Category-Precision@K evaluation metric, which measures whether the search results belong to the same disease category as the query image, reflecting the semantic matching capability at the category level.

[0075] Figure 5 The qualitative comparison results generated for radiology reports on the MIMIC-CXR dataset were used to compare and verify the present invention with the existing mainstream radiology report generation method SEI; statements that accurately describe disease progression are marked in bold, while errors or omissions are marked with underlines. Detailed Implementation

[0076] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0077] A method for automatically generating radiology reports based on patient-specific prior knowledge, wherein the patient-specific prior knowledge consists of patient clinical context information and previous medical images, wherein the clinical context information includes the patient's indications and medical history. Specifically, it includes the following steps:

[0078] (1) Special token missing encoding strategy

[0079] Due to limitations in data storage and management mechanisms, some patients' contextual information (such as indications and medical history) may be missing or incomplete. Nevertheless, this information is of significant reference value in clinical decision-making. To uniformly handle complete and incomplete input data, this invention proposes a special token missing encoding strategy. Specifically, the patient's indication content {ind-data} and medical history content {hx-data} are encoded in the following form:

[0080] t i ="[INDICATION]{ind-data i}"+"[SEP]"+"[HISTORY]{hx-data i}" (1)

[0081] Here, "[INDICATION]", "[SEP]", and "[HISTORY]" represent special tokens used to clearly distinguish whether the information originates from the patient's indications or medical history. This strategy can uniformly encode various input scenarios into standardized strings, simplifying the feature extraction process for patient context information. Subsequently, the standardized text is input into a text encoder to extract clinical context features. Where B is the batch size, s represents the number of tokens, and d is the dimension of each token.

[0082] (2) Spatiotemporal Fusion Network

[0083] Although radiological reports can be generated based solely on the current image, the model is prone to producing hallucinatory content when describing disease progression (e.g., "As compared to the previous radiograph, the patient has received a new right internal jugular vein catheter."). To mitigate this problem, this invention designs a spatiotemporal fusion network (STF), such as... Figure 1 As shown in (A), this method is used to model disease progression and extract spatiotemporal visual features. The specific process is as follows: First, the current medical image and previous medical images are input into the visual encoder to extract the corresponding visual feature representations, denoted as V. c and V p V c V represents the visual characteristics of current medical imaging. p This represents the visual features of previous medical images; subsequently, the visual features V... c and V p Add view position code E respectively view (·), Projector head P proj (·) and time location code E time (·), to enhance the model's ability to model differences in imaging perspective and time series differences, the processing can be expressed as:

[0084]

[0085] Next, the enhanced visual feature V cur and V pri The data is input into a Spatiotemporal Fusion Network (STF) to construct a difference mapping between current and previous medical images, model the evolution trajectory of diseases, and output spatiotemporal visual features. The calculation process of the spatiotemporal fusion network STF can be represented as follows:

[0086]

[0087] Where LN(·) is the layer normalization operation, FFN(·) represents the feedforward neural network, ATTN(Q,K,V) represents the cross-attention mechanism for spatiotemporal feature fusion, and d represents the dimension of the query vector and key vector in the attention mechanism; for samples without prior medical images, the enhanced visual feature V cur Considered as spatiotemporal visual features Where B is the batch size and p represents the number of visual patches.

[0088] (3) Attention-enhanced hierarchical fusion network

[0089] Existing radiology report generation methods generally rely on the last hidden state of the visual encoder as input, which may lead to the loss of low-level local details such as lesion morphology, limiting the model's ability to capture fine-grained medical features. To overcome this limitation, this invention designs an attention-enhanced layer fusion network based on CBAM, such as... Figure 1 As shown in (B). Specifically, firstly, the current medical image is input into the visual encoder, and layer-by-layer feature extraction is performed on it; assuming the encoder contains an L-layer neural network structure, the output multi-layer hidden states are denoted as... Subsequently, the multi-layered hidden state H is input into the channel attention CAM(·) and spatial attention SAM(·) modules of CBAM to achieve dynamic highlighting of important semantics. The output feature is denoted as:

[0090]

[0091] Next, the above features The input is fed into the Channel FusionProjector (CFP) module for feature compression and transformation. The output features are denoted as:

[0092] H out =CFP(H) cbam )=Conv2D(BN(ReLU(Conv2D(H cbam (5)

[0093] Where Conv2D(·) represents a 2D convolution operation, BN(·) represents batch normalization, and ReLU(·) represents the activation function. The output is a feature fused from multiple hidden states. This achieves feature simplification and unification while preserving multi-layered information; finally, it fuses the features H of multiple hidden states. out After projection head transformation and the addition of temporal position encoding, followed by layer normalization, a hierarchical visual representation is obtained that simultaneously integrates fine-grained local details and high-level semantic features.

[0094] (4) Progressive fusion network based on prior perception

[0095] Inspired by the hierarchical structure of visual representation and the diagnostic workflow of radiologists, this invention proposes a progressive fusion network based on prior perception, aiming to gradually integrate clinical context features in a coarse-to-fine manner. Spatiotemporal visual features and hierarchical visual features This results in radiological reports that are more clinically relevant. The fusion network is built on the Perceiver architecture, a general architecture that iteratively compresses the input Q into a set of compact, learnable latent representations P through a cross-attention mechanism, achieving unified fusion of multimodal information. Accordingly, the prior-aware, progressive fusion process can be formally represented as follows:

[0096]

[0097]

[0098]

[0099] in, Let N be the initial implicit representation, and N be the number of implicit representations. It is a compressed representation of the clinical context. Represents the spatiotemporal features of integrated clinical semantics. This is to achieve the final multimodal features by fusing patient-specific prior knowledge with visual semantics. It's worth noting that, since the first phase of training only targets instance-level cross-modal alignment tasks, the final multimodal features... It is only used in the second phase for the radiology report generation task.

[0100] (5) Two-stage training strategy

[0101] To simplify the overall training difficulty, this invention designs a two-stage training strategy, such as... Figure 2 As shown: In the first stage, prior-guided contrastive pre-training is performed, using self-supervised learning on medical images and corresponding radiology reports. Prior knowledge is introduced to enhance the semantic alignment of image-text matching, thereby improving the accuracy of medical image-text retrieval. In the second stage, prior-aware radiology report generation is performed. A progressive fusion network based on prior awareness is used to gradually integrate patient-specific prior knowledge and hierarchical visual features in a coarse-to-fine manner, improving the generated radiology reports in terms of clinical semantic accuracy and language expression quality. The loss functions used in each stage and their effects are described below.

[0102] a. First stage: Prior-guided comparative pre-training

[0103] This stage aims to improve the model's ability to align clinical semantics with visual features, i.e., cross-modal representation capability. In clinical practice, radiologists first conduct an initial clinical assessment by combining the patient's symptoms (indications) and medical history (history). Subsequently, they integrate previous and current images to assess the disease progression trend. To simulate this diagnostic process, this invention designs a priori-guided contrastive pre-training method, the process of which is as follows:

[0104] First, the compressed clinical context representation is obtained using formulas (6) and (7). And spatiotemporal features that integrate clinical semantics This allows us to use patient information from both clinical and extracorporeal perspectives to guide the extraction of spatiotemporal features.

[0105] Then, splicing and After global average pooling and L2 normalization, global visual features are obtained.

[0106] Next, the similarity (logits) between medical images and radiology reports is calculated. The calculation process can be expressed as follows:

[0107]

[0108] in, Let τ be the global textual feature of the radiology report, and τ be the temperature coefficient; similarly, the similarity between the radiology report and the medical image can be calculated. The calculation process can be expressed as follows:

[0109]

[0110] Considering that a patient may have multiple medical images during a single visit, all medical images and radiology reports associated with the same visit are considered as positive sample pairs, thus generating multiple positive sample pairs. Therefore, a true matching label matrix is ​​defined. The calculation process for indicating the matching relationship between medical image and radiology report pairs can be expressed as follows:

[0111]

[0112] in, It is an indicator function, when y i =y j The value is 1 if the condition is met, and 0 otherwise.

[0113] Finally, the loss function for prior-guided contrastive pre-training is the cross-entropy loss between q and p, and its calculation process can be expressed as follows:

[0114]

[0115] In summary, the training objective function for the first stage is L. align It effectively improves the model's ability to align clinical semantics with visual features by integrating patient contextual information and guiding spatiotemporal visual feature learning.

[0116] b. Second stage: Generation of radiological reports based on prior perception

[0117] This phase focuses on the task of generating radiology reports, optimizing the clinical accuracy and language quality of these reports. To better align with the diagnostic workflow of radiologists, this invention designs a priori-aware radiology report generation method, the overall process of which is as follows: First, the compressed clinical context representation obtained in step 5 is... Spatiotemporal representation incorporating clinical semantics And the final multimodal features that integrate patient-specific prior knowledge and visual semantics. These sequences are concatenated to form the input sequence for the text generator. This input sequence helps provide the text generator with rich and diverse semantic context, thereby guiding the text generator to output more personalized and clinically relevant radiology reports.

[0118] Then, the radiology report generated by the text generator is minimized. Radiology reports written by actual radiologists i The cross-entropy loss function between the two is used to supervise the learning of the text generator, and its optimization objective can be expressed as:

[0119]

[0120] Where K is the maximum length of the text generator output. B represents all tokens generated up to the k-th token, and B is the batch size.

[0121] Experimental data

[0122] a. Experiment Setup Instructions

[0123] Dataset Introduction. (1) MIMIC-CXR is a large-scale, publicly available dataset of chest X-ray images and free-text radiology reports. This invention organizes the data based on “study id” to retrieve the latest previous images corresponding to each sample, thereby supporting modeling of disease evolution. (2) MIMIC-ABN is a variant dataset derived from MIMIC-CXR, focusing on the description of abnormalities in radiology reports. Following the common practice in existing studies, only the “Findings” section of the report is used as the actual radiology reports written by radiologists, and samples with invalid or empty content are removed. All experiments strictly follow the official division of training, validation and test sets. For relevant data statistics, see Table 1, where “#Image” and “#Report” represent the number of medical images and radiology reports, respectively; “%PI”, “%Indication” and “%History” represent the proportion of samples containing previous images, indication information and medical history information, respectively.

[0124] Table 1. Data distribution of training, validation, and test sets in the MIMIC-CXR and MIMIC-ABN datasets.

[0125]

[0126]

[0127] Evaluation Metrics. Following standard evaluation protocols from existing research, the quality of model-generated radiology reports was assessed across two dimensions: Natural Language Generation (NLG) and Clinical Efficacy (CE). The NLG metric primarily evaluates the linguistic similarity between the generated and actual radiology reports. The evaluation metrics used included BLEU-n (Bn), METEOR (MTR), and ROUGE-L (RL). The CE metric primarily evaluates the clinical relevance of the generated radiology report. Specifically, based on the 14 observation labels provided by CheXpert, the micro-means of Precision (P), Recall (R), and F1 score (F1) of the model-predicted labels were calculated. All evaluation metrics were implemented using publicly available tools: the NLG metric was implemented using the pycocoevalcap toolkit, and the CE metric was implemented using f1chexbert. Higher scores on all of the above metrics indicate better radiology report quality.

[0128] Implementation details. (1) MIMIC-CXR: In the first stage, the AdamW optimizer is used for training, with 30 training epochs and an initial learning rate of 5e-5. In the second stage, the model is initialized based on the pre-trained weights from the first stage and further fine-tuned for 30 epochs. During the fine-tuning stage, the learning rate of the text generator is set to 5e-5, and the learning rate of the other parameters is set to 5e-6. (2) MIMIC-ABN: Since this dataset is derived from MIMIC-CXR, this invention uses the weights pre-trained on MIMIC-CXR for initialization and uses the AdamW optimizer with a learning rate of 5e-6 for 30 epochs for fine-tuning. (3) General settings: The image encoder uses RAD-DINO, the text encoder uses CXR-BERT, and the text generator uses DistilGPT2. The unified feature dimension is set to d=768, and L=3 STF modules are used (e.g., Figure 1 As shown in the figure, the number of implicit representations is N=128, and the maximum generation length of radiology reports is K=100 tokens. During training, the ReduceLROnPlateau learning rate scheduling strategy is used, with a scheduling patience value set to 5; at the same time, an early stopping strategy is enabled, with a patience value set to 15, to prevent model overfitting.

[0129] b. Explanation of main experimental results

[0130] Table 2 compares the performance of current state-of-the-art methods on the MIMIC-CXR (M-CXR) and MIMIC-ABN (M-ABN) datasets.

[0131]

[0132]

[0133] Comparison with State-of-the-Art Methods. This invention compares the proposed PriorRG method with 11 state-of-the-art (SOTA) methods, categorized into seven groups: knowledge graph methods (KiUT and METransformer), contrastive learning frameworks (CoFE), retrieval enhancement methods (DCG), memory alignment methods (CMN and MAN), large language model methods (R2GenGPT, Med-LLM, and R2-LLM), clinical context-driven models (SEI), and longitudinal data methods (HERGen). The comparison results on the MIMIC-CXR and MIMIC-ABN datasets are shown in Table 2, where “Δ” indicates the performance improvement of PriorRG compared to the best baseline method; “◇” indicates that the result was reproduced from the original code, and other results are directly cited from the original paper; all best results are indicated in bold, and second-best results are indicated by underline. Table 2 shows that PriorRG significantly outperforms existing methods in both natural language generation quality and clinical relevance. On the MIMIC-CXR dataset, PriorRG achieved a 4.0% improvement over the best baseline on BLEU-4 (B-4), reflecting the model's advantage in continuous n-gram matching. Simultaneously, it improved by 1.3% and 2.0% on METEOR (MTR) and ROUGE-L (RL) metrics, respectively, demonstrating superior semantic consistency and linguistic expressiveness. In terms of clinical effectiveness, PriorRG's micro-mean F1 score improved by 3.8%, indicating greater reliability in clinical accuracy. On the MIMIC-ABN dataset, PriorRG also demonstrated a leading advantage, with a 5.9% improvement on BLEU-1 (B-1) and a 1.1% improvement on the F1 score, further validating its superior ability to describe abnormal states. PriorRG's performance improvement is primarily attributed to its "stepwise fusion mechanism of prior perception," which integrates patient-specific clinical context, spatiotemporal dynamic features, and hierarchical visual semantics from coarse to fine. By simulating the diagnostic process of radiologists, PriorRG can generate radiology reports that are both consistent with individual conditions and fluently expressed, thus achieving a comprehensive improvement in both linguistic quality and clinical accuracy.

[0134] Table 3. Accuracy of 14 common chest diseases on the MIMIC-CXR dataset.

[0135]

[0136] Accuracy for 14 Common Diseases. Table 3 shows the accuracy results of the CheXpert tool for assessing 14 common chest diseases. Compared with the representative clinical context method SEI, PriorRG achieved higher recall (R) in 12 of the diseases, especially in clinically relevant but low-frequency conditions such as Pneumonia, Pneumothorax, and Fractures. This result indicates that PriorRG not only has strong overall recognition capabilities but also effectively reduces the risk of missing low-frequency key diseases, even though the model does not explicitly incorporate class imbalance handling mechanisms during training.

[0137] Expert Preference Consistency Assessment Based on a Large Language Model. To systematically evaluate the clinical consistency of generated radiology reports and their alignment with expert preferences, this invention employs two key indicators: "#Matched Findings" and the "GREEN" score. The "GREEN" score comprehensively considers both the number of critical clinical errors and the number of matched clinical manifestations (#Matched Findings), providing a more comprehensive and objective measure of the clinical validity and expert preference consistency of radiology reports. Specifically, the pre-trained large language model GREEN-RadLlama2-7B is used to automatically extract the two evaluation indicators, and this method (PriorRG) is compared with three representative methods: CMN, CvT2DistilGPT2, and SEI. Figure 3 As shown, PriorRG significantly outperforms existing methods in both the "#Matched Findings" and "GREEN" scores, further validating its comprehensive advantages in generating clinically reliable radiology reports with a style close to that of experts.

[0138] c. Explanation of ablation test results

[0139] Analysis of the effect of prior-guided comparison pre-training (Phase 1). To verify the effectiveness of Phase 1, this invention conducted a medical image-text retrieval task on the MIMIC-5x200 dataset. Referring to the experimental setup in existing work, five common diseases were randomly selected from the MIMIC-CXR test set: Atelectasis, Cardiomegaly, Consolidation, Edema, and Pleural Effusion. 200 samples were collected for each disease, resulting in a total MIMIC-5x200 dataset containing 1000 samples. In this retrieval task, the system needs to retrieve the top K most relevant radiological reports from the dataset based on a given query image. This invention uses two evaluation metrics to measure retrieval performance: Cat-P@K (Category-Precision@K, which measures whether the retrieval results belong to the same disease category as the query image, reflecting category-level semantic matching ability); and Stu-P@K (Study-Precision@K, which evaluates whether the retrieval results come from the same radiological study as the query image, reflecting patient-level semantic alignment effect). Figure 4 As shown, PriorRG significantly outperforms current mainstream methods (BiomedCLIP and BioViL-T) and several of their variants in both of the above metrics, indicating that the prior-guided contrastive pre-training strategy proposed in this invention can effectively utilize patient-specific clinical context to enhance the discriminative ability and semantic consistency of multimodal representations, thereby significantly improving cross-modal retrieval performance. Furthermore, the results in Table 4 further demonstrate that, compared to the control model without the first-stage pre-training (PriorRG vs.(f)), PriorRG exhibits significant performance improvements in both retrieval and generation tasks, validating the crucial role of the first stage in improving the clinical accuracy and linguistic fluency of generated radiology reports.

[0140] PriorRG's performance in generating radiology reports (Phase II) was analyzed. As shown in Table 4, PriorRG significantly outperformed the baseline model (variant (a)) without Phase II in both Natural Language Generation (NLG) and Clinical Efficacy (CE) metrics. This result strongly suggests that Phase II plays a crucial role in improving the quality of radiology report generation.

[0141] Table 4 shows the ablation experimental results on the MIMIC-CXR dataset. “CC”, “PI”, and “HS” represent clinical context information, previous images, and the implicit state of the image encoder, respectively.

[0142]

[0143] The role of patient-specific prior knowledge. Compared with the variant (variant (b)) in Table 4 that does not incorporate patient-specific prior knowledge, PriorRG achieved significant improvements in both NLG and CE metrics. This result indicates that patient-specific prior knowledge (including clinical context information and previous imaging) has a positive driving effect on the radiology report generation task, significantly enhancing the personalization and clinical relevance of the generated content.

[0144] The role of patient-specific clinical context. In medical image and text retrieval tasks, such as... Figure 4 As shown, the complete PriorRG model was compared with its variant without clinical context (CC) (PriorRG w / oCC). Experimental results show that introducing clinical context effectively enhances the model's understanding and modeling of spatiotemporal visual features, thereby improving the alignment between visual and clinical semantics. In the radiology report generation task, based on variant (a) in Table 4, version (c) with clinical context incorporated achieved significant improvements in both NLG and CE metrics. This result indicates that patient-specific clinical context information helps the model more accurately capture individualized clinical features, thereby improving both the fluency of radiology report language and its diagnostic accuracy.

[0145] The role of previous images. In medical image retrieval tasks, such as... Figure 4 As shown, PriorRG significantly outperforms the variant without prior images (PriorRG w / o PI) on the Study-Precision@K metric, indicating that the model can effectively model the disease evolution process by comparing current and prior images, thereby improving patient-level retrieval accuracy. In contrast, this advantage is less pronounced in category-level retrieval, possibly because the disease evolution status interferes with the extraction of global visual representations, weakening the consistency of disease categories. In the radiology report generation task, the results in Table 4(c) compared to (e), and PriorRG compared to (d), further validate the positive role of prior images on the NLG and CE metrics, highlighting their key value in improving the clinical accuracy of generated radiology reports.

[0146] The role of hierarchical visual features. As shown in Table 4, PriorRG achieves significant improvements in both NLG and CE metrics compared to variant (e), validating the crucial role of hierarchical visual features extracted from multi-layer hidden states in the radiology report generation task. This feature indicates the ability to capture finer-grained local information in images, helping the model generate radiology reports with richer clinical details and more fluent language.

[0147] Table 5 Comparison of radiology report generation performance of different progressive fusion strategies on the MIMIC-CXR dataset.

[0148]

[0149] The impact of progressive fusion strategies on radiology report generation. Table 5 compares three progressive fusion strategies used for radiology report generation: LastOnly (using only the final multimodal features). The text generator input methods include: Fine-to-Coarse (integrating clinical context, multi-layered implicit states, and spatiotemporal visual features sequentially from fine to coarse granularity), and Coarse-to-Fine (the strategy adopted by PriorRG, fusing clinical context, spatiotemporal visual features, and multi-layered implicit states sequentially from coarse to fine granularity). Both Fine-to-Coarse and Coarse-to-Fine concatenate the outputs of the three Perceiver modules before inputting them into the text generator. Experimental results show that PriorRG outperforms the other two variants in both NLG and CE metrics, indicating that the coarse-to-fine information fusion strategy is more conducive to generating fluent and clinically accurate radiology reports. Although Fine-to-Coarse has a slight advantage in F1 score, it lags behind PriorRG in terms of language expression and content consistency. LastOnly performed the worst, validating the crucial role of intermediate fusion representation in improving the fluency of radiology reports.

[0150] d. Explanation of Qualitative Experimental Results

[0151] Figure 5This paper presents a qualitative comparison of PriorRG and SEI's radiology report generation results on the MIMIC-CXR test set. The color bars in the figure indicate the correspondence between the generated radiology report and the actual radiology report across different clinical findings: more colors indicate a richer range of clinical findings covered in the generated report; the length of the color bar reflects the accuracy and detail of the generated text's description of a particular clinical finding. As can be seen from the figure, PriorRG's generated radiology report outperforms SEI in terms of language expression, information coverage, and clinical relevance, requiring minimal manual modification. For example, in Case 1, the radiologist only needs to add a brief description of pleural effusion to finalize the radiology report. Furthermore, PriorRG effectively utilizes previous imaging to model disease progression, accurately generating descriptions such as "There is unchanged cardiomegaly," demonstrating its ability to perceive longitudinal clinical changes. Although radiology reports primarily rely on imaging information, the introduction of clinical context also plays a crucial guiding role. For example, in Case 1, the "respiratory failure" prompting model focuses on manifestations related to respiratory dysfunction, such as cardiomegaly, pleural effusion, and pulmonary edema; while in Case 2, the "Dobbhofftube placement" complaint-guided model determines the specific location of the catheter and its potential complications. In summary, PriorRG, by integrating patient-specific clinical context and previous imaging information, not only improves the personalization and clinical consistency of radiology reports but also more closely reflects actual diagnostic and treatment processes, demonstrating the high quality and clinical usability of the generated radiology reports.

Claims

1. A method for automatically generating radiology reports based on patient-specific prior knowledge, characterized in that, Specifically, the following steps are included: Step 1: Establish a special token missing encoding strategy. This involves introducing a special token to explicitly encode missing patient clinical context information, including patient indications and medical history. This allows the text encoder to uniformly process both complete and incomplete clinical context inputs, thereby extracting robust contextual features T. c ∈ò B×s×d ; Step 2: Establish a spatiotemporal fusion network (STF). Input current medical images and historical patient images into the network, establish a difference mapping relationship between them, and use it to model the disease evolution process and extract time-dependent spatiotemporal visual features. Step 3: Establish an attention-enhanced hierarchical fusion network to effectively integrate the multi-layered hidden states of the current medical image extracted by the visual encoder, thereby extracting multi-layered, semantically rich hierarchical visual features. Step 4: Establish a progressive fusion network of prior perceptions, gradually integrating patient-specific prior knowledge and hierarchical visual features in a coarse-to-fine manner. Specifically, this includes fusing the clinical context features T extracted in Step 1. c The spatiotemporal visual features extracted in step 2 and the hierarchical visual features extracted in step 3 This generates multimodal features for generating radiology reports; Step 5: Construct a two-stage training strategy based on prior knowledge to improve the accuracy of medical image and text retrieval and the quality of radiology report generation. Specifically, it includes the following two stages: The first stage: Perform prior-guided image and text alignment comparison pre-training. By introducing patient-specific prior knowledge, a self-supervised comparison learning method is used to pre-train the current medical image and radiology report, optimize the alignment ability of the current medical image and radiology report in the shared semantic space, thereby improving the accuracy of medical image and text retrieval. The second stage involves training the radiology report generation based on prior perception. Building upon the model parameters from the first stage, the multimodal fusion features generated in step 4 for the radiology report generation task are further input into the text generator for supervised training. This improves the accuracy of the final generated radiology report in terms of clinical semantic expression and the fluency of language organization.

2. The method for automatically generating radiology reports based on patient-specific prior knowledge according to claim 1, characterized in that, The specific method for step 1 is as follows: The patient's indications {ind-data} and patient's medical history {hx-data} are encoded in the following format: t i ="[INDICATION]{ind-data i }"+"[SEP]"+"[HISTORY]{hx-data i }" (1) Here, "[INDICATION]", "[SEP]", and "[HISTORY]" all represent special tokens used to clearly distinguish information originating from patient indications and patient medical history. Multiple input information types are uniformly encoded into standardized strings, simplifying the feature extraction process for patient clinical context information. Subsequently, the standardized strings are input into a text encoder to extract clinical context features. Where B is the batch size, s represents the number of tokens, and d is the dimension of each token.

3. The method for automatically generating radiology reports based on patient-specific prior knowledge according to claim 1, characterized in that, The specific method for step 2 is as follows: First, the current medical image and previous medical images are input into the visual encoder to extract the corresponding visual feature representations, denoted as V respectively. c and V p V c V represents the visual characteristics of current medical imaging. p Indicates the visual characteristics of previous medical images; Subsequently, the visual feature V c and V p Add view position code E respectively view (·), Projector head P proj (·) and time location code E time (·), which aims to enhance the visual encoder's ability to perceive and model imaging perspective differences and time-series information, is represented as follows: Next, the enhanced visual feature V mentioned above... cur and V pri The data is input into a Spatiotemporal Fusion Network (STF) to model the semantic differences between current and previous images and to characterize the evolution of the disease over time, thereby outputting spatiotemporal visual features. The calculation process of the spatiotemporal fusion network STF is expressed as follows: Where LN(·) is the layer normalization operation, FFN(·) represents the feedforward neural network, ATTN(Q,K,V) represents the cross-attention mechanism for spatiotemporal feature fusion, and d represents the dimension of the query vector and key vector in the attention mechanism; for samples without prior medical images, the enhanced visual feature V cur Considered as spatiotemporal visual features Where B is the batch size and p represents the number of visual patches.

4. The method for automatically generating radiology reports based on patient-specific prior knowledge according to claim 1, characterized in that, The specific method for step 3 is as follows: First, the current medical image is input into the visual encoder, where it undergoes layer-by-layer feature extraction. Assuming the visual encoder contains an L-layer neural network structure, the output multi-layer hidden states are denoted as... Subsequently, the multi-layered hidden state H is input into the channel attention CAM(·) and spatial attention SAM(·) modules of CBAM to achieve dynamic highlighting of important semantics, and the output feature is denoted as: Next, the above feature H cbam The input is fed into the Channel Fusion Projector (CFP) module for feature compression and transformation. The output features are denoted as: OR out =CFP(H cbam )=Conv2D(BN(ReLU(Conv2D(H cbam )))) (5) Where Conv2D(·) represents a 2D convolution operation, BN(·) represents batch normalization, and ReLU(·) represents the activation function; the output is a feature H that fuses multiple hidden states. out ∈ò B×p×d This allows for feature simplification and unification while retaining multiple layers of information; Finally, the feature H that integrates multiple hidden states. out After projection head transformation and the addition of temporal position encoding, followed by layer normalization, a hierarchical visual representation is obtained that simultaneously integrates fine-grained local details and high-level semantic features.

5. The method for automatically generating radiology reports based on patient-specific prior knowledge according to claim 1, characterized in that, The prior-aware progressive fusion network described in step 4 is built on the Perceiver architecture. Perceiver(P,Q) is a general architecture that iteratively compresses the input Q into a set of compact, learnable latent representations P through a cross-attention mechanism to achieve unified fusion of multimodal information. Specifically, the computation process of the prior-aware progressive fusion network is as follows: in, Let T be the initial implicit representation, N be the number of implicit representations, and T be the number of implicit representations. c This represents the clinical context features obtained in step 1. It is a compressed representation of the clinical context. Represents the spatiotemporal features of integrated clinical semantics. To achieve the final multimodal features that integrate patient-specific prior knowledge with visual semantics.

6. The method for automatically generating radiology reports based on patient-specific prior knowledge according to claim 1, characterized in that, The specific method for step 5 is as follows: a. First stage: Prior-guided comparative pre-training, the process is as follows: First, the compressed clinical context representation is obtained using formulas (6) and (7). And spatiotemporal features that integrate clinical semantics This allows for the introduction of patient clinical context information into the visual encoding process, enabling prior knowledge to guide the extraction of spatiotemporal visual features. Then, splicing and After global average pooling and L2 normalization, global visual features are obtained. Next, the similarity (logits) between medical images and radiology reports is calculated. The calculation process can be expressed as follows: in, For global textual features of radiology reports, τ This is the temperature coefficient; similarly, it calculates the similarity between radiology reports and medical images. The calculation process can be expressed as follows: Considering that a patient may have multiple medical images during a single visit, all medical images and radiology reports associated with the same visit are considered as positive sample pairs, thus generating multiple positive sample pairs; therefore, a true matching label matrix is ​​defined. The calculation process for indicating the matching relationship between medical image and radiology report pairs can be expressed as follows: in, It is an indicator function, when y i =y j The value is 1 if the condition is met, and 0 otherwise. Finally, the loss function for prior-guided contrastive pre-training is the cross-entropy loss between q and p, and its calculation process can be expressed as follows: b. Second stage: Generation of prior-perceived radiological reports, the process is as follows: First, the compressed clinical context representation obtained in step 5 is... Spatiotemporal representation incorporating clinical semantics And the final multimodal features that integrate patient-specific prior knowledge and visual semantics. The sequences are concatenated to form the input sequence for the text generator; this input sequence helps to provide the text generator with rich and diverse semantic context, thereby guiding the text generator to output more personalized and clinically relevant radiology reports. Then, the radiology report generated by the text generator is minimized. Radiology reports written by actual radiologists i The cross-entropy loss function between the two is used to supervise the learning of the text generator, and its optimization objective can be expressed as: Where K is the maximum length of the text generator output. B represents all tokens generated up to the k-th token, and B is the batch size.

7. The automatic radiology report generation system according to any one of claims 1 to 6, comprising: The token missing encoding module is used in step 1 to represent missing patient clinical context information using a special token, so that the text encoder can uniformly process complete and incomplete input cases, in order to extract personalized text features. The spatiotemporal fusion network module is used to integrate the patient's previous medical images in step 2, construct the difference mapping relationship between the current medical images and previous medical images, and model the evolution trajectory of the disease to extract spatiotemporal visual features. Attention-enhanced hierarchical fusion network module is used to efficiently integrate multi-layer hidden states in the visual encoder in step 3, thereby facilitating the extraction of hierarchical visual features; A priori-aware progressive fusion network module is used to progressively integrate the personalized text features extracted in step 1 in step 4. Spatiotemporal visual features extracted in step 2 and the hierarchical visual features generated in step 3 A two-stage training module is used in step 5 to optimize the consistency of multimodal representation between medical images and radiology reports, as well as the quality of radiology report generation.

8. The radiology report automatic generation device according to any one of claims 1 to 6, comprising: Memory, used to store computer programs; A processor for executing the computer program to implement the method for automatically generating radiological reports based on patient-specific prior knowledge.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it is capable of automatically generating radiological reports according to any one of claims 1 to 6 based on the patient-specific prior knowledge method.