Chest radiograph report generation method and system based on cross-modal alignment and significant semantic region
By combining cross-modal alignment with salient semantic regions, and utilizing recognition networks and language generation models to generate chest X-ray reports, the problems of time-consuming and inaccurate report generation in existing technologies are solved, achieving efficient and accurate chest X-ray report generation.
Patent Information
- Application Number
- CN202510880363.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing technology for generating chest X-ray reports has problems such as time-consuming and labor-intensive manual writing by doctors, large individual differences, insufficient understanding of subtle image features by artificial intelligence methods, and insufficient fusion of multimodal information. These problems result in poor report accuracy and consistency, and inability to effectively utilize comprehensive patient information.
Using the method of cross-modal alignment and salient semantic regions, the recognition network processes image and text information to generate salient regions and salient maps, and uses the mask image modeling module and language generation model to generate accurate radiology reports.
It improves the efficiency and accuracy of chest X-ray report generation, can better capture subtle abnormalities, combine multimodal information to generate detailed medical reports, and improves the language and medical accuracy of the reports.
Smart Images

Figure CN120748604A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a chest X-ray report generation method and system based on cross-modal alignment and significant semantic regions. Background Art
[0002] In medicine, chest X-rays are a widely used diagnostic tool, and chest X-ray reports are crucial for doctors to accurately assess a patient's condition. With the advancement of medical imaging technology, the traditional practice of relying on doctors to manually write chest X-ray reports has gradually exposed numerous problems. For one thing, doctors face a massive daily workload, and manually writing reports consumes significant time and energy, leading to fatigue and compromising the accuracy and consistency of reports. Individual differences in the interpretation of chest X-ray images can lead to significant discrepancies in the descriptions of similar conditions, impacting the accuracy and standardization of diagnoses. For example, different doctors have varying criteria and methods for determining whether lung markings are thickened, or describing the size and morphology of nodules.
[0003] On the other hand, with the gradual application of artificial intelligence technology in the medical field, the use of computers to assist in the generation of chest X-ray reports has become a research hotspot. However, some existing chest X-ray report generation methods based on artificial intelligence often have problems with insufficient understanding of subtle features and complex semantics in images. For example, some methods have difficulty accurately identifying subtle abnormalities in early lung lesions, or have difficulties converting image features into accurate and comprehensive text reports, resulting in the generated reports being unable to provide doctors with sufficiently detailed and accurate diagnostic information. In addition, in actual clinical applications, chest X-ray reports not only need to accurately describe what is seen in the image, but also need to be combined with other clinical information such as the patient's medical history and symptoms. However, most current chest X-ray report generation technologies fail to fully consider the fusion of multimodal information and cannot effectively utilize the patient's comprehensive information to generate more accurate reports.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore includes information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0005] The purpose of this application is to provide a chest X-ray report generation method and system based on cross-modal alignment and significant semantic regions, which at least to some extent overcomes the problems existing in the prior art. The recognition network is used to process image and text information, and the significant regions and saliency maps of medical semantic information are generated through steps such as image block segmentation and feature mapping. The mask image modeling module performs mask probability adjustment and other processing on the chest X-ray image and the significant region image to generate reconstructed image block features for capturing subtle abnormalities. The language generation model generates the target radiology report based on the saliency map and image block features through training, matrix multiplication and other operations. The report quality assessment index value is calculated using natural language generation indicators and medical accuracy indicators to evaluate the effectiveness of report generation. The recognition network is trained and optimized by randomly masking image block features, completing and using a preset loss function. With the help of multiple modules working together, key information is extracted from images and text to generate a radiology report that accurately describes the medical observations therein.
[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0007] According to one aspect of the present application, a chest X-ray report generation method based on cross-modal alignment and salient semantic regions is provided, comprising: obtaining chest X-ray image information of a target patient and corresponding initial radiology report text information, and a radiology report generation model, wherein the radiology report generation model comprises a network model for identifying salient regions of medical semantic information, a mask image modeling module guided by the salient regions, and a language generation model guided by a saliency map; processing the chest X-ray image information of the target patient and the corresponding initial radiology report text information based on the network model for identifying salient regions of medical semantic information to generate salient regions and a saliency map of the medical semantic information; processing the chest X-ray image of the target patient and the salient region image of the medical semantic information based on the mask image modeling module guided by the saliency region to generate image block features reconstructed by the mask image; processing the saliency map and the image block features reconstructed by the mask image based on the language generation model guided by the saliency map to generate target radiology report information; processing the target radiology report information to generate an evaluation result of chest X-ray report generation.
[0008] Another aspect of the present application is a chest X-ray report generation device based on cross-modal alignment and salient semantic regions, characterized in that it includes: an acquisition module for acquiring chest X-ray image information of a target patient and corresponding initial radiology report text information, and a radiology report generation model, wherein the radiology report generation model includes a network model for identifying salient regions of medical semantic information, a mask image modeling module guided by the salient regions, and a language generation model guided by the salient map; a processing module for processing the chest X-ray image information of the target patient and the corresponding initial radiology report text information based on the network model for identifying salient regions of medical semantic information to generate salient regions and a salient map of medical semantic information; processing the chest X-ray image of the target patient and the salient region image of the medical semantic information based on the mask image modeling module guided by the salient regions to generate image block features reconstructed by the mask image; processing the saliency map and the image block features reconstructed by the mask image based on the language generation model guided by the saliency map to generate target radiology report information; processing the target radiology report information to generate an evaluation result of chest X-ray report generation.
[0009] According to another aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a second processor, the method for generating a chest X-ray report based on cross-modal alignment and significant semantic regions is implemented.
[0010] The present application provides a chest X-ray report generation method and system based on cross-modal alignment and significant semantic regions. The server obtains the chest X-ray image of the target patient, the initial radiology report text and the radiology report generation model (including the recognition network, the mask image modeling module, and the language generation model), uses the recognition network to process the image and text information, and generates the significant regions and significant maps of the medical semantic information through steps such as image block division and feature mapping.
[0011] The masked image modeling module performs mask probability adjustments on chest X-ray images and salient region images to generate reconstructed image block features for capturing subtle abnormalities. The language generation model generates the target radiology report through training, matrix multiplication, and other operations based on the saliency map and image block features. Report quality assessment indicators are calculated using natural language generation metrics and medical accuracy metrics to evaluate the effectiveness of report generation. The recognition network is trained and optimized through random masking of image block features, completion, and a preset loss function. This method leverages the collaborative work of multiple modules to extract key information from images and text, and has high application value in the field of medical imaging report generation.
[0012] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A flowchart of a chest X-ray report generation method based on cross-modal alignment and significant semantic regions provided by an embodiment of the present application is shown;
[0014] Figure 2 A schematic structural diagram of a chest X-ray report generation device based on cross-modal alignment and significant semantic regions provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0015] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0016] The following combination Figure 1 The following describes a chest X-ray report generation method based on cross-modal alignment and significant semantic regions according to an exemplary embodiment of the present application. In one embodiment, the present application also proposes a chest X-ray report generation method and system based on cross-modal alignment and significant semantic regions. Figure 1 The flowchart of a chest X-ray report generation method based on cross-modal alignment and significant semantic regions according to an embodiment of the present application is schematically shown. Figure 1 As shown, the method is applied to the server and includes:
[0017] S101, obtaining chest X-ray image information of a target patient and corresponding initial radiology report text information and a radiology report generation model.
[0018] In one embodiment, it is assumed that there is a patient whose chest X-ray image clearly shows the outlines and some details of organs such as the heart and lungs, and the initial radiology report text is "no obvious abnormalities in the heart and lungs, and the chest is symmetrical." This information will provide key data support for subsequent model training. Starting from the chest X-ray image information of the target patient, the image block features are randomly masked according to a preset ratio (such as 50%). When extracting image features, the VisionTransformer (ViT)
[181] structure is used as the image encoder to convert the input chest image I∈R H×W×C Divided into N v non-overlapping image patches These image patch features are randomly selected and masked to obtain randomly masked image patch features. This step aims to simulate information that may be missing in real-world scenarios, allowing the model to learn how to extract key features from partial information. The randomly masked image patch features are then supplemented based on the masked features. During the supplementation process, the supplemented image patches serve as important data for subsequent training. This operation allows the model to attempt to restore the original features of the image patches when faced with incomplete information, helping to improve the model's understanding and learning of image features.
[0019] This network model is a cross-modal learning model that aims to identify significant regions related to medical semantics by learning the association between images and text. In terms of image feature processing, the ViT structure is used as the image encoder to flatten the image blocks and map them into visual embeddings through linear transformation. The images are then input into the Transformer model together with the visual position embeddings to obtain image block features. In terms of text feature processing, BioClinicalBERT is used to tokenize radiology reports into subword tokens, which are then input into the Transformer model after linear transformation and position embedding to obtain text token features. Subsequently, the image block features and text token features are processed through fine-grained cross-modal semantic alignment (FCSA) and mapping-before-aggregation (MbA) strategies to construct a common semantic space.
[0020] The completed image blocks and the real image blocks are processed based on the preset loss function to generate a network model for identifying the salient areas of medical semantic information. The preset loss function is: Among them, x i and Represent the completed image block and the real image block respectively, N v Represents the number of image blocks. During training, the completed image blocks x i and real image patches Input into this loss function for calculation. The loss value obtained in each calculation reflects the gap between the current output of the model and the actual situation. The process of model training is to continuously adjust its own parameters so that the loss value gradually decreases. Through the backpropagation algorithm, the gradient of the loss function will be calculated, and then the parameters in the model such as ViT structure, linear transformation layer and Transformer model will be updated, so that the model can be optimized in the direction of more accurately restoring the image block features. For example, in each iteration, according to the gradient information, the linear transformation function f in the image encoder is adjusted. v The weights and Transformer model trans v Fine-tune the parameters in so that the image block features output by the model are closer to the features of the real image blocks.
[0021] Preset loss function Lv This loss function directly impacts the model's learning of image patch features. If the loss value consistently decreases, it indicates that the model is improving, better capturing detailed information in the image and, consequently, more accurately identifying salient regions of medical semantic information. Conversely, if the loss value does not decrease significantly during training or fluctuates more, this may indicate a training issue, such as an inappropriate learning rate or an inappropriate model structure, requiring adjustments to the training process. The results of this loss function are not only used to optimize the network model used to identify salient regions of medical semantic information, but also indirectly influence the performance of the entire radiology report generation model. Accurately identifying salient regions is fundamental to the effective functioning of the subsequent mask image modeling module, which is guided by the salient regions, and the language generation model, which is guided by the saliency map. If the network model can accurately identify salient regions through loss function optimization, the mask image modeling module will be better able to capture subtle anomalies, and the language generation model will be able to generate more accurate and realistic radiology reports.
[0022] The mask image modeling model guided by salient regions belongs to the image modeling module based on the Transformer structure, which is used to process chest X-ray images to capture subtle abnormalities and generate more representative image block features. The input chest X-ray image is processed using VisionTransformer (ViT) as the visual encoder. The image is divided into multiple non-overlapping image blocks, which are input into the Transformer model after linear transformation and position embedding to obtain image block embedding features. These features are used for mask image reconstruction tasks and also provide visual information for subsequent radiology report generation. During the processing, key parameters include the mask probability adjustment parameters of the image blocks. For the i-th image block The mask probability calculation formula is: Where I(.) is the indicator function. If the image block belongs to the salient region S, the mask probability increases (Set to 0.35). This method increases the probability of masking salient regions, helping the model better learn subtle anomaly features.
[0023] The language generation model guided by saliency map is a Transformer-based language generation model that aims to generate radiology reports containing sparse clinical semantics using saliency map and image patch features. The standard Transformer structure is used, with 3 layers and 8 attention heads in the experiment, and the dimension of the hidden state is set to 512. The model receives the saliency map And the image block features after mask image reconstruction As input. First, the saliency map M Sand image patch features E I Perform matrix multiplication and normalization to obtain the discriminant representation ω′=Norm(M S E I ), which captures the pathological clues of each chest radiograph image patch. ω' is then added as an additional input to the language model, and through a self-attention mechanism, salient semantics are gradually incorporated into the text report representation. After training, the model is able to generate target radiology report information based on this input information.
[0024] S102, based on the network model for identifying significant areas of medical semantic information, the chest X-ray image information of the target patient and the corresponding initial radiology report text information are processed to generate significant areas and a significant map of the medical semantic information.
[0025] In one embodiment, the chest X-ray image information of the target patient is divided and processed to generate several non-overlapping image blocks. All image blocks are mapped based on linear transformation to generate visual embedding features. The chest X-ray image information of the target patient is divided and divided into several non-overlapping image blocks. This is the basis for subsequent processing. Through this division, the overall information of the image can be decomposed into multiple local information units, which facilitates the model to perform feature extraction and analysis separately. After division, all image blocks are mapped based on linear transformation. Take the input chest image I∈R H×W×C For example, after linear transformation f v , mapping each image block into a series of visual embeddings, the formula is expressed as in Represents the [CLS] tag in the Transformer structure. This step converts the raw data of the image block into a feature vector form that the model can understand and process, preparing for subsequent feature extraction and model training.
[0026] Based on the network model for identifying the salient regions of medical semantic information, the visual embedding features and visual position embedding features are processed to generate image block features. The obtained visual embedding features are combined with the learnable visual position embedding Input together to the Transformer model trans v Through the multi-layer attention mechanism and nonlinear transformation of the Transformer model, these features are deeply processed and fused to generate image block features E containing rich image information. v , the calculation formula is Among them, E v is the image block feature, are several non-overlapping image blocks, is the visual embedding feature, Embed features for visual positions. These image block features not only contain the visual content information of the image block, but also incorporate position information, allowing the model to perceive the positional relationship of the image block in the entire image, helping to more accurately represent the image features.
[0027] The initial radiology report text information is tokenized to generate several subword tokens. All subword tokens are mapped based on linear transformation to generate text embedding features. The text embedding features and text position embedding features are processed based on a network model for identifying significant regions of medical semantic information to generate text token features. The initial radiology report text information is tokenized and BioClinicalBERT is used to tokenize the radiology report into a series of subword tokens. where N t Represents the number of subwords in the report, and V is the size of the vocabulary. Then, all subword tokens are mapped based on linear transformation, and the linear transformation f t Map subword tokens to text embeddings, and we get in Represents the [CLS] tag in the Transformer structure. This step converts text information into a vector form that the model can process, laying the foundation for subsequent extraction of semantic features of the text.
[0028] Embed the above text embedding features with the learnable text position embedding Input together to the Transformer model trans t After being processed by the Transformer model, a text word feature E containing text semantics and position information is generated. t , the formula is Among them, E t is the image block feature, is a number of subword lemmas, is the text embedding feature, Embedding features for text positions. These text word-meta features can accurately represent the semantic information in radiology reports and provide textual support for cross-modal learning and salient region identification.
[0029] The image block features and text word features are flattened and linearly projected to local visual embedding features and local text embedding features respectively, and then aggregated through the maximum pooling operation to generate image block global features and text word global features. v and text word features E t Flatten and linearly project to local visual embedding features and local text embedding features This step further adjusts the dimension and representation of the features through linear projection, making them more suitable for subsequent aggregation operations.
[0030] Then, the local features of the two modalities are aggregated into their respective global features through the MaxPooling operation. The formula is as follows: where f v ′ and f t ′ represents the image block tag feature mapping function and the text word tag feature mapping function respectively. The maximum pooling operation can retain the most significant feature information in each modality, and the global feature of the image block generated in this way is and global features of text terms Key information is highlighted, providing more representative feature representation for the subsequent construction of a common semantic space.
[0031] The global features of the image block and the global features of the text word are processed to generate the target semantic space. Next, the bidirectional cross-modal contrast loss L is used. bi To learn the common semantic space, the calculation formula is: L bi =λ v L v→t +λ t L t→v , where λ v and λ t Is the weight parameter of the cross-modal contrast loss in the two directions. Cross-modal contrast loss L in the image-text direction v→t The formula is: Cross-modal contrast loss L in text-image direction t→v It can be expressed as: Among them, L v→t is the transmembrane contrast loss in the image-text direction, L t→v is the transmembrane contrast loss in the text-image direction, N is the minimum batch size for training, Represents the cosine similarity between the global features of the image block and the global features of the text word, Represents the cosine similarity between the global features of the image block and the global features of the text word, Represents the cosine similarity between the global features of text words and the global features of image blocks, represents the cosine similarity between the global features of text words and image patches, and τ is the temperature parameter. By calculating the bidirectional cross-modal contrastive loss, the model can learn the similarity between images and text in a common semantic space, making the image patch features and text word features comparable in this space, thereby achieving cross-modal semantic alignment.
[0032] Based on the target semantic space, the global features and local features of the image block are processed to generate the salient regions and salient maps of medical semantic information. S Finally, since the aggregation operation is performed in the learned common semantic space, the visual image blocks that are most similar to the global features (top k%) can be regarded as containing parts with discriminative semantic information. These parts usually represent the content that distinguishes a specific chest X-ray image from other chest X-ray images, such as lesions, abnormal organs, surgical traces, and implanted devices. SISRNet uses cosine similarity to identify local features with discriminative semantics, and then reversely locates their corresponding local image blocks. Specifically, by comparing the cosine similarity of the local features of each image block with the global features, image blocks with higher similarity are found. The areas corresponding to these image blocks are semantically rich salient areas.
[0033] The calculated saliency map M S is a numerical vector reflecting the importance of each image block. Image blocks corresponding to locations with large values in this vector (i.e., high similarity to global features) are marked, and these marked image blocks constitute salient regions. These salient regions are of great significance in clinical diagnosis. They contain key medical information and are crucial for the subsequent generation of accurate radiology reports. Ultimately, salient regions and saliency maps of medical semantic information are generated. These results provide key input for subsequent mask image modeling and language generation models, enabling the models to pay more attention to important areas in the image when generating reports, thereby improving the accuracy and reliability of the reports.
[0034] S103: Processing the chest X-ray image of the target patient and the salient region image of the medical semantic information based on the mask image modeling module guided by the salient region, and generating image block features after mask image reconstruction.
[0035] In one embodiment, a mask probability adjustment process is performed on the chest X-ray image of the target patient based on a mask image modeling module guided by a salient region to generate a mask probability of the image block. A mask image modeling module guided by a salient region processes the salient region image of medical semantic information to generate an adjusted mask probability of the image block. By increasing the probability of masking the salient region, the model is helped to learn subtle abnormal features. In this process, a pre-trained salient region recognition network is used to input a chest X-ray image to obtain a salient region S with rich medical-related semantic content. For the i-th image block The mask probability adjustment formula is: Among them, I(.) is the indicator function, if the image block Belonging to the salient area S, its mask probability increases (set to 0.35); p i The initial mask probability follows a uniform distribution U(0,1). This method adjusts the mask probability of the chest X-ray image and generates the mask probability of the image block. Simultaneously, the image is processed for salient regions of medical semantic information to further determine the adjusted mask probability of the image block. This process, also based on the aforementioned formula, strengthens the masking operation on salient regions, allowing the model to focus more on feature learning in abnormal areas.
[0036] The image block features are processed based on the mask probability of the image block and the adjusted mask probability of the image block to generate a masked image block. Based on the mask probability of the image block and the adjusted mask probability of the image block, the image block features are processed according to the corresponding rules. Specifically, based on the mask probability determined above, a selective masking operation is performed on the input image block features, and the portion of the image block that meets the masking conditions is processed to obtain a masked image block. This process enables the model to simulate scenarios in which some information is missing in real-world situations, prompting the model to learn how to extract key features from the remaining information, especially focusing on areas that may contain subtle anomalies.
[0037] The mask image block and the real image block are processed to generate pixels of the reconstructed mask image block. The generated mask image block and the real image block are processed, and the image decoder is used to reconstruct the pixels of the mask image block through the mean square error (MSE) loss function. The calculation formula is Among them, L I To reconstruct the pixels of the mask image block, and z i Denote the masked image patch and the real image patch, respectively, and M is the number of masked image patches. By minimizing this loss function, the model continuously adjusts its parameters to make the pixels of the reconstructed image patch as close as possible to those of the real image patch, thereby improving the restoration of image features and laying the foundation for subsequent generation of more accurate image patch features.
[0038] S104 , processing the saliency map and the image block features reconstructed by the mask image based on the language generation model guided by the saliency map to generate target radiology report information.
[0039] In one embodiment, a language generation model guided by a saliency map is trained based on a preset loss function to generate a trained language generation model. The generation of radiology reports presents certain challenges. The sentences are semantically dispersed and mostly contain normal descriptions, making it difficult for the decoder to capture abnormal information. To solve this problem, SISRNet uses a Transformer-based language generation model and trains it using a preset loss function. The preset loss function is Among them, L Rrepresents the loss function used to train the language model for generating radiology reports, l and V represent the length of the radiology report and the vocabulary size, respectively. Represents the probability of selecting the jth word in the vocabulary at the i-th position in the generated report. During the training process, the model continuously adjusts its own parameters to minimize the loss function so that the predicted word element As close as possible to the real word y ij , thereby learning the ability to accurately generate radiology reports and ultimately generating a trained language generation model.
[0040] The saliency map and the image block features reconstructed by the mask image are matrix multiplied and normalized to generate discriminant representation information. The discriminant representation information is used to characterize the pathological clues of each chest X-ray image block, integrating the information of the saliency map and the image block features, and providing key semantic guidance for the subsequent generation of radiology reports in the language model. The saliency map obtained after the mask image modeling module guided by the saliency region is processed. And the image block features after mask image reconstruction First, perform matrix multiplication on the two, and then process them through the normalization layer to obtain the discriminant representation information ω′. The calculation formula is ω′=Norm(M s E1). The discriminative representation integrates information from the saliency map and image patch features, capturing pathological clues within each chest X-ray image patch. For example, in actual chest X-ray image analysis, this computational approach can combine the importance of different image regions (represented by the saliency map) with the characteristics of the image patches themselves, providing critical semantic guidance for subsequent radiology report generation within the language model, enabling the model to focus more on information from important regions when generating reports.
[0041] S105: Process the target radiology report information to generate an evaluation result of chest X-ray report generation.
[0042] In one embodiment, a report quality assessment index value is generated based on the target radiology report information and the reference real report information using natural language generation indicators and medical accuracy indicators. To comprehensively evaluate the quality of radiology reports generated by the SISRNet model, it is necessary to calculate and process the target radiology report information and the reference real report information using natural language generation indicators (NLG) and medical accuracy indicators (CE). Commonly used natural language generation indicators include BLEU, METEOR, and ROUGE, which are mainly used to evaluate the similarity between the generated radiology report and the reference real report at the linguistic level. Among them, BLEU is a precision-based metric that analyzes n-grams with a maximum length of 4 and calculates penalties or rewards based on how well the generated text matches the reference text in length, vocabulary choice, and order. METEOR is a precision-and-recall metric that extends the BLEU-1 metric. It measures the matching of unigrams in the generated text with those in the reference text based on exact form, stem form, and meaning. It calculates the harmonic mean of the precision and recall of the unigrams, with a bias towards recall, and also uses a multiplicative factor to reward unigrams that match the order of the reference text. ROUGE-L is a recall-based metric that measures the length of the longest common subsequence between two texts. It calculates the weighted harmonic mean of the precision and recall, with a bias towards recall. Its calculation does not require continuous matching between texts, but rather requires sequential matching of n-grams that reflect sentence-level word order. However, these NLG metrics primarily focus on lexical similarity and are difficult to reflect the medical accuracy of reports.
[0043] The medical accuracy metric, CE, focuses on the medical relevance and accuracy of the generated content. It is used to measure the accuracy of the generated radiology report in terms of medical facts, that is, whether the generated report correctly describes the key medical information contained in the input chest X-ray image. The rule-based tagger CheXpert is used to extract 14 tags related to chest pathology and supporting devices for the generated and reference reports. If the two reports are medically semantically equivalent, the CE score obtained by the CheXpert tagger will be high. The target radiology report information and the reference real report information are substituted into the above two types of indicators for calculation, thereby obtaining the report quality assessment index values. These index values comprehensively reflect the degree of difference between the generated report and the real report in terms of language and medical content.
[0044] The report quality assessment index values are analyzed and interpreted to generate evaluation results for chest X-ray report generation. The chest X-ray report generation evaluation results are used to characterize the language accuracy, medical accuracy, and overall quality level of reports generated by the radiology report generation model. Analysis of NLG index values can reveal the fluency of the generated report in terms of language expression, grammatical correctness, and the degree of vocabulary matching with the real report. For example, a high BLEU index score indicates that the generated report is more similar to the real report in terms of vocabulary selection and order; the METEOR index can more finely reflect the matching of unigrams based on various forms, and its score can reflect the accuracy of the generated report in vocabulary comprehension and application.
[0045] Analysis of the medical accuracy indicator CE can directly determine the accuracy of the generated report's description of medical information. A high CE score indicates that the generated report can accurately describe the key medical information in the chest X-ray image and has a high reliability in medical diagnosis. Conversely, a low CE score indicates that the report lacks medical accuracy. The combined analysis results of the NLG and CE indicators can comprehensively characterize the language accuracy, medical accuracy, and overall quality level of the reports generated by the radiology report generation model. If the various NLG indicators perform well in terms of language accuracy and the medical accuracy indicator CE also reaches a high level, then the report generated by the model can be considered to be of high quality. If one of the indicators performs poorly, such as the NLG indicator indicating problems with language expression or the CE indicator indicating inaccurate description of medical information, it means that the model has room for improvement in the corresponding aspect. The final evaluation results will provide an important basis for model optimization and improvement.
[0046] The server obtains the target patient's chest X-ray image, initial radiology report text, and radiology report generation model (including recognition network, mask image modeling module, and language generation model), uses the recognition network to process image and text information, and generates salient areas and salient maps of medical semantic information through steps such as image block segmentation and feature mapping.
[0047] The masked image modeling module performs mask probability adjustments on chest X-ray images and salient region images to generate reconstructed image block features for capturing subtle abnormalities. The language generation model generates the target radiology report through training, matrix multiplication, and other operations based on the saliency map and image block features. Report quality assessment indicators are calculated using natural language generation metrics and medical accuracy metrics to evaluate the effectiveness of report generation. The recognition network is trained and optimized through random masking of image block features, completion, and a preset loss function. This method leverages the collaborative work of multiple modules to extract key information from images and text, and has high application value in the field of medical imaging report generation.
[0048] In one embodiment, Figure 2 As shown, the present application also provides a chest X-ray report generation device based on cross-modal alignment and significant semantic regions, comprising:
[0049] Acquisition module 201 is used to acquire chest X-ray image information of a target patient, the corresponding initial radiology report text information, and a radiology report generation model, wherein the radiology report generation model includes a network model for identifying salient regions of medical semantic information, a mask image modeling module guided by the salient regions, and a language generation model guided by the saliency map;
[0050] The processing module 202 is used to process the chest X-ray image information of the target patient and the corresponding initial radiology report text information based on the network model for identifying the salient areas of medical semantic information to generate the salient areas and salient map of the medical semantic information; to process the chest X-ray image of the target patient and the salient area image of the medical semantic information based on the mask image modeling module guided by the salient areas to generate image block features reconstructed by the mask image; to process the salient map and the image block features reconstructed by the mask image based on the language generation model guided by the salient map to generate target radiology report information; and to process the target radiology report information to generate an evaluation result of chest X-ray report generation.
[0051] Each embodiment in this application is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the evaluation of the chest X-ray report generation method based on cross-modal alignment and significant semantic regions, the electronic device, the electronic device, and the readable storage medium embodiment, since they are basically similar to the above-mentioned chest X-ray report generation method based on cross-modal alignment and significant semantic regions, the description is relatively simple, and the relevant parts can be referred to the partial description of the above-mentioned chest X-ray report generation method based on cross-modal alignment and significant semantic regions.
Claims
1. A chest radiograph report generation method based on cross-modal alignment and salient semantic regions, characterized in that: include: Obtaining chest X-ray image information of the target patient, the corresponding initial radiology report text information, and a radiology report generation model, wherein the radiology report generation model includes a network model for identifying salient regions of medical semantic information, a mask image modeling module guided by the salient regions, and a language generation model guided by the saliency map; Based on the network model for identifying salient regions of medical semantic information, the target patient's chest X-ray image information and the corresponding initial radiology report text information are processed to generate salient regions and saliency maps of the medical semantic information; Based on the mask image modeling module guided by the salient region, the chest X-ray image of the target patient and the salient region image of the medical semantic information are processed to generate image block features after the mask image reconstruction; Based on the language generation model guided by the saliency map, the saliency map and the image block features reconstructed by the mask image are processed to generate the target radiology report information; The target radiology report information is processed to generate an evaluation result of chest X-ray report generation.
2. The method according to claim 1, wherein Based on the network model for identifying salient regions of medical semantic information, the target patient's chest X-ray image information and the corresponding initial radiology report text information are processed to generate salient regions and salient maps of medical semantic information, including: Divide and process the chest X-ray image information of the target patient to generate several non-overlapping image blocks; Map all image blocks based on linear transformation to generate visual embedding features; The visual embedding features and visual position embedding features are processed based on a network model for identifying salient regions of medical semantic information to generate image patch features. Tokenize the initial radiology report text information to generate several subword tokens; Map all subword tokens based on linear transformation to generate text embedding features; The text embedding features and text position embedding features are processed based on a network model for identifying significant regions of medical semantic information to generate text word features. The image block features and text word features are flattened and linearly projected into local visual embedding features and local text embedding features respectively, and then aggregated through the maximum pooling operation to generate image block global features and text word global features; Process the global features of image blocks and text words to generate the target semantic space; The global features and local features of the image blocks are processed based on the target semantic space to generate the salient regions and saliency maps of medical semantic information.
3. The method according to claim 2, wherein The target patient's chest X-ray image information and the corresponding initial radiology report text information are processed based on a network model for identifying salient regions of medical semantic information to generate salient regions and a salient map of the medical semantic information, further comprising: The calculation formula for generating image block features is: Among them, E v is the image block feature, are several non-overlapping image blocks, is the visual embedding feature, Embed features for visual positions; The calculation formula for generating text word features is: Among them, E t is the image block feature, is a number of subword lemmas, is the text embedding feature, Embed features for text positions; The calculation formula for generating the target semantic space is: Among them, λ v and λ t is the weight parameter of the transmembrane state contrast loss in two directions, L v→t is the transmembrane contrast loss in the image-text direction, L t→v is the transmembrane contrast loss in the text-image direction, N is the minimum batch size for training, Represents the cosine similarity between the global features of the image block and the global features of the text word, Represents the cosine similarity between the global features of the image block and the global features of the text word, Represents the cosine similarity between the global features of the text word and the global features of the image block, represents the cosine similarity between the global features of the text word and the global features of the image block, and τ is the temperature parameter.
4. The method according to claim 2, wherein The mask image modeling module guided by the salient regions processes the chest X-ray image of the target patient and the salient region image of the medical semantic information to generate image block features after mask image reconstruction, including: Based on the mask image modeling module guided by the salient region, the mask probability adjustment processing is performed on the chest X-ray image of the target patient to generate the mask probability of the image block; The salient region image of medical semantic information is processed based on the salient region guided mask image modeling module to generate the mask probability after image block adjustment; Processing the image block features based on the mask probability of the image block and the mask probability after adjustment of the image block to generate a mask image block; Processing the mask image block and the real image block to generate pixels of the reconstructed mask image block; Generate features of the image block after mask image reconstruction based on pixels of the reconstructed mask image block; The calculation formula for generating the pixels of the reconstructed mask image block is: Among them, L I To reconstruct the pixels of the mask image block, and z i represent the mask image blocks and the real image blocks respectively, and M is the number of mask image blocks.
5. The method according to claim 4, wherein The language generation model guided by the saliency map processes the saliency map and the image patch features reconstructed from the mask image to generate the target radiology report information, including: Training the language generation model guided by the saliency map based on a preset loss function to generate a trained language generation model; Matrix multiplication and normalization are performed on the saliency map and the image patch features reconstructed from the mask image to generate discriminant representation information. This discriminant representation information is used to characterize the pathological clues of each chest X-ray image patch. It integrates the information of the saliency map and image patch features, providing key semantic guidance for the subsequent generation of radiology reports in the language model. The discriminant representation information is processed based on the trained language generation model to generate target radiology report information; The calculation formula for generating discriminant representation information is: ω′=Norm(M s E1); The default loss function is: L R represents the loss function used to train the language model for generating radiology reports, l and V represent the length of the radiology report and the vocabulary size, respectively. Represents the probability of selecting the jth word in the vocabulary at the i-th position in the generated report.
6. The method according to claim 1, wherein Process the target radiology report information to generate chest X-ray report evaluation results, including: Based on the target radiology report information and the reference real report information, the report quality assessment index value is generated by calculating and processing the natural language generation index and the medical accuracy index; The report quality assessment index values are analyzed and interpreted to generate the assessment results of chest X-ray report generation. The assessment results of chest X-ray report generation are used to characterize the language accuracy, medical accuracy and overall quality level of the reports generated by the radiology report generation model.
7. The method according to claim 1, wherein Obtain a network model for identifying salient regions of medical semantic information, including: Obtaining features of randomly masked image blocks of a preset proportion; Completing the features of the randomly masked image blocks based on the mask mark features to generate the completed image blocks; The completed image blocks and the real image blocks are processed based on a preset loss function to generate a network model for identifying salient areas of medical semantic information; The default loss function is: Among them, x i and Represent the completed image block and the real image block respectively, N v Represents the number of image blocks.
8. A chest X-ray report generation device based on cross-modal alignment and significant semantic regions, characterized in that: The device comprises: An acquisition module, configured to acquire chest X-ray image information of a target patient, the corresponding initial radiology report text information, and a radiology report generation model. The radiology report generation model comprises a network model for identifying salient regions of medical semantic information, a mask image modeling module guided by salient regions, and a language generation model guided by saliency maps. A processing module is used to process the chest X-ray image information of the target patient and the corresponding initial radiology report text information based on a network model for identifying salient regions of medical semantic information to generate salient regions and a salient map of the medical semantic information; to process the chest X-ray image of the target patient and the salient region image of the medical semantic information based on a mask image modeling module guided by the salient region to generate image block features reconstructed by the mask image; to process the salient map and the image block features reconstructed by the mask image based on a language generation model guided by the salient map to generate target radiology report information; and to process the target radiology report information to generate an evaluation result of chest X-ray report generation.
9. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the chest X-ray report generation method based on cross-modal alignment and significant semantic areas as described in any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the second processor, the method for generating a chest X-ray report based on cross-modal alignment and significant semantic regions according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Semantics-based medical imaging report template generation method
CN109545302A
Chest X-ray image report generation method based on cross-modal network
CN117558394A
Chest radiograph report generation method based on knowledge enhancement cross-modal semantic association
CN118553369A
Ultrasonic image pre-training method based on vision-language multi-mode contrast learning
CN118821900A
Radiology clinical medical image report generation method and system
CN119028509A