Radiology report generation method and system based on visual collaborative enhancement and cross-modal fusion network

By using a visual collaboration enhancement module and a cross-modal fusion engine, the problems of unbalanced data distribution and cross-modal information fusion in radiology report generation were solved, achieving efficient identification of abnormal regions and semantic-level feature alignment of reports, and generating high-quality radiology reports.

CN120853786APending Publication Date: 2025-10-28DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510732129.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing radiology report generation tasks suffer from imbalanced data distribution and cross-modal information fusion challenges, resulting in poor performance of models in identifying abnormal regions and aligning image and text features.

Method used

A visual collaborative enhancement module is used to model from both global and local perspectives. An abnormal feature is identified by combining a recalibration attention mechanism. A multi-level fusion of visual and textual information is achieved through a cross-modal information fusion device. An accurate radiological report is generated using a Transformer encoder and decoder.

Benefits of technology

It alleviates attention bias caused by unbalanced data distribution, achieves semantic-level feature alignment and refinement, and generates more accurate and consistent radiological reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853786A_ABST
    Figure CN120853786A_ABST
Patent Text Reader

Abstract

The invention discloses a radiology report generation method and system based on visual collaborative enhancement and a cross-modal fusion network, and belongs to the technical field of natural language processing. According to the invention, a visual collaborative enhancement module is designed for modeling visual features from global and local perspectives to enhance the recognition of abnormal lesions in a radiology image, so that the attention deviation of an abnormal region caused by unbalanced data distribution is relieved. Meanwhile, a cross-modal information fusion device is provided, the module utilizes a novel double cross-modal communication component to promote multi-level fusion of visual and text information, the problem of modal isomerism is solved, and semantic-level feature alignment and refinement are achieved. According to the method, the problem that a model cannot capture key focus features due to unbalanced data distribution in a radiology image is solved, and the problem that effective alignment and fusion are difficult due to the fact that feature spaces of different modal information of image and text information are different is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of natural language processing, and relates to a radiology report generation technology, particularly a radiology report generation method and system based on visual collaborative enhancement and cross-modal fusion networks. Background Technology

[0002] Radiology reports serve as a bridge between medical imaging and clinical decision-making, and their accuracy and timeliness directly impact patient treatment outcomes. Traditionally, the generation of radiology reports has relied on the expertise and experience of radiologists, a labor-intensive process susceptible to human factors such as fatigue and subjective judgment. Furthermore, advancements in medical imaging technology, leading to increased image resolution and complexity, place even greater demands on radiologists' analytical abilities.

[0003] With the continuous development of natural language processing and deep learning technologies, deep neural network models have been widely applied to the task of automating radiology report generation. However, this automation process still faces many challenges, mainly in the following two aspects:

[0004] (1) Imbalanced Data Distribution: The imbalanced data distribution problem mainly stems from the fact that in medical images, the number of normal samples is usually far greater than that of abnormal samples, while abnormal regions often occupy only a small portion of the image and may exhibit blurry or indistinct features. This poses a significant challenge to models in accurately identifying and describing abnormal regions. Previous research methods typically segment the entire image into multiple patches and utilize convolutional neural networks to extract visual feature representations of the image. However, since the distribution of abnormal regions is usually irregular, this method has significant limitations in identifying abnormal lesions in images.

[0005] (2) Cross-modal information fusion problem: The radiology report generation task involves two modalities: images and text. Alignment between images and text is crucial for effective cross-modal interaction. However, due to the differences in feature spaces between images and text, and the complex correspondence between them, achieving accurate alignment remains a challenge. Specifically, medical images typically contain rich visual information, such as the location, shape, and texture features of lesions, while text reports are abstractions and summaries of this visual information. Since images and text belong to different modalities, their feature representations differ fundamentally, making direct alignment of information from these two modalities highly complex. Summary of the Invention

[0006] To address the aforementioned problems, this invention proposes a method and system for generating radiological reports based on visual collaborative enhancement and cross-modal fusion networks. The aim is to generate corresponding descriptive text reports from radiological images. This invention designs a visual collaborative enhancement module to model visual features from both global and local perspectives, enhancing the identification of abnormal lesions in radiological images and mitigating attentional bias towards abnormal regions caused by data imbalance. Simultaneously, a radiological report generation module is developed. This module primarily designs a cross-modal information fusion device. This fusion device utilizes a dual cross-modal communication component to promote multi-level fusion of visual and textual information, solving the modal heterogeneity problem and achieving semantic-level feature alignment and refinement.

[0007] The technical solution adopted in the present invention is as follows:

[0008] On the one hand, this invention provides a method for generating radiological reports based on visual collaborative enhancement and cross-modal fusion networks, the steps of which are as follows:

[0009] Preprocessing of radiological images and text reports in electronic medical record datasets;

[0010] Visual features are obtained from preprocessed radiographic images using a visual encoder based on a convolutional neural network; and the preprocessed text report is vectorized using word2vec to obtain text report features.

[0011] The visual features are fed into the visual co-enhancement module, which includes a visual encoder, a global pooling layer, a linear layer, batch normalization, a Tanh activation function, a Dropout layer, a gating unit, and a softmax. The visual co-enhancement module identifies fine-grained visual features related to abnormal features in medical images to obtain visual co-enhancement features.

[0012] The text report features and the visual collaborative features are input into the radiology report generation module, which includes a multi-head masking attention layer, a normalization layer, a multi-head attention layer, a cross-modal information integrator, a feedforward layer, and a linear layer. Cross-modal feature alignment is achieved by interacting with information representations between different modalities, thereby generating accurate and consistent radiology reports.

[0013] Furthermore, the radiological images and text reports in the electronic medical record dataset are preprocessed, including: converting the radiological images and text reports in the electronic medical record dataset into embedded representations that can be learned by a deep neural network model; standardizing image size and normalizing image grayscale; removing stop words from the text reports and constructing a vocabulary.

[0014] Furthermore, the visual encoder based on the convolutional neural network is a pre-trained convolutional neural network ResNet-101.

[0015] Furthermore, visual features are obtained from the preprocessed radiographic images using a convolutional neural network-based visual encoder, including:

[0016] Given a batch of radiological images Where, N I This represents the number of medical images; visual characteristics are denoted as... The calculation process is as follows:

[0017]

[0018] Among them, v i ∈R H×W×C This represents the extracted block-level visual feature representation, where H, W, and C correspond to the height, width, and number of channels of the radiographic image, respectively, and f RIF This represents a visual encoder.

[0019] Furthermore, by identifying fine-grained visual features related to abnormal features in medical images through the visual collaboration enhancement module, visual collaboration features are obtained, including:

[0020] After performing preliminary visual extraction on multiple visual blocks to obtain visual feature representations V, average pooling is used to capture global visual feature information, denoted as V.

[0021]

[0022] Among them, b i ∈R HW×C The aggregated visual feature vector b i The dimension is HW×C, and AvgPool(·) refers to the average pooling operation;

[0023] A recalibrated attention mechanism is employed to identify local and global feature representations in a fine-grained manner. This recalibrated attention mechanism combines convolutional operations and bottom-up attention, and the specific process is as follows:

[0024] L=Dropout(Tanh(BatchNorm(Linear(V))))#(3)

[0025] G=Dropout(Tanh(BatchNorm(Linear(B))))#(4)

[0026] Where L and G represent fine-grained representations of local and global visual features, respectively, Tanh(·) represents the activation function, Linear(·) and BatchNorm(·) represent linear operation and batch normalization operation, and Dropout(·) refers to the regularization operation;

[0027] A bottom-up attention mechanism is used to integrate fine-grained global and local visual features, thereby capturing the local feature representations most relevant to the global features. Attention weights are then calculated based on these local feature representations to focus on anomalous regions.

[0028] S=L⊙G#(5)

[0029] W = softmax(GU(S))#(6)

[0030] Where S is the local visual feature score relative to the global visual features, W represents the weight contribution of each local visual feature to the global visual features, and GU(·) represents the gating unit.

[0031] Local representation of the entire image Img v i Encoding is performed through an attention mechanism, v g It is aggregated from the detected regions, and the process is as follows:

[0032]

[0033] Among them, V g This represents the output of the visual collaboration enhancement module, W. i It is the weight contribution parameter of each local visual feature to the global visual feature, and (||·||2) refers to the L2 regularization function.

[0034] Furthermore, the radiology cross-modal integrator includes an attention mechanism and a dual cross-modal communication component. The dual cross-modal communication component includes a linear layer, a GELU activation function, a normalization layer, and a regularization function. The Transformer decoder includes a masked multi-head self-attention layer, a multi-head attention layer, a feedforward neural network, residual connections, and layer normalization. The text report features and the visual collaborative features are input into the radiology report generation module. Cross-modal feature alignment is achieved through the interaction of information representations between different modalities, generating accurate and consistent radiology reports, including:

[0035] Initialize a shared memory unit Where η is the number of storage units; an attention mechanism is applied to embed text features. and visual feature embedding V g Cross-modal feature communication is possible; these features are mapped to a unified shared representation space, producing a text representation M.g and visual representation M i The calculation process is defined as follows:

[0036]

[0037] Among them, W v W k W q It is a trainable parameter, q r This represents the query value of the text feature, k. m This represents the key value of the memory unit, q. v This represents the query value for visual features;

[0038] Since the feature information in the shared representation space is continuously updated during the loop, F(·) is introduced to represent the loop process:

[0039] X = F(M; M) g M i )#(11)

[0040] Where X represents the final feature representation containing both visual and textual feature information;

[0041] The implementation process of the dual cross-modal communication component is as follows:

[0042] Z *,i =X *,i +W2Φ(W1(Norm(X *,i )))#(12)

[0043] U *,i =Z j,* +W4Φ(W3(Norm(Z j,* )))#(13)

[0044] Where Z and U represent the updated feature matrices after horizontal and vertical interactions, respectively; i ranges from the first row to the dim row of the matrix, and j ranges from the first column to the dim column of the matrix; Norm(·) represents the LayerNorm operation along a specific dimension, Φ(·) represents the GELU activation function; and the weight matrices W1, W2, W3, and W4 are learnable parameters.

[0045] Introducing the Transformer encoder, visual features generated by the visual co-enhancement module. The input is fed into the encoder; the encoder uses a multi-head self-attention mechanism to input V. g Learn to obtain feature representations that include global contextual information.

[0046]

[0047] Among them, f en (·) represents the encoder, and H is the output of the Transformer encoder;

[0048] Y i =RCMI(V g ;R;M)#(15)

[0049] Among them, Y I This represents the multi-level cross-modal feature representation obtained through a cross-modal information fusion unit;

[0050] Since recursive relation generation is an autoregressive process, the decoder is based on previously generated words. The next word is generated sequentially from the encoder's output. Until a complete radiology report is generated; the specific generation process is as follows:

[0051]

[0052] in, f represents the final generated report. de (·) indicates the decoder.

[0053] In another aspect, the present invention also provides a radiology report generation system based on visual collaborative enhancement and cross-modal fusion networks, comprising:

[0054] The preprocessing module preprocesses the radiological images and text reports in the electronic medical record dataset;

[0055] The feature representation module uses a visual encoder based on a convolutional neural network to obtain visual features from the preprocessed radiographic image; and it uses word2vec to vectorize the preprocessed text report to obtain text report features.

[0056] The visual collaboration enhancement module identifies fine-grained visual features related to abnormal features in medical images based on the visual features, and obtains visual collaboration features; the visual collaboration enhancement module includes a visual encoder, a global pooling layer, a linear layer, batch normalization, a Tanh activation function, a Dropout layer, a gating unit, and a softmax;

[0057] The radiology report generation module combines the text report features and the visual collaborative features to achieve cross-modal feature alignment through the interaction of information representations between different modalities, thereby generating accurate and consistent radiology reports. The radiology report generation module includes a multi-head mask attention layer, a normalization layer, a cross-modal information integrator, a feedforward layer, and a linear layer.

[0058] Compared with existing technologies, this invention alleviates two challenges in most current radiology report generation models: imbalanced data distribution and insufficient cross-modal information fusion. The beneficial effects of this invention are:

[0059] 1) This invention models visual features from both global and local perspectives using a visual collaborative enhancement module to enhance the identification of abnormal lesions in radiological images, thereby alleviating attention bias in abnormal areas caused by unbalanced data distribution.

[0060] 2) This invention utilizes a novel cross-modal information fusion device through a radiology report generation module to promote multi-level fusion of visual and textual information, solve the problem of modal heterogeneity, and achieve semantic-level feature alignment and refinement. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 This is a schematic diagram of a radiology report generation method based on visual collaborative enhancement and cross-modal fusion network in an embodiment of the present invention;

[0063] Figure 2 This is a schematic diagram of the neural network in an embodiment of the present invention. Detailed Implementation

[0064] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0065] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0066] The technical method of this invention involves inputting radiological images into a pre-trained convolutional neural network (ResNet-101) to capture preliminary visual feature representations, and preprocessing and vectorizing text reports. A visual co-enhancement module is designed to extract fine-grained visual feature representations from medical images and identify abnormal lesion regions. This module first collaboratively learns the image feature representations from both global and local perspectives. The global and local perspectives are captured respectively by designing a recalibrated attention mechanism that integrates convolutional operations and bottom-up attention. Furthermore, to align feature information across different modalities, this invention develops a radiological report generation module that utilizes a novel cross-modal information fusion engine to promote multi-level fusion of visual and textual information, addressing the heterogeneity of data from different modalities and thus better integrating visual and textual information.

[0067] like Figure 1 , 2 As shown in the figure, a radiology report generation method based on visual collaborative enhancement and cross-modal fusion network in an embodiment of the present invention specifically includes the following steps:

[0068] S1: Preprocess the radiological images and text reports in the electronic medical record dataset, and then embed the preprocessed data into vectorized form.

[0069] The preprocessed data includes images and text. Image preprocessing includes operations such as standardization and resizing, and grayscale normalization. Text preprocessing includes cleaning and removal of stop words.

[0070] Vectorized embedding of text refers to initial vectorization, which transforms preprocessed text into a numerical representation.

[0071] S2: Input the vectorized radiographic image into ResNet-101 to extract abstract visual features; and vectorize the text report using word2vec.

[0072] Among them, word vectorization refers to generating semantic word vectors through word2vec to capture contextual relationships.

[0073] First, given a batch of radiomedical images Where, N I This represents the number of medical images. Using ResNet-101 as the visual encoder, visual features are learned from radiological images, denoted as... The calculation process is as follows:

[0074]

[0075] Among them, v i ∈R H×W×C This represents the extracted block-level visual feature representation, where H, W, and C correspond to the height, width, and number of channels of the radiographic image, respectively. RIF This represents a visual encoder.

[0076] S3: Obtain the visual features The data is fed into the visual collaboration enhancement module to identify fine-grained visual features in medical images that are associated with abnormal features.

[0077] like Figure 2 As shown, the visual co-enhancement module includes a visual encoder, a global pooling layer, a linear layer, batch normalization, a Tanh activation function, a Dropout layer, a gating unit, and a softmax function. This module learns feature representations of the image from both global and local perspectives and enables them to work collaboratively. More specifically, this module aims to improve the accuracy of anomaly detection by comprehensively understanding the content of the image through the integration of local and global features. Local features focus on the details of the image, while global features provide overall background information.

[0078] After preliminary visual feature extraction (V) from multiple visual blocks, average pooling is used to capture global visual feature information, denoted as V. The calculation process is as follows:

[0079]

[0080] Among them, b i ∈R HW×C The aggregated visual feature vector b i The dimension is HW×C, and AvgPool(·) refers to the average pooling operation.

[0081] Because ResNet-101 uses a convolutional neural network (CNN) to capture visual features, its large receptive field may overlook fine-grained details in the image, reducing the model's ability to identify subtle lesions. Therefore, this invention designs a recalibrated attention mechanism that combines convolutional operations with bottom-up attention to identify local and global feature representations with fine-grained detail. The specific process is as follows:

[0082] L=Dropout(Tanh(BatchNorm(Linear(V))))#(3)

[0083] G=Dropout(Tanh(BatchNorm(Linear(B))))#(4)

[0084] Where L and G represent fine-grained representations of local and global visual features, respectively, Tanh(·) represents the activation function, Linear(·) and BatchNorm(·) represent linear operation and batch normalization operation, and Dropout(·) refers to the regularization operation.

[0085] Then, a bottom-up attention mechanism is used to integrate fine-grained global and local visual features, thereby capturing local feature representations most relevant to the global features. Based on these representations, we compute attention weights to focus on anomalous regions, as follows:

[0086] s=L⊙G#(5)

[0087] W = softmax(GU(S))#(6)

[0088] Where S is the local visual feature score relative to the global visual features, and W represents the weight contribution of each local visual feature to the global visual features. GU(·) represents a gating unit, which filters out irrelevant information based on the salience or importance of the input information, thus retaining only the key visual information.

[0089] Finally, the local representation v of the entire image Img. i Encoding is performed through an attention mechanism. Specifically, v g It is aggregated from the detected regions, and the process is as follows:

[0090]

[0091] Among them, V g This represents the output of the visual collaboration enhancement module, W. i It is the weight contribution parameter of each local visual feature to the global visual feature, and (||·||2) refers to the L2 regularization function.

[0092] S4: Input text report features and visual collaborative features into the radiology report generation module. By using a cross-modal information fusion device to interact with information representations between different modalities, cross-modal feature alignment is achieved, thereby generating accurate and consistent radiology reports.

[0093] In radiology report generation, effectively fusing visual and textual features is crucial for producing high-quality reports. However, their representations differ significantly: visual features emphasize spatial structure and detail, while textual features convey semantic content. These modality-specific differences can lead to the loss of critical information when models try to capture the relationships between multimodal features. Therefore, cross-modal feature interaction and alignment is a key challenge in generating accurate radiology reports.

[0094] like Figure 2 As shown, the radiology report generation module includes a multi-head masking attention layer, a normalization layer, a multi-head attention layer, a cross-modal information integrator, a feedforward layer, and a linear layer. In this embodiment, the radiology report generation module incorporates a radiology cross-modal integrator, which includes an attention mechanism and a dual cross-modal communication component. The dual cross-modal communication component includes a linear layer, a GELU activation function, a normalization layer, and a regularization function. This module facilitates semantic fusion of visual and textual features at multiple levels, enhances modal synergy, and ensures alignment and refinement of cross-modal information at the semantic level.

[0095] First, we initialize a shared memory unit. Where η is the number of storage units. Then, an attention mechanism is applied to embed the text features. and visual feature embedding V g Cross-modal feature communication is possible. These features are mapped to a unified shared representation space, producing a text representation M. g and visual representation M i The calculation process is defined as follows:

[0096]

[0097] Among them, W v W k W q It is a trainable parameter, q r This represents the query value of the text feature, k. m This represents the key value of the memory unit, q. v This represents the query value for visual features.

[0098] Since the feature information in the shared representation space is continuously updated during the loop, we introduce F(·) to represent the loop process:

[0099] X = F(M; M) g M i )#(11)

[0100] Where X represents the final feature representation that includes visual and textual feature information.

[0101] This embodiment then designs a dual cross-modal communication component to enhance communication and interaction between different modalities, particularly addressing the problem of losing critical cross-modal information during cyclic updates. The dual cross-modal communication component employs horizontal and vertical communication interaction mechanisms to facilitate information exchange and ensure consistency between modalities.

[0102] The dual cross-modal communication component processes X through horizontal and vertical communication interaction mechanisms to achieve information exchange between modalities and avoid information loss or inconsistency. In horizontal interaction, key information in visual features (such as lesion size and location) undergoes deep interaction with corresponding text features; in vertical interaction, information from different modalities is further aligned in a shared representation space, enhancing their relevance. The specific implementation process of the dual cross-modal communication component is as follows:

[0103] Z *,i =X *,i +W2Φ(W1(Norm(X *,i 000#(12)

[0104] U *,i =Z j,* +W4Φ(W3(Norm(Z j,* )))#(13)

[0105] Here, Z and U represent the updated feature matrices after horizontal and vertical interactions, respectively. The range of i is from row 1 to row dim, and the range of j is from column 1 to column dim. Norm(·) represents the LayerNorm operation along a specific dimension, and Φ(·) represents the GELU activation function. The weight matrices W1, W2, W3, and W4 are learnable parameters.

[0106] This embodiment also introduces a Transformer encoder, and the Transformer decoder includes a masked multi-head self-attention layer, a multi-head attention layer, a feedforward neural network, residual connections, and layer normalization; visual features generated by the visual co-enhancement module. It is input into the encoder. The encoder uses a multi-head self-attention mechanism to input V. g Learn to obtain feature representations that include global contextual information.

[0107]

[0108] Among them, f en (·) represents the encoder, and H is the output of the Transformer encoder.

[0109] Y I =RCMI(V g ;R;M)#(15)

[0110] Among them, Y I This represents the multi-level cross-modal feature representation obtained through the Radiological Cross-Modal Information Fusion (RCMI) fusion device.

[0111] Since the recursive relation generation report is an autoregressive process, the decoder is based on previously generated words.

[0112] The next word is generated sequentially from the encoder's output. This continues until a complete radiology report is generated. The specific generation process is as follows:

[0113]

[0114] in, f represents the final generated report. de (·) represents the decoder. Therefore, in order to generate a complete report, the above process will be repeated iteratively until the generation is complete.

[0115] In the above embodiments, the present invention models visual features from both global and local perspectives through a visual collaboration enhancement module to enhance the identification of abnormal lesions in radiological images, thereby alleviating attention bias towards abnormal regions caused by data imbalance. Simultaneously, the present invention utilizes a novel cross-modal information fusion processor in the radiological report generation module to promote multi-level fusion of visual and textual information, solving the modal heterogeneity problem and achieving semantic-level feature alignment and refinement.

[0116] In another embodiment, corresponding to the above-described radiology report generation method, the present invention also provides a radiology report generation system based on visual collaborative enhancement and cross-modal fusion networks, comprising:

[0117] The preprocessing module preprocesses the radiological images and text reports in the electronic medical record dataset;

[0118] The feature representation module uses a visual encoder based on a convolutional neural network to obtain visual features from the preprocessed radiographic image; and it uses word2vec to vectorize the preprocessed text report to obtain text report features.

[0119] The visual collaboration enhancement module identifies fine-grained visual features related to abnormal features in medical images based on the visual features, and obtains visual collaboration features; the visual collaboration enhancement module includes a visual encoder, a global pooling layer, a linear layer, batch normalization, a Tanh activation function, a Dropout layer, a gating unit, and a softmax;

[0120] The radiology report generation module combines the text report features and the visual collaborative features to achieve cross-modal feature alignment through the interaction of information representations between different modalities, thereby generating accurate and consistent radiology reports. The radiology report generation module includes a multi-head mask attention layer, a normalization layer, a cross-modal information integrator, a feedforward layer, and a linear layer.

[0121] In the above embodiments, the present invention models visual features from both global and local perspectives through a visual collaboration enhancement module to enhance the identification of abnormal lesions in radiological images, thereby alleviating attention bias towards abnormal regions caused by data imbalance. Simultaneously, the present invention utilizes a novel cross-modal information fusion processor in the radiological report generation module to promote multi-level fusion of visual and textual information, solving the modal heterogeneity problem and achieving semantic-level feature alignment and refinement.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating radiological reports based on visual collaborative enhancement and cross-modal fusion networks, characterized in that, The steps are as follows: Preprocessing of radiological images and text reports in electronic medical record datasets; Visual features are obtained from preprocessed radiographic images using a visual encoder based on a convolutional neural network; Furthermore, the preprocessed text report is word vectorized using word2vec to obtain the text report features; The visual features are fed into the visual co-enhancement module, which includes a visual encoder, a global pooling layer, a linear layer, batch normalization, a Tanh activation function, a Dropout layer, a gating unit, and a softmax. The visual co-enhancement module identifies fine-grained visual features related to abnormal features in medical images to obtain visual co-enhancement features. The text report features and the visual collaborative features are input into the radiology report generation module, which includes a multi-head masking attention layer, a normalization layer, a multi-head attention layer, a cross-modal information integrator, a feedforward layer, and a linear layer. Cross-modal feature alignment is achieved by interacting with information representations between different modalities, thereby generating accurate and consistent radiology reports.

2. The method according to claim 1, characterized in that, The radiological images and text reports in the electronic medical record dataset are preprocessed, including: converting the radiological images and text reports in the electronic medical record dataset into embedded representations that can be learned by a deep neural network model; standardizing image size and normalizing image grayscale; removing stop words from the text reports and constructing a vocabulary.

3. The method according to claim 1, characterized in that, The visual encoder based on a convolutional neural network is a pre-trained convolutional neural network, ResNet-101.

4. The method according to claim 3, characterized in that, Visual features are obtained from preprocessed radiographic images using a convolutional neural network-based visual encoder, including: Given a batch of radiological images Where, N I This represents the number of medical images; visual characteristics are denoted as... The calculation process is as follows: Among them, v i ∈R H×W×C This represents the extracted block-level visual feature representation, where H, W, and C correspond to the height, width, and number of channels of the radiographic image, respectively, and f RIF This represents a visual encoder.

5. The method according to claim 1, characterized in that, The visual collaboration enhancement module identifies fine-grained visual features related to abnormal features in medical images to obtain visual collaboration features, including: After performing preliminary visual extraction on multiple visual blocks to obtain visual feature representations V, average pooling is used to capture global visual feature information, denoted as V. Among them, b i ∈R HW×C The aggregated visual feature vector b i The dimension is HW×C, and AvgPool(·) refers to the average pooling operation; A recalibrated attention mechanism is employed to identify local and global feature representations in a fine-grained manner. This recalibrated attention mechanism combines convolutional operations and bottom-up attention, and the specific process is as follows: L=Dropout(Tanh(BatchNorm(Linear(V))))#(3) G=Dropout(Tanh(BatchNorm(Linear(B))))#(4) Where L and G represent fine-grained representations of local and global visual features, respectively, Tanh(·) represents the activation function, Linear(·) and BatchNorm(·) represent linear operation and batch normalization operation, and Dropout(·) refers to the regularization operation; A bottom-up attention mechanism is used to integrate fine-grained global and local visual features, thereby capturing the local feature representations most relevant to the global features. Attention weights are then calculated based on these local feature representations to focus on anomalous regions. S=L⊙G#(5) W = softmax(GU(S))#(6) Where S is the local visual feature score relative to the global visual features, W represents the weight contribution of each local visual feature to the global visual features, and GU(·) represents the gating unit. Local representation of the entire image Img v i Encoding is performed through an attention mechanism, v g It is aggregated from the detected regions, and the process is as follows: Among them, V g This represents the output of the visual collaboration enhancement module, W. i It is the weight contribution parameter of each local visual feature to the global visual feature, and (||·||2) refers to the L2 regularization function.

6. The method according to claim 1, characterized in that, The radiology cross-modal integrator includes an attention mechanism and a dual cross-modal communication component. The dual cross-modal communication component includes a linear layer, a GELU activation function, a normalization layer, and a regularization function. The Transformer decoder includes a masked multi-head self-attention layer, a multi-head attention layer, a feedforward neural network, residual connections, and layer normalization. The text report features and the visual collaborative features are input into the radiology report generation module. Cross-modal feature alignment is achieved through the interaction of information representations between different modalities, generating accurate and consistent radiology reports, including: Initialize a shared memory unit Where η is the number of storage units; an attention mechanism is applied to embed text features. and visual feature embedding V g Cross-modal feature communication is possible; these features are mapped to a unified shared representation space, producing a text representation M. g and visual representation M i The calculation process is defined as follows: Among them, W v W k W q It is a trainable parameter, q r This represents the query value of the text feature, k. m This represents the key value of the memory unit, q. v This represents the query value for visual features; Since the feature information in the shared representation space is continuously updated during the loop, F(·) is introduced to represent the loop process: X=F(M;M g ,M i )#(11) Where X represents the final feature representation containing both visual and textual feature information; The implementation process of the dual cross-modal communication component is as follows: Where Z and U represent the updated feature matrices after horizontal and vertical interactions, respectively; i ranges from the first row to the dim row of the matrix, and j ranges from the first column to the dim column of the matrix; Norm(·) represents the LayerNorm operation along a specific dimension, Φ(·) represents the GELU activation function; and the weight matrices W1, W2, W3, and W4 are learnable parameters. Introducing the Transformer encoder, visual features generated by the visual co-enhancement module. The input is fed into the encoder; the encoder uses a multi-head self-attention mechanism to input V. g Learn to obtain feature representations that include global contextual information. Among them, f en (·) represents the encoder, and H is the output of the Transformer encoder; Y I =RCMI(V g (R)#(15) Among them, Y I This represents the multi-level cross-modal feature representation obtained through a cross-modal information fusion unit; Since recursive relation generation is an autoregressive process, the decoder is based on previously generated words. The next word is generated sequentially from the encoder's output. Until a complete radiology report is generated; the specific generation process is as follows: in, f represents the final generated report. de (·) indicates the decoder.

7. A radiology report generation system based on visual collaborative enhancement and cross-modal fusion networks, characterized in that, include: The preprocessing module preprocesses the radiological images and text reports in the electronic medical record dataset; The feature representation module uses a visual encoder based on a convolutional neural network to obtain visual features from the preprocessed radiographic images; Furthermore, the preprocessed text report is word vectorized using word2vec to obtain the text report features; The visual collaboration enhancement module identifies fine-grained visual features related to abnormal features in medical images based on the visual features, and obtains visual collaboration features; the visual collaboration enhancement module includes a visual encoder, a global pooling layer, a linear layer, batch normalization, a Tanh activation function, a Dropout layer, a gating unit, and a softmax; The radiology report generation module combines the text report features and the visual collaborative features to achieve cross-modal feature alignment through the interaction of information representations between different modalities, thereby generating accurate and consistent radiology reports. The radiology report generation module includes a multi-head mask attention layer, a normalization layer, a cross-modal information integrator, a feedforward layer, and a linear layer.

Citation Information

Cited By

  • Cross-modal medical image segmentation method and device, electronic equipment and storage medium

    CN122199542A

  • Cross-modal medical image segmentation methods, devices, electronic equipment and storage media

    CN122199542B

  • Radiology image report generation method based on explicit visual evidence

    CN122290851A