Medical report generation method, system and terminal based on cross-view semantic alignment

CN117809794BActive Publication Date: 2026-09-25SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311691412.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2026-09-25
Estimated Expiration
2043-12-08

AI Technical Summary

Technical Problem

[0005]本发明的主要目的在于提供一种基于跨视图语义对齐的医学报告生成方法、系统、终端及计算机可读存储介质,旨在解决现有技术中没有考虑多个视图之间细粒度的语义对齐,无法有效整合信息,导致生成的医学报告内容不准确的问题

Benefits of technology

[0046]本发明中,获取多个不同视角的图像,并将所有所述图像输入预先训练完成的视图自适应网络,所述视图自适应网络对所有所述图像进行压缩融合处理,得到多视图全局表示;将所述多视图全局表示输入预先训练完成的跨视图语义对齐网络,所述跨视图语义对齐网络对所述多视图全局表示进行兴趣提取处理,得到感兴趣文本局部表征;根据所述感兴趣文本局部表征获取医学文本内容,并根据所述医学文本内容生成医学报告。本发明通过跨视图语义对齐网络,实现了多个图像之间的细粒度语义对齐,从而提高了对多个图像的特征提取能力和信息整合能力,使得生成的医学报告内容更加准确。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117809794B_ABST
    Figure CN117809794B_ABST
Patent Text Reader

Abstract

The application discloses a medical report generation method, system and terminal based on cross-view semantic alignment, and the method comprises the following steps: acquiring multiple images of different perspectives, and inputting all the images into a pre-trained view adaptive network; the view adaptive network performs compression and fusion processing on all the images to obtain a multi-view global representation; the multi-view global representation is input into a pre-trained cross-view semantic alignment network, the cross-view semantic alignment network performs interest extraction processing on the multi-view global representation to obtain a local representation of a text of interest; medical text content is acquired according to the local representation of the text of interest, and a medical report is generated according to the medical text content. Through the cross-view semantic alignment network, the application realizes fine-grained semantic alignment between multiple images, thereby improving the feature extraction capability and information integration capability of the multiple images, and making the generated medical report content more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, and in particular to a method, system, terminal, and computer-readable storage medium for generating medical reports based on cross-view semantic alignment. Background Technology

[0002] With the development of technology, large-scale labeled medical image datasets have promoted the development of deep learning-based medical image understanding. In recent years, deep learning-based visual language (VL) representation learning has been pre-trained on a large number of naturally occurring paired image texts. Even with limited labels, it can perform well in various downstream VL tasks in the field of natural language.

[0003] When transferring VL representation learning from the natural language domain to the medical domain, the text and images in the natural language domain mostly present a one-to-one correspondence. Medical image examination requires obtaining a different number of views based on each patient's unique clinical attributes and personalized clinical needs. Furthermore, pathology in medical images typically occupies only a small proportion, so it is crucial to accurately capture the visual information corresponding to the pathology. While existing works have attempted to address these issues, they have only considered the correspondence between a single view and text, neglecting fine-grained semantic alignment between multiple views. This results in ineffective information integration and inaccurate medical reports.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main objective of this invention is to provide a medical report generation method, system, terminal, and computer-readable storage medium based on cross-view semantic alignment, aiming to solve the problem that existing technologies do not consider fine-grained semantic alignment between multiple views, thus failing to effectively integrate information and resulting in inaccurate medical report content.

[0006] To achieve the above objectives, the present invention provides a medical report generation method based on cross-view semantic alignment, the method comprising the following steps:

[0007] Multiple images from different perspectives are acquired, and all of the images are input into a pre-trained view adaptation network. The view adaptation network compresses and fuses all the images to obtain a multi-view global representation.

[0008] The multi-view global representation is input into a pre-trained cross-view semantic alignment network, which performs interest extraction processing on the multi-view global representation to obtain local representations of the text of interest.

[0009] Medical text content is obtained based on the local representation of the text of interest, and a medical report is generated based on the medical text content.

[0010] Optionally, the medical report generation method based on cross-view semantic alignment, wherein acquiring multiple images from different perspectives and inputting all the images into a pre-trained view adaptation network, wherein the view adaptation network performs compression and fusion processing on all the images to obtain a multi-view global representation, specifically includes:

[0011] Receives images from multiple different perspectives input by the user;

[0012] All the images are input into a pre-trained view adaptation network, which performs representation compression processing on each image to obtain a private subspace corresponding to each image.

[0013] All the private subspaces are merged to obtain a common subspace, which is then used as the global representation of the multi-view.

[0014] Optionally, the medical report generation method based on cross-view semantic alignment, wherein the training process of the cross-view semantic alignment network specifically includes:

[0015] Obtain the sample image set and sample report text set;

[0016] The trained multi-view global representation corresponding to the sample image set is obtained based on the view adaptation network;

[0017] The trained multi-view global representation is segmented to obtain multiple trained multi-view local representations;

[0018] The sample report text set is input into a pre-trained natural language model to obtain multiple word text representations, wherein the word text representations correspond one-to-one with the trained multi-view local representations;

[0019] Calculate the global loss and local loss based on all the trained multi-view local representations and all the word text representations;

[0020] The cross-view semantic alignment network is trained based on the global loss and the local loss.

[0021] Optionally, the medical report generation method based on cross-view semantic alignment, wherein calculating the global loss and local loss based on all the trained multi-view local representations and all the word text representations specifically includes:

[0022] Based on each of the trained multi-view local representations and the corresponding word text representations, a first attention matrix corresponding to each of the trained multi-view local representations and a second attention matrix corresponding to each of the word text representations are calculated;

[0023] Each of the trained multi-view local representations is updated based on each of the first attention matrices to obtain multiple updated multi-view local representations;

[0024] Each word text representation is updated based on each of the second attention matrices to obtain multiple updated word text representations;

[0025] The updated multi-view local representations and the updated word text representations are weighted and fused separately to obtain the updated multi-view global representation and the updated global text representation.

[0026] The global loss is calculated based on the updated multi-view global representation and the updated global text representation, and the local loss is calculated based on all the updated multi-view local representations and all the updated word text representations.

[0027] Optionally, the medical report generation method based on cross-view semantic alignment, wherein the step of inputting the multi-view global representation into a pre-trained cross-view semantic alignment network, and the cross-view semantic alignment network performing interest extraction processing on the multi-view global representation to obtain a local representation of the text of interest, specifically includes:

[0028] The multi-view global representation is input into the pre-trained cross-view semantic alignment network, and the multi-view global representation is segmented into multiple view local representations;

[0029] The attention matrix is ​​calculated based on the local representation of each view and the corresponding actual word text representation, wherein the actual word text representation is obtained after the cross-view semantic alignment network is trained;

[0030] All attention matrices are converted into visual graphs, and the sub-regions of interest are obtained based on the visual graphs;

[0031] Based on the sub-region of interest, select the corresponding local representation of the text of interest from all the actual word text representations.

[0032] Optionally, the medical report generation method based on cross-view semantic alignment, wherein obtaining medical text content based on the local representation of the text of interest and generating a medical report based on the medical text content specifically includes:

[0033] Based on the local representation of the text of interest, the corresponding medical text content is selected from a preset text library;

[0034] Semantic analysis was performed on the medical text content to obtain multiple medical keywords;

[0035] Enter all the aforementioned medical keywords into the preset medical report template to generate a medical report.

[0036] Optionally, the medical report generation method based on cross-view semantic alignment further includes:

[0037] The text representation of each actual word is updated according to each attention matrix;

[0038] We perform a weighted fusion of all updated actual word text representations to obtain a global representation of the actual text.

[0039] Based on the actual text global representation, image text retrieval or zero-point text classification is performed to obtain the corresponding retrieval results or classification results.

[0040] Furthermore, to achieve the above objectives, the present invention also provides a medical report generation system based on cross-view semantic alignment, wherein the medical report generation system based on cross-view semantic alignment includes:

[0041] The global representation acquisition module is used to acquire images from multiple different perspectives and input all the images into a pre-trained view adaptation network. The view adaptation network performs compression and fusion processing on all the images to obtain a multi-view global representation.

[0042] The interest representation acquisition module is used to input the multi-view global representation into a pre-trained cross-view semantic alignment network. The cross-view semantic alignment network performs interest extraction processing on the multi-view global representation to obtain the local representation of the text of interest.

[0043] The medical report generation module is used to obtain medical text content based on the local representation of the text of interest, and to generate a medical report based on the medical text content.

[0044] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a medical report generation program based on cross-view semantic alignment stored in the memory and executable on the processor, wherein when the medical report generation program based on cross-view semantic alignment is executed by the processor, it implements the steps of the medical report generation method based on cross-view semantic alignment as described above.

[0045] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a medical report generation program based on cross-view semantic alignment, which, when executed by a processor, implements the steps of the medical report generation method based on cross-view semantic alignment as described above.

[0046] In this invention, multiple images from different perspectives are acquired, and all images are input into a pre-trained view adaptation network. The view adaptation network compresses and fuses all images to obtain a multi-view global representation. This multi-view global representation is then input into a pre-trained cross-view semantic alignment network, which performs interest extraction processing on the multi-view global representation to obtain a local representation of the text of interest. Medical text content is then obtained based on this local representation of the text of interest, and a medical report is generated based on the medical text content. This invention, through the cross-view semantic alignment network, achieves fine-grained semantic alignment between multiple images, thereby improving the feature extraction and information integration capabilities of multiple images, resulting in more accurate medical report content. Attached Figure Description

[0047] Figure 1 This is a flowchart of a preferred embodiment of the medical report generation method based on cross-view semantic alignment of the present invention;

[0048] Figure 2 This is a schematic diagram of the overall process of a preferred embodiment of the medical report generation method based on cross-view semantic alignment of the present invention;

[0049] Figure 3 This is a schematic diagram of the technical route of a preferred embodiment of the medical report generation method based on cross-view semantic alignment of the present invention;

[0050] Figure 4 This is a partial flowchart of a preferred embodiment of the medical report generation method based on cross-view semantic alignment of the present invention;

[0051] Figure 5 This is a schematic diagram illustrating the principle of a preferred embodiment of the medical report generation system based on cross-view semantic alignment of the present invention;

[0052] Figure 6 This is a schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0054] The preferred embodiment of the medical report generation method based on cross-view semantic alignment described in this invention aims to perform fine-grained semantic alignment of multi-view images through text guidance, thereby facilitating downstream VL tasks for various data with limited labels, such as report generation, image text retrieval, and text classification. Therefore, this invention proposes a text-guided cross-view medical semantic alignment framework (TCSA). TCSA includes a report text and image encoding module for processing image text, a cross-modal attention alignment module (i.e., a cross-view semantic alignment network) for performing cross-view semantic alignment, and a local loss and global loss module for optimizing the model.

[0055] like Figure 1 and Figure 2 As shown, the medical report generation method based on cross-view semantic alignment includes the following steps:

[0056] Step S10: Acquire images from multiple different perspectives and input all the images into a pre-trained view adaptation network. The view adaptation network performs compression and fusion processing on all the images to obtain a multi-view global representation.

[0057] Specifically, multiple images from different perspectives (e.g., front-back, left-right, and axis images) are first acquired, and then the acquired multiple images are input into a pre-trained view adaptation network. By simultaneously compressing and fusing multiple images, a multi-view global representation corresponding to all images is obtained, and then the multi-view global representation is processed.

[0058] Furthermore, the acquisition of multiple images from different perspectives, and inputting all of these images into a pre-trained view adaptation network, wherein the view adaptation network performs compression and fusion processing on all the images to obtain a multi-view global representation, specifically includes:

[0059] The system receives multiple images from different viewpoints input by the user; inputs all the images into a pre-trained view adaptation network, which performs representation compression processing on each image to obtain a private subspace corresponding to each image; merges all the private subspaces to obtain a common subspace, and uses the common subspace as a global representation of the multi-view system.

[0060] Specifically, such as Figure 2As shown, the system first receives multiple images from different perspectives input by the user, and then simultaneously inputs these images into a pre-trained View Adaptive Network (VAN). The VAN extracts view representations for each image and then compresses the view representations of each image into their respective low-dimensional latent spaces. The low-dimensional latent space of each image is the private subspace corresponding to that image. Then, the image encoder uses an attention mechanism in the channel dimension of the image to fuse all private subspaces, thus obtaining a common subspace for representing the common image representation of multiple images. Therefore, the common subspace can also be called the multi-view global representation corresponding to multiple images.

[0061] Step S20: Input the multi-view global representation into a pre-trained cross-view semantic alignment network. The cross-view semantic alignment network performs interest extraction processing on the multi-view global representation to obtain local representations of the text of interest.

[0062] Specifically, after obtaining the multi-view global representations corresponding to multiple images, the multi-view global representations are input into the cross-view semantic alignment network (CMAA) pre-trained by the technician. The cross-view semantic alignment network performs fine-grained semantic alignment on the multi-view global representations through an attention mechanism, and then obtains the corresponding local representation of the text of interest.

[0063] Furthermore, the training process of the cross-view semantic alignment network specifically includes:

[0064] Obtain a set of sample images and a set of sample report texts; obtain a trained multi-view global representation corresponding to the set of sample images based on the view adaptation network; segment the trained multi-view global representation to obtain multiple trained multi-view local representations; input the set of sample report texts into a pre-trained natural language model to obtain multiple word text representations, wherein each word text representation corresponds one-to-one with a trained multi-view local representation; calculate global loss and local loss based on all trained multi-view local representations and all word text representations; and complete the training of the cross-view semantic alignment network based on the global loss and the local loss.

[0065] Specifically, when training the cross-view semantic alignment network of the present invention, a preset database is used. The database contains a set of sample images (which also contains images from multiple different perspectives) and a set of sample report texts for training. Therefore, at the beginning of training, the set of sample images and the set of sample texts in the database are first obtained. The pre-trained view adaptation network is used to compress and fuse all training images in the set of sample images to obtain the training multi-view global representation (also known as the training common subspace, the specific steps of which are the same as above and will not be repeated here) corresponding to all training images in the set of sample images. After obtaining the training multi-view global representation, the training multi-view global representation is segmented and uniformly divided into multiple training multi-view local representations.

[0066] Next, the sample report text set needs to be processed. First, the sample report text set is input into a pre-trained natural language model (such as the BERT model). The natural language model will extract word text representations from the sample report text set (equivalent to...). Figure 2 The text representations are local representations of the training multi-view semantic alignment network (VL). Each word text representation corresponds one-to-one with the training multi-view local representations. After obtaining multiple word text representations, an attention matrix is ​​calculated for each training multi-view local representation and its corresponding word text representation based on the attention mechanism. Then, each training multi-view local representation and word text representation is updated based on the attention matrix. The updated training multi-view local representations and word text representations are used to calculate the local and global losses. The local and global losses are minimized using gradient descent. The network parameters of the VL network are optimized based on the minimized local and global losses, bringing similar samples closer together and pushing dissimilar samples further apart. This allows similar samples to cluster in the feature space, which is beneficial for subsequent VL downstream tasks.

[0067] It should be noted that, in a preferred embodiment of the present invention, three databases are used to train the cross-view semantic alignment network, namely the chest X-ray database MIMIC-CXR2.0.0, the CheXpert dataset, and the CheXpert 5×200 dataset, as shown below. Figure 3 As shown, the CheXpert dataset and the CheXpert 5×200 dataset are used as validation samples for downstream tasks to fine-tune the cross-view semantic alignment network and verify its performance (equivalent to...). Figure 3 (Model fine-tuning in the process).

[0068] During the training phase of the cross-view semantic network, the MIMIC-CXR2.0.0 dataset was used to train the network, which contains 377,110 images corresponding to 227,835 X-ray cases. In the model fine-tuning phase, the CheXpert and CheXpert 5×200 datasets were used for downstream tasks, and the cross-view semantic network was fine-tuned during these tasks to validate its performance. The CheXpert dataset contains 224,316 chest X-rays from 65,240 patients. Based on the distribution of studies with different numbers of views in the CheXpert dataset, we extracted 3,000 samples (5,000 chest X-rays) from the CheXpert dataset, including radiological studies with different numbers of views. CheXpert 5×200 is a subset of CheXpert containing 1,000 chest X-rays from 1,000 studies.

[0069] Further, the step of calculating the global loss and local loss based on all the trained multi-view local representations and all the word text representations specifically includes:

[0070] Based on each trained multi-view local representation and its corresponding word text representation, a first attention matrix and a second attention matrix are calculated for each trained multi-view local representation. Each trained multi-view local representation is updated based on each first attention matrix to obtain multiple updated multi-view local representations. Each word text representation is updated based on each second attention matrix to obtain multiple updated word text representations. All updated multi-view local representations and all updated word text representations are weighted and fused to obtain updated multi-view global representations and updated global text representations. A global loss is calculated based on the updated multi-view global representations and the updated global text representations, and a local loss is calculated based on all updated multi-view local representations and all updated word text representations.

[0071] Specifically, during the training process of the cross-semantic alignment network, corresponding local and global losses are calculated to optimize the network parameters. In a preferred embodiment of the invention, each trained multi-view local representation and its corresponding word text representation are calculated according to the attention mechanism to obtain a first attention matrix corresponding to each trained multi-view local representation and a second attention matrix corresponding to each word text representation. Subsequently, the first and second attention matrices are used to update each trained multi-view local representation and each word text representation (i.e., modality fusion). After the update is completed, the updated multi-view global representation corresponding to each trained multi-view local representation and the updated word text representation corresponding to each word text representation are obtained. Then, the corresponding local loss can be calculated based on all updated multi-view global representations and all updated word text representations. Finally, all updated multi-view local representations are weighted and fused to obtain the updated multi-view global representation (equivalent to...). Figure 2 The visual global representation in the image is used to weight and fuse all the updated local text representations to obtain the updated global text representation (equivalent to...). Figure 2 The text global representation is used to update the multi-view global representation and the updated text local representation. Then, the corresponding global loss is calculated based on the updated multi-view global representation and the updated text local representation. By using text guidance to align fine-grained semantics between multiple images from different perspectives, the extraction of consistent and complementary features between images is promoted. This enables better learning of multi-view features and improves the accuracy and robustness of the cross-semantic alignment network.

[0072] Further, the step of inputting the multi-view global representation into a pre-trained cross-view semantic alignment network, wherein the cross-view semantic alignment network performs interest extraction processing on the multi-view global representation to obtain local representations of the text of interest, specifically includes:

[0073] The multi-view global representation is input into the pre-trained cross-view semantic alignment network, and the multi-view global representation is segmented into multiple view local representations. An attention matrix is ​​calculated for each view local representation and its corresponding actual word text representation, wherein the actual word text representation is obtained after training the cross-view semantic alignment network. All attention matrices are converted into visualizations, and regions of interest are obtained from the visualizations. Corresponding text local representations of interest are selected from all the actual word text representations based on the regions of interest.

[0074] Specifically, after the cross-view semantic alignment network is trained, it can be used to perform interest reminder processing on the multi-view global representations corresponding to multiple different perspectives of the user-inputted image that require cross-view semantic alignment. First, the multi-view global representation is uniformly divided into multiple view local representations. Then, the actual word text features corresponding to each view local representation are obtained in the cross-view semantic alignment network. Next, the corresponding attention matrix is ​​calculated based on each view local representation and the corresponding actual word text features. That is, the third attention matrix corresponding to each view local representation and the fourth attention matrix corresponding to each actual word text feature are calculated. Finally, all the third attention matrices and fourth attention matrices are converted into a visualization. The region with the highest brightness in the visualization is obtained. This region is the sub-region of interest. The actual word text representation corresponding to the sub-region of interest is then obtained and used as the text representation of interest.

[0075] like Figure 4 As shown, after obtaining the text features of interest corresponding to the multi-view global representation, for example, if the obtained text features of interest are "Pneumothorax", then "Pneumothorax" can be embedded into the attention map (that is, the visualization above) to weight the multi-view global representation, so as to obtain the "Pneumothorax" weighted visual representation. This makes the cross-view semantic alignment network pay more attention to the regions related to pneumothorax, thereby capturing more comprehensive feature information related to pneumothorax.

[0076] It should be noted that the actual word text features are actually trained on the cross-semantic alignment network by training the network on the sample report text set. After training, the relevant text parameters are left behind. These text parameters can become the actual word text features, so that when applying the cross-semantic alignment network later, there is no need to input the text report again. Only multi-view images need to be input for semantic alignment.

[0077] Step S30: Obtain medical text content based on the local representation of the text of interest, and generate a medical report based on the medical text content.

[0078] Specifically, after obtaining the local representation of the text of interest, the corresponding medical text content is obtained based on the local representation of the text of interest. The medical text content can be sentences or corresponding articles, and then the corresponding medical report is generated based on these medical text contents.

[0079] Furthermore, the step of obtaining medical text content based on the local representation of the text of interest, and generating a medical report based on the medical text content, specifically includes:

[0080] Based on the local representation of the text of interest, corresponding medical text content is selected from a preset text library; semantic analysis is performed on the medical text content to obtain multiple medical keywords; all the medical keywords are filled into a preset medical report template to generate a medical report.

[0081] Specifically, based on the local representation of the text of interest, the corresponding medical text content is selected from a pre-set text library. The text library can be a database consisting of multiple medical reports or medical professional books. Then, semantic analysis is performed on the medical text content, and the medical keywords corresponding to the local representation of the text of interest are segmented to obtain multiple medical keywords. Finally, the medical keywords are filled into the medical report template set in the background by the technicians to generate the corresponding medical report.

[0082] Furthermore, the medical report generation method based on cross-view semantic alignment also includes:

[0083] The text representation of each actual word is updated according to each attention matrix; all updated text representations of actual words are weighted and fused to obtain a global representation of the actual text; image text retrieval or zero-point text classification is performed based on the global representation of the actual text to obtain the corresponding retrieval results or classification results.

[0084] Specifically, in another preferred embodiment of the present invention, an attention matrix can be used to update the actual word text representation, and the updated actual word text representation can be weighted and fused to obtain the actual text global representation. The actual text global representation can be used for image text retrieval tasks or zero-shot text classification tasks, and the corresponding retrieval results or classification results can be obtained after completion. Finally, the cross-semantic alignment network can be fine-tuned according to the retrieval results or classification results to improve the network performance.

[0085] In the corresponding preferred embodiment, the retrieval results of the image-text retrieval task using TCSA on the CheXpert 5×200 dataset are compared with existing processing frameworks as shown in Table 1:

[0086] Table 1

[0087]

[0088] Among them, ConVIRT and GLoRIA are existing processing frameworks. The evaluation metric is the Top-K average progress of the image text retrieval task. Top-K refers to the accuracy of retrieving the top K texts that contain texts registered with the current image. K refers to the number of texts. In Table 1, K takes the values ​​of 5, 10, and 100.

[0089] As shown in Table 1, in the image-text retrieval task, the query image is used as input, and the closest matching text is located based on the similarity of their representations. The Top-K mean precision metric is used to evaluate whether the selected reports belong to the same category as the query image, and the accuracy of the top K retrieved reports is calculated. Table 1 shows the results of the image-text retrieval task on CheXpert 5×200, indicating that TCSA (the text-guided cross-view medical semantic alignment framework proposed in this invention) achieves the best results.

[0090] Furthermore, since medical reports are typically generated based on multiple views, while baseline models can only process a single view, an information gap exists between the report and the image, resulting in poor performance in medical image text retrieval tasks. TCSA utilizes text-guided comprehensive analysis of any number of views to eliminate the information gap between the report and the medical image; therefore, TCSA achieves the best results in image text retrieval tasks.

[0091] In a preferred embodiment, the classification results of TCSA on the CheXPert and CheXpert 5×200 datasets for zero-point text classification tasks are compared with existing processing frameworks, as shown in Table 2:

[0092] Table 2

[0093]

[0094] Zero-point text classification task refers to classifying a category without seeing any data for that category. The evaluation metric used in Table 2 is AUROC.

[0095] As shown in Table 2, TCSA has significant improvements over GLoRIA and ConVIRT, thanks to the integration of multi-view information.

[0096] It is worth noting that this invention was pre-trained on MIMIC-CXR and zero-shot classification was performed on the external datasets CheXpert and CheXpert5×200 to verify the generalization performance of TCSA.

[0097] Furthermore, such as Figure 5 As shown, based on the above-described medical report generation method based on cross-view semantic alignment, the present invention also provides a medical report generation system based on cross-view semantic alignment, wherein the medical report generation system based on cross-view semantic alignment includes:

[0098] The global representation acquisition module 51 is used to acquire images from multiple different perspectives and input all the images into a pre-trained view adaptation network. The view adaptation network performs compression and fusion processing on all the images to obtain a multi-view global representation.

[0099] The interest representation acquisition module 52 is used to input the multi-view global representation into a pre-trained cross-view semantic alignment network, and the cross-view semantic alignment network performs interest extraction processing on the multi-view global representation to obtain the local representation of the text of interest.

[0100] The medical report generation module 53 is used to obtain medical text content based on the local representation of the text of interest, and generate a medical report based on the medical text content.

[0101] Furthermore, such as Figure 6 As shown, based on the above-mentioned medical report generation method and system based on cross-view semantic alignment, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 6 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0102] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a medical report generation program 40 based on cross-view semantic alignment, which can be executed by the processor 10 to implement the medical report generation method based on cross-view semantic alignment in this application.

[0103] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the medical report generation method based on cross-view semantic alignment.

[0104] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components 10-30 of the terminal communicate with each other via a system bus.

[0105] In one embodiment, when processor 10 executes medical report generation program 40 based on cross-view semantic alignment in memory 20, the following steps are performed:

[0106] Multiple images from different perspectives are acquired, and all of the images are input into a pre-trained view adaptation network. The view adaptation network compresses and fuses all the images to obtain a multi-view global representation.

[0107] The multi-view global representation is input into a pre-trained cross-view semantic alignment network, which performs interest extraction processing on the multi-view global representation to obtain local representations of the text of interest.

[0108] Medical text content is obtained based on the local representation of the text of interest, and a medical report is generated based on the medical text content.

[0109] The process of acquiring multiple images from different perspectives and inputting all of these images into a pre-trained view adaptation network, wherein the view adaptation network performs compression and fusion processing on all the images to obtain a multi-view global representation, specifically includes:

[0110] Receives images from multiple different perspectives input by the user;

[0111] All the images are input into a pre-trained view adaptation network, which performs representation compression processing on each image to obtain a private subspace corresponding to each image.

[0112] All the private subspaces are merged to obtain a common subspace, which is then used as the global representation of the multi-view.

[0113] The training process of the cross-view semantic alignment network specifically includes:

[0114] Obtain the sample image set and sample report text set;

[0115] The trained multi-view global representation corresponding to the sample image set is obtained based on the view adaptation network.

[0116] The trained multi-view global representation is segmented to obtain multiple trained multi-view local representations;

[0117] The sample report text set is input into a pre-trained natural language model to obtain multiple word text representations, wherein the word text representations correspond one-to-one with the trained multi-view local representations;

[0118] Calculate the global loss and local loss based on all the trained multi-view local representations and all the word text representations;

[0119] The cross-view semantic alignment network is trained based on the global loss and the local loss.

[0120] Specifically, calculating the global and local losses based on all the trained multi-view local representations and all the word text representations includes:

[0121] Based on each of the trained multi-view local representations and the corresponding word text representations, a first attention matrix corresponding to each of the trained multi-view local representations and a second attention matrix corresponding to each of the word text representations are calculated;

[0122] Each of the trained multi-view local representations is updated based on each of the first attention matrices to obtain multiple updated multi-view local representations;

[0123] Each word text representation is updated based on each of the second attention matrices to obtain multiple updated word text representations;

[0124] The updated multi-view local representations and the updated word text representations are weighted and fused separately to obtain the updated multi-view global representation and the updated global text representation.

[0125] The global loss is calculated based on the updated multi-view global representation and the updated global text representation, and the local loss is calculated based on all the updated multi-view local representations and all the updated word text representations.

[0126] Specifically, the step of inputting the multi-view global representation into a pre-trained cross-view semantic alignment network, wherein the cross-view semantic alignment network performs interest extraction processing on the multi-view global representation to obtain local representations of the text of interest, includes:

[0127] The multi-view global representation is input into the pre-trained cross-view semantic alignment network, and the multi-view global representation is segmented into multiple view local representations;

[0128] The attention matrix is ​​calculated based on the local representation of each view and the corresponding actual word text representation, wherein the actual word text representation is obtained after the cross-view semantic alignment network is trained;

[0129] All attention matrices are converted into visual graphs, and the sub-regions of interest are obtained based on the visual graphs;

[0130] Based on the sub-region of interest, select the corresponding local representation of the text of interest from all the actual word text representations.

[0131] The step of obtaining medical text content based on the local representation of the text of interest and generating a medical report based on the medical text content specifically includes:

[0132] Based on the local representation of the text of interest, the corresponding medical text content is selected from a preset text library;

[0133] Semantic analysis was performed on the medical text content to obtain multiple medical keywords;

[0134] Enter all the aforementioned medical keywords into the preset medical report template to generate a medical report.

[0135] The medical report generation method based on cross-view semantic alignment further includes:

[0136] The text representation of each actual word is updated according to each attention matrix;

[0137] We perform a weighted fusion of all updated actual word text representations to obtain a global representation of the actual text.

[0138] Based on the actual text global representation, image text retrieval or zero-point text classification is performed to obtain the corresponding retrieval results or classification results.

[0139] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a medical report generation program based on cross-view semantic alignment, the medical report generation program based on cross-view semantic alignment implementing the steps of the medical report generation method based on cross-view semantic alignment as described above when executed by a processor.

[0140] In summary, this invention provides a medical report generation method, system, and terminal based on cross-view semantic alignment. The method includes: acquiring multiple images from different perspectives and inputting all the images into a pre-trained view adaptation network; the view adaptation network compresses and fuses all the images to obtain a multi-view global representation; inputting the multi-view global representation into a pre-trained cross-view semantic alignment network; the cross-view semantic alignment network extracts interest from the multi-view global representation to obtain a local representation of text of interest; obtaining medical text content based on the local representation of text of interest; and generating a medical report based on the medical text content. This invention achieves fine-grained semantic alignment between multiple images through a cross-view semantic alignment network, thereby improving the feature extraction and information integration capabilities of multiple images, resulting in more accurate generated medical report content.

[0141] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0142] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0143] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for generating medical reports based on cross-view semantic alignment, characterized in that, The medical report generation method based on cross-view semantic alignment includes: Multiple images from different perspectives are acquired, and all of the images are input into a pre-trained view adaptation network. The view adaptation network compresses and fuses all the images to obtain a multi-view global representation. The process involves acquiring images from multiple different viewpoints and inputting all of these images into a pre-trained view adaptation network. The view adaptation network then performs compression and fusion processing on all the images to obtain a multi-view global representation. Specifically, this includes: Receives images from multiple different perspectives input by the user; All the images are input into a pre-trained view adaptation network, which performs representation compression processing on each image to obtain a private subspace corresponding to each image. All the private subspaces are merged to obtain a common subspace, which is then used as the global representation of the multi-view. The multi-view global representation is input into a pre-trained cross-view semantic alignment network, which performs interest extraction processing on the multi-view global representation to obtain local representations of the text of interest. The training process of the cross-view semantic alignment network specifically includes: Obtain the sample image set and sample report text set; The trained multi-view global representation corresponding to the sample image set is obtained based on the view adaptation network; The trained multi-view global representation is segmented to obtain multiple trained multi-view local representations; The sample report text set is input into a pre-trained natural language model to obtain multiple word text representations, wherein the word text representations correspond one-to-one with the trained multi-view local representations; Calculate the global loss and local loss based on all the trained multi-view local representations and all the word text representations; The training of the cross-view semantic alignment network is completed based on the global loss and the local loss. The calculation of global and local losses based on all the trained multi-view local representations and all the word text representations specifically includes: Based on each of the trained multi-view local representations and the corresponding word text representations, a first attention matrix corresponding to each of the trained multi-view local representations and a second attention matrix corresponding to each of the word text representations are calculated; Each of the trained multi-view local representations is updated based on each of the first attention matrices to obtain multiple updated multi-view local representations; Each word text representation is updated based on each of the second attention matrices to obtain multiple updated word text representations; The updated multi-view local representations and the updated word text representations are weighted and fused separately to obtain the updated multi-view global representation and the updated global text representation. The global loss is calculated based on the updated multi-view global representation and the updated global text representation, and the local loss is calculated based on all the updated multi-view local representations and all the updated word text representations; Medical text content is obtained based on the local representation of the text of interest, and a medical report is generated based on the medical text content.

2. The medical report generation method based on cross-view semantic alignment according to claim 1, characterized in that, The process of inputting the multi-view global representation into a pre-trained cross-view semantic alignment network, wherein the cross-view semantic alignment network performs interest extraction processing on the multi-view global representation to obtain local representations of text of interest, specifically includes: The multi-view global representation is input into the pre-trained cross-view semantic alignment network, and the multi-view global representation is segmented into multiple view local representations; The attention matrix is ​​calculated based on the local representation of each view and the corresponding actual word text representation, wherein the actual word text representation is obtained after the cross-view semantic alignment network is trained; All attention matrices are converted into visual graphs, and the sub-regions of interest are obtained based on the visual graphs; Based on the sub-region of interest, select the corresponding local representation of the text of interest from all the actual word text representations.

3. The medical report generation method based on cross-view semantic alignment according to claim 1, characterized in that, The step of obtaining medical text content based on the local representation of the text of interest, and generating a medical report based on the medical text content, specifically includes: Based on the local representation of the text of interest, the corresponding medical text content is selected from a preset text library; Semantic analysis was performed on the medical text content to obtain multiple medical keywords; Enter all the aforementioned medical keywords into the preset medical report template to generate a medical report.

4. The medical report generation method based on cross-view semantic alignment according to claim 2, characterized in that, The medical report generation method based on cross-view semantic alignment also includes: The text representation of each actual word is updated according to each attention matrix; We perform a weighted fusion of all updated actual word text representations to obtain a global representation of the actual text. Based on the actual text global representation, image text retrieval or zero-point text classification is performed to obtain the corresponding retrieval results or classification results.

5. A medical report generation system based on cross-view semantic alignment, characterized in that, The medical report generation system based on cross-view semantic alignment is used to implement the medical report generation method based on cross-view semantic alignment as described in any one of claims 1-4, wherein the medical report generation system based on cross-view semantic alignment includes: The global representation acquisition module is used to acquire images from multiple different perspectives and input all the images into a pre-trained view adaptation network. The view adaptation network performs compression and fusion processing on all the images to obtain a multi-view global representation. The interest representation acquisition module is used to input the multi-view global representation into a pre-trained cross-view semantic alignment network. The cross-view semantic alignment network performs interest extraction processing on the multi-view global representation to obtain the local representation of the text of interest. The medical report generation module is used to obtain medical text content based on the local representation of the text of interest, and to generate a medical report based on the medical text content.

6. A terminal, characterized in that, The terminal includes: a memory, a processor, and a medical report generation program based on cross-view semantic alignment stored in the memory and executable on the processor. When the medical report generation program based on cross-view semantic alignment is executed by the processor, it implements the steps of the medical report generation method based on cross-view semantic alignment as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a medical report generation program based on cross-view semantic alignment, which, when executed by a processor, implements the steps of the medical report generation method based on cross-view semantic alignment as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Image classification method and device, computer readable storage medium and computer equipment

    CN110321920A

  • Cross-modal image-text retrieval method based on multi-level semantic alignment

    CN116821391A