Chest X-ray film multi-mode pre-training method and system based on graph perception learning

By constructing a multi-round question-answering dictionary and graph-aware pre-training method, high-quality local and global descriptive text is generated, which solves the problems of scarce chest X-ray data and insufficient semantic comparison, and achieves higher analysis accuracy and generalization ability.

CN120911533AInactive Publication Date: 2025-11-07NANCHANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511438882.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to acquire image-text data from chest X-rays, resulting in data scarcity. Furthermore, the semantic comparison between images and text is insufficient, affecting the accuracy and generalization ability of models in complex analysis tasks.

Method used

By constructing a multi-round question-answering dictionary to generate local and global descriptive text, and combining graph-aware pre-training methods, a correlation graph structure of local and global features is constructed to deeply explore the cross-modal correlation between different parts of chest X-rays and diseases, thereby reducing the modal differences between images and text.

Benefits of technology

It improves the accuracy and generalization ability of chest X-ray analysis, effectively solves the problem of insufficient data, and accurately captures subtle visual differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911533A_ABST
    Figure CN120911533A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, and provides a chest X-ray film multi-modal pre-training method and system based on graph perception learning, and the method comprises the steps: carrying out the data generation through constructing a multi-round question and answer dictionary, generating a local descriptive text of each lesion part in a chest X-ray film from the three levels of disease classification, classification certainty and corresponding lesion parts, and carrying out the recognition of the local descriptive text. According to the method, a global description text is automatically generated, the problem of insufficient data is effectively avoided, the quality and consistency of text description are improved, and a correlation graph structure between local and global features is constructed through graph perception pre-training based on a global-to-local graph perception learning method, so that the accuracy of text description is improved. The cross-modal relevance between each part of the chest radiograph and a disease is deeply mined, tiny visual differences which are difficult to recognize are more accurately captured, modal differences between images and texts are reduced, and the accuracy and generalization ability of chest X-ray radiograph analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, in particular to a chest X-ray film multi-modal pre-training method and system based on graph perception learning. BACKGROUND

[0002] With the rapid development of deep learning, there are extensive applications in various fields, especially in the field of image analysis. Because of the complex image structure and the difficulty of analysis, the chest X-ray film often requires sufficient analysis experience and more time cost for manual analysis. Therefore, using a model to assist in analyzing the chest X-ray film is indispensable.

[0003] In the prior art, the model is used to assist in analyzing the chest X-ray film, which often faces two major problems. First, the image-text pair data acquisition is difficult, which leads to the challenge of data scarcity in model training. Because the manually annotated image-text pair data is limited by the differences in professional terms and expression methods, the quality and consistency of the text description are difficult to guarantee, which makes it time-consuming and expensive to obtain high-quality image-text pair data. In addition, although the existing data enhancement method alleviates the problem of data scarcity to some extent, the quality of the generated image-text pair is still limited, which is difficult to meet the demand of deep learning model for large-scale high-quality data. Second, the existing method has deficiencies in the semantic comparison between image and text, especially in the small visual feature differences of specific disease categories. Such differences are often difficult to identify and analyze without in-depth clinical knowledge, which leads to poor performance of the model in complex analysis tasks.

[0004] Therefore, how to avoid the influence of data scarcity and image-text comparison to improve the precision and generalization ability of the model in analyzing the chest X-ray film has become a problem to be solved. SUMMARY

[0005] Therefore, the chest X-ray film multi-modal pre-training method and system based on graph perception learning is proposed, which generates data by constructing a multi-round question and answer dictionary. From the aspects of disease classification, classification certainty and corresponding lesion site, local descriptive text of each lesion site in the chest X-ray film is generated, and global descriptive text is automatically generated. The problem of data scarcity is effectively avoided, and the quality and consistency of the text description are improved. Through graph perception pre-training, based on the global-to-local graph perception learning method, the correlation graph structure between local and global features is constructed, the cross-modal relevance between each part of the chest film and the disease is deeply mined, the small visual differences that are difficult to identify are more accurately captured, and the modal difference between image and text is reduced. The present application improves the precision and generalization ability of analyzing the chest X-ray film.

[0006] The application provides a chest X-ray film multi-modal pre-training method based on graph perception learning, which comprises the following steps: An X-ray image-label data set and an X-ray image-text data set are obtained and input into a data generation module, and the data generation module is based on a large language model; Label data in the X-ray image-label data set and the X-ray image-text data set are extracted, a multi-round question and answer dictionary is constructed according to the label data, global description text is obtained, the label data comprises normal labels and lesion labels, the multi-round question and answer dictionary is based on normal part description text and lesion part description text, and the global description text is generated based on local description text; Image data in the X-ray image-label data set and the X-ray image-text data set are subjected to image enhancement processing, and then the image data subjected to image enhancement processing and the global description text are input into a graph perception pre-training module, and the graph perception pre-training module is based on a global-local architecture; Text feature extraction is performed on the global description text according to a text encoder, so that global text features and local text features are obtained, and image feature extraction is performed on the image data according to an image encoder, so that global visual features are obtained; Global contrast learning is performed according to the global text features and the global visual features, and graph perception learning is performed according to the global text features and the local text features, so that global contrast learning enhanced features and graph perception enhanced features are obtained; Cross-modal similarity contrast loss optimization is performed according to the global contrast learning enhanced features and the graph perception enhanced features.

[0007] In summary, according to the above-mentioned chest X-ray multi-modal pre-training method based on graph perception learning, by constructing a multi-turn question and answer dictionary for data generation, from the three aspects of disease classification, classification certainty and corresponding lesion site, local descriptive text of each lesion site in the chest X-ray is generated, and global descriptive text is automatically generated, effectively avoiding the problem of insufficient data, and improving the quality and consistency of the text description, and through graph perception pre-training, based on the graph perception learning method from global to local, the correlation graph structure between local and global features is constructed, the cross-modal correlation between each part of the chest X-ray and the disease is deeply mined, the difficult-to-identify micro visual differences are more accurately captured, and the modal difference between the image and the text is reduced, the precision and generalization ability of analyzing the chest X-ray are improved. Specifically, the chest X-ray image-label data set and the chest X-ray image-text data set are obtained and input into the data generation module, the data generation module based on a large language model extracts label data in the chest X-ray image-label data set and the chest X-ray image-text data set, constructs a multi-turn question and answer dictionary according to the label data to obtain global descriptive text, the label data includes normal labels and lesion labels, the multi-turn question and answer dictionary is based on normal site description text and lesion site description text, and the global description text is generated based on local description text. By constructing a multi-turn question and answer dictionary for data generation, from the three aspects of disease classification, classification certainty and corresponding lesion site, local descriptive text of each lesion site in the chest X-ray is generated, and global descriptive text is automatically generated, effectively avoiding the problem of insufficient data, and improving the quality and consistency of the text description, and through graph perception pre-training, based on the graph perception learning method from global to local, the correlation graph structure between local and global features is constructed, the cross-modal correlation between each part of the chest X-ray and the disease is deeply mined, the difficult-to-identify micro visual differences are more accurately captured, and the modal difference between the image and the text is reduced, the precision and generalization ability of analyzing the chest X-ray are improved. Specifically, the chest X-ray image-label data set and the chest X-ray image-text data set are obtained and input into the data generation module, the data generation module based on a large language model extracts label data in the chest X-ray image-label data set and the chest X-ray image-text data set, constructs a multi-turn question and answer dictionary according to the label data to obtain global descriptive text, the label data includes normal labels and lesion labels, the multi-turn question and answer dictionary is based on normal site description text and lesion site description text, and the global description text is generated based on local description text. By constructing a multi-turn question and answer dictionary for data generation, from the three aspects of disease classification, classification certainty and corresponding lesion site, local descriptive text of each lesion site in the chest X-ray is generated, and global descriptive text is automatically generated, effectively avoiding the problem of insufficient data, and improving the quality and consistency of the text description, and through graph perception pre-training, based on the graph perception learning method from global to local, the correlation graph structure between local and global features is constructed, the cross-modal correlation between each part of the chest X-ray and the disease is deeply mined, the difficult-to-identify micro visual differences are more accurately captured, and the modal difference between the image and the text is reduced, the precision and generalization ability of analyzing the chest X-ray are improved.

[0008] Further, the step of extracting the label data in the chest X-ray image-label data set and the chest X-ray image-text data set, and constructing a multi-round question and answer dictionary according to the label data to obtain the global description text specifically comprises: extracting the label data in the chest X-ray image-label data set and the chest X-ray image-text data set, the label data comprising normal labels and lesion labels; generating normal site description texts according to the standard anatomical site list and the normal labels; generating lesion site description texts according to the disease category-lesion anatomical site mapping relationship table and the lesion labels; constructing a multi-round question and answer dictionary according to the normal site description texts and the lesion site description texts, the multi-round question and answer dictionary being constructed based on a large language model; performing description text expansion according to the multi-round question and answer dictionary to obtain a plurality of local description texts, and performing report text supplement according to the local description texts to obtain a global description text.

[0009] Further, the step of constructing a multi-round question and answer dictionary according to the label data further comprises: In the first round of question and answer, the multi-round question and answer dictionary generates a standard anatomical site list according to a standard chest X-ray report, the standard anatomical sites in the standard anatomical site list corresponding one-to-one to the normal labels; In the second round of question and answer, the multi-round question and answer dictionary generates a disease category-lesion anatomical site mapping relationship table according to the disease category and the lesion labels; In the third round of question and answer, the multi-round question and answer dictionary generates normal site description texts according to the standard anatomical sites by a large language model, the normal site description texts corresponding one-to-one to the standard anatomical sites; In the fourth round of question and answer, the multi-round question and answer dictionary generates lesion site description texts according to the lesion anatomical sites by a large language model, the lesion site description texts corresponding one-to-one to the disease categories; In the fifth round of question and answer, the multi-round question and answer dictionary is iteratively optimized multiple times to generate an expanded description text list for each disease category according to a large language model and form a final multi-round question and answer dictionary.

[0010] Further, the step of performing graph perception learning according to the global text features and the local text features specifically comprises: constructing a graph structure according to the global text features, the local text features, and the global visual features, the graph structure comprising nodes and edges, the nodes comprising global text feature nodes, local text feature nodes, and global visual feature nodes, the local text feature nodes being bridge nodes, the bridge nodes being used to connect the global text feature nodes and the global visual feature nodes; performing local perception enhancement processing according to the graph convolutional neural network and the graph structure.

[0011] Further, the step of constructing the graph structure according to the global text feature, the local text feature and the global visual feature specifically comprises: constructing the global text feature, the local text feature and the global visual feature as nodes of the graph structure, and constructing edges of the graph structure according to types of the nodes, wherein the edges of the graph structure include local text feature bidirectional edges, global text-local text feature unidirectional edges, global visual-local text feature unidirectional edges and global text-global visual feature bidirectional edges, the local text feature bidirectional edges are used for the local text feature nodes to obtain enhanced information representation, the global text-local text feature unidirectional edges are used for global-local text information supplement, the global visual-local text feature unidirectional edges are used for cross-modal advanced semantic information fusion, and the global text-global visual feature bidirectional edges are used for cross-modal global information fusion.

[0012] Further, the step of performing local perception enhancement processing according to the graph convolutional neural network and the graph structure specifically comprises: performing local perception enhancement processing according to the graph convolutional neural network and the graph structure, wherein the graph convolutional neural network comprises two graph convolutional layers and one ReLU activation layer, and a graph convolution algorithm of the graph convolutional layer is specifically as follows: , , wherein, denotes a node feature matrix, denotes an identity matrix, denotes an activation function, denotes a degree matrix of the node, denotes a trainable parameter matrix of the convolutional layer, denotes an adjacency matrix; a specific algorithm of the graph convolutional neural network is as follows: , wherein, denotes the graph convolutional neural network, denotes a graph convolution operation, denotes a local text feature node, a global text feature node and a global visual feature node respectively.

[0013] Further, the step of performing cross-modal similarity contrast loss optimization according to the global contrast learning enhanced feature and the graph perception enhanced feature specifically comprises: a specific algorithm of the cross-modal similarity contrast loss optimization is as follows: , wherein, represents a cross-modal loss, represents a cosine similarity calculation, respectively represent a local text feature node, a global text feature node and a global visual feature node.

[0014] The application provides a chest X-ray multi-modal pre-training system based on graph perception learning, comprising: A data acquisition module is configured to acquire a chest X-ray image-label dataset and a chest X-ray image-text dataset and input the data into a data generation module, wherein the data generation module is based on a large language model; The data generation module is configured to extract label data from the chest X-ray image-label dataset and the chest X-ray image-text dataset, construct a multi-turn question and answer dictionary according to the label data, and obtain global description text, wherein the label data comprises normal labels and lesion labels, the multi-turn question and answer dictionary is based on normal part description text and lesion part description text, and the global description text is generated based on local description text; An image enhancement module is configured to perform image enhancement processing on image data in the chest X-ray image-label dataset and the chest X-ray image-text dataset, and then input the image data after the image enhancement processing and the global description text into a graph perception pre-training module, wherein the graph perception pre-training module is based on a global-local architecture; The graph perception pre-training module is configured to perform text feature extraction on the global description text according to a text encoder to obtain global text features and local text features, and perform image feature extraction on the image data according to an image encoder to obtain global visual features; Global contrast learning is performed according to the global text features and the global visual features, and graph perception learning is performed according to the global text features and the local text features to obtain global contrast learning enhanced features and graph perception enhanced features, respectively; Cross-modal similarity contrast loss optimization is performed according to the global contrast learning enhanced features and the graph perception enhanced features.

[0015] The application further provides a storage medium storing one or more programs, wherein the programs are executed by a processor to implement the chest X-ray multi-modal pre-training method based on graph perception learning as described above.

[0016] The application further provides a computer device comprising a memory and a processor, wherein: The memory is configured to store a computer program; The processor is used to execute the computer program stored in the memory, and the chest X-ray multi-modal pre-training method based on graph perception learning is realized. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 The flow chart of the chest X-ray multi-modal pre-training method based on graph perception learning for the first embodiment of the present application is shown in Figure 2 The structure schematic diagram of the chest X-ray multi-modal pre-training system based on graph perception learning for the second embodiment of the present application is shown in Figure 3 The model framework diagram of the first embodiment of the present application is shown in Figure 4 The data generation framework diagram of the first embodiment of the present application is shown in Figure 5 The graph construction framework diagram of the first embodiment of the present application is shown in

[0018] The following specific embodiments will further illustrate the present application in combination with the above-mentioned drawings. DETAILED DESCRIPTION

[0019] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. Several embodiments of the present application are shown in the drawings. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0020] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can be a middle element. When an element is referred to as being "connected" to another element, it can be directly connected to the other element or there can be a middle element. The terms "vertical", "horizontal", "left", "right", and similar expressions used herein are for illustrative purposes only.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terminology used in the description of the present application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0022] Referring to Figure 1 , the flow chart of the chest X-ray multi-modal pre-training method based on graph perception learning for the first embodiment of the present application is shown, and the chest X-ray multi-modal pre-training method based on graph perception learning includes steps S01 to S06, wherein: Step S01: Obtain the chest X-ray image-label data set and the chest X-ray image-text data set and input the data generation module; It should be noted that in the embodiment, the data generation module is based on a large language model, and the total model logical architecture in the embodiment is described with reference to Figure 3 .

[0023] Step S02: Extract the label data in the chest X-ray image-label data set and the chest X-ray image-text data set, construct a multi-round question and answer dictionary according to the label data, and obtain global description text; It should be noted that in the embodiment, the specific logical framework of data generation is described with reference to Figure 4 , the label data includes normal labels and lesion labels, the multi-round question and answer dictionary is based on normal part description text and lesion part description text, and the global description text is generated based on local description text. Extract the label data in the chest X-ray image-label data set and the chest X-ray image-text data set, and the label data includes normal labels and lesion labels; Generate normal part description text according to the standard anatomical part list and the normal label; Generate lesion part description text according to the disease category-disease anatomical part mapping relationship table and the lesion label; Construct a multi-round question and answer dictionary according to the normal part description text and the lesion part description text, and the multi-round question and answer dictionary is constructed based on a large language model; According to the multi-round question and answer dictionary, the description text is expanded to obtain a plurality of local description texts, and the report text is supplemented according to the local description text to obtain global description text.

[0024] Step S03: Perform image enhancement processing on the image data in the chest X-ray image-label data set and the chest X-ray image-text data set, and then input the image data after image enhancement processing and the global description text into the image perception pre-training module; It should be noted that in the embodiment, the image perception pre-training module is constructed based on a global-local architecture, and the specific logical framework of the image perception pre-training module for image construction in the embodiment is described with reference to Figure 5 .

[0025] Step S04: Extract text features from the global description text according to the text encoder to obtain global text features and local text features, respectively, and extract image features from the image data according to the image encoder to obtain global visual features; Step S05: performing global contrast learning according to the global text feature and the global visual feature, and performing graph perception learning according to the global text feature and the local text feature, to obtain a global contrast learning enhanced feature and a graph perception enhanced feature respectively; It should be noted that in the embodiment, the graph structure is constructed according to the global text feature, the local text feature and the global visual feature, the graph structure includes nodes and edges, the nodes include a global text feature node, a local text feature node and a global visual feature node, the local text feature node is a bridge node, and the bridge node is used to connect the global text feature node and the global visual feature node. Performing local perception enhancement processing according to the graph convolutional neural network and the graph structure; Taking the disease label Atelectasis as an example, the graph construction process is decomposed to obtain a disease anatomical site set of the disease label Atelectasis: , Among them, denotes the disease anatomical site set of the disease label Atelectasis, denote the disease anatomical sites lung, mediastinum, pleura and airway respectively; Then, a local text feature group of the disease label Atelectasis is obtained: , Among them, denotes the local text feature group of the disease label Atelectasis, denote the local text features of the disease anatomical sites lung, mediastinum, pleura and airway respectively; The graph structure and the adjacency matrix of the disease label Atelectasis are constructed: , , Among them, denotes the graph structure of the disease label Atelectasis, denotes the node of the graph structure, denotes the edge of the graph structure, denote the local text feature node, the global text feature node and the global visual feature node of the disease label Atelectasis respectively, denotes the adjacency matrix of the disease label Atelectasis; The global text features, local text features and global visual features are constructed as nodes of a graph structure, and edges of the graph structure are constructed according to the types of the nodes, the edges of the graph structure including local text feature bidirectional edges, global text-local text feature unidirectional edges, global visual-local text feature unidirectional edges, and global text-global visual feature bidirectional edges, the local text feature bidirectional edges being used for local text feature nodes to obtain enhanced information representation, the global text-local text feature unidirectional edges being used for global-local text information supplement, the global visual-local text feature unidirectional edges being used for cross-modal advanced semantic information fusion, and the global text-global visual feature bidirectional edges being used for cross-modal global information fusion; According to the graph convolutional neural network and the graph structure, local perception enhancement processing is performed, the graph convolutional neural network including two graph convolutional layers and one ReLU activation layer, and a graph convolution algorithm of the graph convolutional layer is specifically as follows: , , wherein, denotes a node feature matrix, denotes a unit matrix, denotes an activation function, denotes a degree matrix of, denotes a parameter matrix trainable for a convolutional layer, denotes an adjacency matrix; A specific algorithm of the graph convolutional neural network is as follows: , wherein, denotes a graph convolutional neural network, denotes a graph convolution operation, denote a local text feature node, a global text feature node and a global visual feature node respectively.

[0026] Step S06: Perform cross-modal similarity comparison loss optimization on the global contrast learning enhanced features and the graph perception enhanced features. It should be noted that a specific algorithm of the cross-modal similarity comparison loss optimization in the embodiment is as follows: , wherein, denotes a cross-modal loss, denotes a cosine similarity calculation, denote a local text feature node, a global text feature node and a global visual feature node respectively.

[0027] In summary, according to the chest X-ray multi-modal pre-training method based on graph perception learning, the data generation is performed by constructing a multi-turn question and answer dictionary, the local descriptive text of each lesion site in the chest X-ray is generated from three aspects of disease classification, classification certainty and corresponding lesion site, and the global descriptive text is automatically generated, the problem of insufficient data is effectively avoided, and the quality and consistency of the text description are improved, and through the graph perception pre-training, the correlation graph structure between the local and global features is constructed based on the graph perception learning method from the global to the local, the cross-modal correlation between the chest X-ray sites and diseases is deeply mined, the small visual differences that are difficult to identify are more accurately captured, and the modal difference between the image and the text is reduced, and the precision and generalization ability of analyzing the chest X-ray are improved. Specifically, a chest X-ray image-label data set and a chest X-ray image-text data set are obtained and input into a data generation module, the data generation module extracts label data in the chest X-ray image-label data set and the chest X-ray image-text data set based on a large language model, constructs a multi-turn question and answer dictionary according to the label data to obtain global descriptive text, the label data includes normal labels and lesion labels, the multi-turn question and answer dictionary is based on normal site descriptive text and lesion site descriptive text, and the global descriptive text is generated based on local descriptive text. Through the construction of the multi-turn question and answer dictionary, the data generation is performed from three aspects of disease classification, classification certainty and corresponding lesion site, the local descriptive text of each lesion site in the chest X-ray is generated, and the global descriptive text is automatically generated, the problem of insufficient data is effectively avoided, and the quality and consistency of the text description are improved. The image data in the chest X-ray image-label data set and the chest X-ray image-text data set is subjected to image enhancement processing, and then the image data and the global descriptive text are input into a graph perception pre-training module, the graph perception pre-training module is constructed based on a global-local architecture, text feature extraction is performed on the global descriptive text by a text encoder to obtain global text features and local text features, image feature extraction is performed on the image data by an image encoder to obtain global visual features, global contrast learning is performed according to the global text features and the image features, graph perception learning is performed according to the global text features and the local text features to obtain global contrast learning enhanced features and graph perception enhanced features, and cross-modal similarity contrast loss optimization is performed according to the global contrast learning enhanced features and the graph perception enhanced features. Through the graph perception pre-training, the correlation graph structure between the local and global features is constructed based on the graph perception learning method from the global to the local, the cross-modal correlation between the chest X-ray sites and diseases is deeply mined, the small visual feature differences that are difficult to identify are more accurately captured, and the modal difference between the image and the text is reduced, and the precision and generalization ability of analyzing the chest X-ray are improved.

[0028] Please refer toFigure 2 Figure 2 shows a structural schematic diagram of a chest X-ray multi-modal pre-training system based on image perception learning according to a second embodiment of the present application, which comprises: A data acquisition module 10 is configured to acquire a chest X-ray image-label data set and a chest X-ray image-text data set and input the data into a data generation module, and the data generation module is based on a large language model; The data generation module 20 is configured to extract label data in the chest X-ray image-label data set and the chest X-ray image-text data set, construct a multi-turn question and answer dictionary according to the label data to obtain global description text, the label data includes normal labels and lesion labels, the multi-turn question and answer dictionary is based on normal part description text and lesion part description text, and the global description text is generated based on local description text; The image enhancement module 30 is configured to perform image enhancement processing on image data in the chest X-ray image-label data set and the chest X-ray image-text data set, and then input the image data after the image enhancement processing and the global description text into an image perception pre-training module, and the image perception pre-training module is based on a global-local architecture; The image perception pre-training module 40 is configured to perform text feature extraction on the global description text according to a text encoder to obtain global text features and local text features respectively, and perform image feature extraction on the image data according to an image encoder to obtain global visual features; Global contrast learning is performed according to the global text features and the global visual features, and image perception learning is performed according to the global text features and the local text features to obtain global contrast learning enhanced features and image perception enhanced features respectively; Cross-modal similarity contrast loss optimization is performed according to the global contrast learning enhanced features and the image perception enhanced features.

[0029] The present application further provides a computer storage medium having one or more programs stored thereon, which programs, when executed by a processor, implement the above-mentioned chest X-ray multi-modal pre-training method based on image perception learning.

[0030] The present application further provides a computer device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the above-mentioned chest X-ray multi-modal pre-training method based on image perception learning.

[0031] Those skilled in the art will appreciate that the logic and / or steps represented in the flow diagrams, or otherwise described herein, can be embodied in

[0032] More specific examples (a non-exhaustive list) of the computer readable medium include the following: a portable computer diskette (magnetic device); a random access memory (RAM); a read-only memory (ROM); an erasable programmable read-only memory (EPROM or Flash memory); an optical fiber device; and a portable compact disc read-only memory (CDROM). Additionally, the computer readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.

[0033] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, can be used: a combination of discrete logic circuits having logic gates for implementing logic functions upon an application of data signals; application specific integrated circuits having logic gates; field programmable gate arrays (FPGAs), programmable logic arrays (PLAs); and / or other implementations such as formed using transparent X (TX) devices.

[0034] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific feature, structure, material or characteristic being described is included in at least one embodiment or example of the application. The illustrative descriptions of the above terms in the specification do not necessarily refer to the same embodiment or example. Moreover, the description of the specific features, structures, materials or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0035] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A chest X-ray multi-modal pre-training method based on graph perception learning, characterized in that, The method comprises the following steps: acquiring a chest X-ray image-label data set and a chest X-ray image-text data set and inputting the data into a data generation module, the data generation module being based on a large language model; extracting label data in the chest X-ray image-label data set and the chest X-ray image-text data set, constructing a multi-round question and answer dictionary according to the label data to obtain global description text, the label data comprising normal labels and lesion labels, the multi-round question and answer dictionary being based on normal part description text and lesion part description text, and the global description text being generated based on local description text; performing image enhancement processing on image data in the chest X-ray image-label data set and the chest X-ray image-text data set, and then inputting the image data after the image enhancement processing and the global description text into a graph perception pre-training module, the graph perception pre-training module being constructed based on a global-local architecture; extracting text features from the global description text according to a text encoder to obtain global text features and local text features respectively, and extracting image features from the image data according to an image encoder to obtain global visual features; performing global contrast learning according to the global text features and the global visual features, and performing graph perception learning according to the global text features and the local text features to obtain global contrast learning enhanced features and graph perception enhanced features respectively; performing cross-modal similarity contrast loss optimization according to the global contrast learning enhanced features and the graph perception enhanced features.

2. The chest X-ray multi-modal pre-training method based on graph perception learning according to claim 1, characterized in that, The step of extracting label data in the chest X-ray image-label data set and the chest X-ray image-text data set and constructing a multi-round question and answer dictionary according to the label data to obtain global description text specifically comprises the following steps: extracting label data in the chest X-ray image-label data set and the chest X-ray image-text data set, the label data comprising normal labels and lesion labels; generating normal part description text according to a standard anatomical part list and the normal labels; generating lesion part description text according to a disease category-lesion anatomical part mapping relationship table and the lesion labels; constructing a multi-round question and answer dictionary according to the normal part description text and the lesion part description text, the multi-round question and answer dictionary being constructed based on a large language model; performing description text expansion according to the multi-round question and answer dictionary to obtain a plurality of local description texts, and then performing report text supplement according to the local description texts to obtain global description text.

3. The chest X-ray multi-modal pre-training method based on graph perception learning according to claim 1, characterized in that, The step of constructing a multi-round question and answer dictionary according to the label data further comprises the following steps after the step: in the first round of questioning, the multi-round question and answer dictionary generates a standard anatomical part list according to a standard chest X-ray report, and the standard anatomical parts in the standard anatomical part list correspond one by one to the normal labels; in the second round of questioning, the multi-round question and answer dictionary generates a disease category-lesion anatomical part mapping relationship table according to disease categories and lesion labels; in the third round of questioning, the multi-round question and answer dictionary generates normal part description text according to a large language model to describe standard anatomical parts, and the normal part description text corresponds one by one to the standard anatomical parts. In the fourth round of question and answer, the multi-round question and answer dictionary generates description text of the lesion site according to the large language model to obtain the lesion site description text, which is one-to-one corresponding to the disease category; In the fifth round of question and answer, the multi-round question and answer dictionary is iteratively optimized to generate an extended description text list for each disease category according to the large language model and form a final multi-round question and answer dictionary.

4. The chest X-ray multi-modal pre-training method based on graph perception learning according to claim 1, characterized in that, The step of performing graph perception learning according to the global text features and the local text features specifically includes: A graph structure is constructed according to the global text features, the local text features and the global visual features, the graph structure including nodes and edges, the nodes including global text feature nodes, local text feature nodes and global visual feature nodes, the local text feature nodes being bridge nodes, and the bridge nodes being used to connect the global text feature nodes and the global visual feature nodes; Local perception enhancement processing is performed according to the graph convolutional neural network and the graph structure.

5. The chest X-ray multi-modal pre-training method based on graph perception learning according to claim 4, characterized in that, The step of constructing a graph structure according to the global text features, the local text features and the global visual features specifically includes: The global text features, the local text features and the global visual features are constructed as nodes of the graph structure, and the edges of the graph structure are constructed according to the types of the nodes, the edges of the graph structure including local text feature bidirectional edges, global text-local text feature unidirectional edges, global visual-local text feature unidirectional edges and global text-global visual feature bidirectional edges, the local text feature bidirectional edges being used for the local text feature nodes to obtain enhanced information representation, the global text-local text feature unidirectional edges being used for global-local text information supplement, the global visual-local text feature unidirectional edges being used for cross-modal high-level semantic information fusion, and the global text-global visual feature bidirectional edges being used for cross-modal global information fusion.

6. The chest X-ray multi-modal pre-training method based on graph perception learning according to claim 4, characterized in that, The step of performing local perception enhancement processing according to the graph convolutional neural network and the graph structure specifically includes: Local perception enhancement processing is performed according to the graph convolutional neural network and the graph structure, the graph convolutional neural network including two graph convolutional layers and one ReLU activation layer, and the graph convolution algorithm of the graph convolutional layer being specifically as follows: , , wherein, denotes a node feature matrix, denotes an identity matrix, denotes an activation function, denotes a degree matrix of, denotes a parameter matrix trainable by a convolution layer, denotes an adjacency matrix; The specific algorithm of the graph convolutional neural network is as follows: , wherein, denotes a graph convolutional neural network, denotes a graph convolution operation, denote a local text feature node, a global text feature node, and a global visual feature node, respectively.

7. The chest X-ray multi-modal pre-training method based on graph perception learning according to claim 1, characterized in that, The step of performing cross-modal similarity contrast loss optimization according to the global contrast learning enhanced features and the graph perception enhanced features specifically includes: The specific algorithm of the cross-modal similarity contrast loss optimization is as follows: , wherein, denotes a cross-modal loss, denotes a cosine similarity computation, denote a global text feature node and a global visual feature node, respectively.

8. A chest X-ray multi-modal pre-training system based on graph perception learning, characterized in that, It includes: A data acquisition module is configured to acquire a chest X-ray image-label dataset and a chest X-ray image-text dataset and input the datasets to a data generation module, the data generation module being based on a large language model; The data generation module is configured to extract label data in the chest X-ray image-label dataset and the chest X-ray image-text dataset, construct a multi-round question and answer dictionary according to the label data to obtain global description text, the label data including normal labels and lesion labels, the multi-round question and answer dictionary being based on normal site description text and lesion site description text, and the global description text being generated based on local description text; An image enhancement module is configured to perform image enhancement processing on image data in the chest X-ray image-label dataset and the chest X-ray image-text dataset, and then input the image data after the image enhancement processing into the global description text input image perception pre-training module, which is constructed based on a global-local architecture; The image perception pre-training module is configured to perform text feature extraction on the global description text according to a text encoder to obtain global text features and local text features respectively, and perform image feature extraction on the image data according to an image encoder to obtain global visual features; Global contrast learning is performed according to the global text features and the global visual features, and image perception learning is performed according to the global text features and the local text features to obtain global contrast learning enhanced features and image perception enhanced features respectively; Cross-modal similarity contrast loss optimization is performed according to the global contrast learning enhanced features and the image perception enhanced features.

9. A storage medium, characterized by The storage medium stores one or more programs, which are executed by the processor to implement the chest X-ray multi-modal pre-training method based on image perception learning according to any one of claims 1-7.

10. A computer device, comprising: The computer device comprises a memory and a processor, wherein: The memory is configured to store a computer program; The processor is configured to execute the computer program stored in the memory to implement the chest X-ray multi-modal pre-training method based on image perception learning according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-modal large model-based traditional Chinese medicine tongue diagnosis analysis system and method

    CN118899077A

  • Remote sensing question answering system based on comparison-generation type pre-training model

    CN119445394A

  • Fine-grained multi-mode prompt learning method based on visual language pre-training model

    CN119538179A

  • Video question and answer method and system based on multi-level alignment

    CN120104831A