The invention relates to the field of
image processing, and provides a chest X-
ray film multi-
modal pre-training method and
system based on graph
perception learning, and the method comprises the steps: carrying out the data generation through constructing a multi-round question and answer dictionary, generating a local descriptive text of each
lesion part in a chest X-
ray film from the three levels of
disease classification, classification certainty and corresponding
lesion parts, and carrying out the recognition of the local descriptive text. According to the method, a global description text is automatically generated, the problem of insufficient data is effectively avoided, the quality and consistency of text description are improved, and a correlation graph structure between local and global features is constructed through graph
perception pre-training based on a global-to-local graph
perception learning method, so that the accuracy of text description is improved. The cross-
modal relevance between each part of the
chest radiograph and a
disease is deeply mined, tiny visual differences which are difficult to recognize are more accurately captured,
modal differences between images and texts are reduced, and the accuracy and generalization ability of chest X-
ray radiograph analysis are improved.