The present application belongs to the technical field of slice image recognition, and particularly relates to a
cervical cell pathological slice recognition method based on multi-
modal learning, which comprises the following steps: acquiring cervical
pathological slice images and clinical text data of a patient, extracting image
modal feature vectors and text
modal feature vectors through an image
encoder and a text
encoder respectively; inputting the two modal features into a cross-modal fusion network constructed based on an asymmetric attention mechanism, guiding image feature enhancement with the text feature as a query, guiding text feature enhancement with the image feature as a query, and fusing to generate multi-modal joint feature representation; and outputting classification results such as normal cells, low-grade lesions, high-grade lesions or squamous
cell carcinoma based on the joint feature representation. Through bidirectional cross-modal attention interaction, the present application realizes effective fusion of
pathological image morphological features and clinical text information, and makes up for the limitations of single-
modal data.