The invention discloses an image-text
pathology recognition method and device based on multi-
modal fusion, and the method comprises the steps: inputting an image block sequence into a UNI-Patch
encoder, outputting the microscopic features of a fixed dimension, segmenting the image block sequence into a plurality of instances, inputting the instances into an MIL
mask generator, learning and generating a feature
mask, and carrying out the recognition of the feature
mask. The method comprises the following steps of: performing point product operation on a feature mask block and a microscopic feature to generate an image block feature, aggregating the image block feature by adopting a Perceeiver-Slide aggregator to generate a slice-level global feature, inputting a text sequence into a BERT-Text
encoder, outputting a text
semantic feature, and performing point product operation on the microscopic feature, the slice-level global feature and the text
semantic feature through a projection layer to generate an image block feature. And after uniformly mapping to a
semantic space with the same dimension, inputting into a CoCa trainer for training, and generating a
pathological report. According to the method, deep semantic fusion of the
pathological image and the clinical text can be realized, and the diagnosis efficiency and accuracy are improved.