The invention relates to a multi-
modal AI data fusion
processing method, device and equipment and a medium, and the method comprises the steps: firstly extracting visual, auditory and text
modal features through a pre-training
encoder, executing dimension alignment, and generating a standard data
feature set with unified dimensions; a cross-
modal semantic graph is constructed based on a
cosine similarity algorithm, and the problem of semantic mismatch of heterogeneous data is solved; residual enhancement is carried out on the map nodes, and
noise interference is eliminated; fusing the optimized features and the semantic topology in combination with a graph convolutional network to generate aggregation graph representation; the fusion features are mapped to a low-dimensional
semantic space through a variational auto-
encoder, and cross-modal correlation essence is captured; the key dimension contribution degree is quantified, a visual report is generated, and
semantic association rules among modals are disclosed, so that the dimension isomerism limitation of a traditional fusion technology is broken through, quantifiable cross-modal
semantic mapping is established, the whole process
traceability from
feature fusion to decision interpretation is realized, and the method is suitable for popularization and application. And the multi-modal decision
black box problem in the fields of
medical diagnosis, automatic driving and the like is effectively solved.