The invention relates to the technical field of data fusion, and discloses an animal scene-oriented adaptive multi-
modal data fusion method, which comprises the following steps of: extracting spatio-temporal characteristics from multi-source heterogeneous data such as visual sense,
auditory sense and physiological sensing, constructing an animal-environment-group ternary spatio-temporal
relation graph, and constructing an animal-environment-group ternary spatio-temporal
relation graph; a
pilot frequency sampling problem is solved through an
adaptive interpolation algorithm, cross-
modal projection alignment is completed in a public
semantic space, unified space-time representation is output, and confidence coefficient weight is dynamically calculated based on uncertainty measurement of each
modal feature. According to the method, accurate alignment of multi-
modal data is realized through the cross-modal space-time
attention network, the multi-modal feature alignment error is reduced compared with that of a traditional LSTM method, the training data volume of a federated element migration
reinforcement learning framework is reduced compared with that of a traditional migration learning method, and the cross-species generalization performance of the model is improved. A multi-level
causal inference engine quantitatively reveals causal association between environmental factors and animal diseases, and in combination with a dynamic
decision tree visualization technology, the decision recognition degree is improved.