The invention discloses a
power equipment image-text fusion labeling method and
system based on a single-double mixing
tower, and the method comprises the steps: carrying out the visual
feature extraction to obtain image features, carrying out the text
feature extraction to obtain text features, mining the deep features of an image and a text, maintaining the
modal specificity, and carrying out the recognition of the image and the text.
Processing the acquired image features and text features by adopting a cross attention mechanism to generate a bidirectional attention matrix, calculating a dynamic
weight value based on the acquired bidirectional attention matrix, and generating weighted image features and weighted text features based on the dynamic
weight value; and performing
dynamic feature fusion on the weighted image feature and the weighted text feature to obtain a fusion result feature, and performing end-to-end multi-
modal labeling based on the fusion result feature, so that the image and the text can be accurately associated, the labeling efficiency and accuracy are improved, adaptive
feature fusion can be realized, the real-time problem of heterogeneous
feature fusion is solved, and the real-time performance of the heterogeneous
feature fusion is improved. The time consumed by multi-
modal alignment is reduced from the minute level of manual intervention to the
millisecond level.