The invention discloses an RGB-T semantic segmentation method and
system based on multi-attention guidance and hierarchical fusion, and relates to the technical field of
image processing. The
system is composed of a double-flow
encoder, a discriminative local texture
perception unit, a semantic-driven cross-
modal fusion unit, a
semantic enhancement unit and a multi-scale layered refinement decoder, and efficient fusion and analysis of multi-
modal features in a complex
traffic scene are achieved. According to the discriminative local texture
perception method, saliency features are learned through multi-attention guidance and a self-adaptive gating mechanism, accurate modeling of shallow texture information is focused, and the distinguishing ability of a
region of interest and a target edge is improved; according to the semantic-driven cross-
modal feature fusion method, efficient aggregation of global contexts is realized through high-level semantic guidance and cross-modal feature interaction, and feature complementarity is enhanced, so that the semantic-driven cross-modal
feature fusion method has significant advantages in analysis of small targets, long-distance targets and boundary regions. The decoder adopts a progressive fusion mode, an additional
edge detection module does not need to be added, and the overall segmentation precision is improved.