This invention discloses an
image fusion method based on shifted window attention and semantic-driven dual adversarial approach. The method includes the following steps: Step 1, constructing an adaptive
feature extraction network based on a dual-
stream shifted window
Transformer as a generator, utilizing a cross-window attention mechanism to achieve deep interaction and fusion of
infrared image features and visible light image features; Step 2, constructing a multi-style dual
discriminator decoupled from target and texture, introducing a target
mask, and establishing brightness discrimination paths for salient
infrared targets and
texture gradient discrimination paths for visible light backgrounds; Step 3, introducing a semantic-driven meta-feature embedding feedback mechanism, utilizing a pre-trained target detection network to extract high-level semantic features, and constructing a
semantic consistency loss to guide generator parameter updates; Step 4, constructing a joint objective function including content loss, dual adversarial loss, and semantic loss based on a
mask weighting strategy, and performing end-to-end training on the network; Step 5, inputting the image to be fused into the trained generator, and outputting the fused image. This invention solves the problems of limited
feature extraction and
modal information conflict in traditional visible light and
infrared fusion methods, significantly improving the detection accuracy of the fused image in
machine vision tasks while balancing infrared high-brightness targets and clear background textures.