The invention relates to the technical field of medical image and
artificial intelligence crossing, and discloses a cross-
modal-based VMama medical
image fusion method and
system combining grouped ACmix
convolution and selective clustering, and the
system comprises an input module, a preprocessing module, a
feature extraction module, a multi-scale fusion module, and an output reconstruction module. According to the scheme, cross-
modal attention is introduced into a visual
state space model for the first time, a new multi-
modal image fusion framework LMACV is obtained, the framework performs linkage optimization on ACmix and VMamba structures in a cross-modal medical image for the first time, feature reconstruction efficiency is enhanced through a selective clustering mechanism, and MSE and PSNR are remarkably superior to existing methods such as MPCT, FATFusion and MATR. By fusing a convolutional network and
state space modeling, the network can effectively capture local texture details, and meanwhile, long-distance
semantic association is reserved; besides, linear state updating and cross-modal dynamic alignment of attention guidance are effectively realized, and compared with the latest MPCT
algorithm, the fusion speed is improved by 37.5%.