The application discloses an image semantic segmentation model which is composed of a designed mixed attention focusing method, an attention rectification residual module (ARRM) and a mixed feature integration module (MFFM), is trained in a deep supervision mode as a whole, is reasonably provided with deformable
convolution, is combined with a constructed multi-scale spatial attention module (MSP) and a double-
pooling attention module (DPA) to be jointly optimized, and the problem of
small target feature segmentation difficulty is solved. In order to reflect the excellent performance of the model on a downstream task, the training between different data sets is completed in a transfer learning mode, and the application range of the model is expanded. Finally, grouping
convolution is added in the
backbone network, the calculation cost is greatly reduced, and the problem of high training cost of the segmentation model is solved under the premise of guaranteeing the segmentation effect.