The invention discloses an
image retrieval method, device and equipment based on multi-
modal semantics, which are applied to the technical field of multi-
modal retrieval. According to the
image retrieval method based on the multi-
modal semantics, a retrieval request comprising a
reference image and a modified text is obtained; based on the
reference image and the modified text, a first global image feature including visual information of the
reference image, an object feature including information of the modified object, and a description feature including
semantic information of the modified text are extracted, respectively. Therefore, relatively complete visual information and
semantic information can be extracted. And integrating the first global image feature, the object feature and the description feature to obtain a retrieval feature. And finally, determining a target image by utilizing the retrieval features, and generating a
retrieval result. According to the method, retrieval is carried out by utilizing the retrieval characteristics fusing the visual information and the
semantic information, so that the retrieval capability of cross-modal information can be enhanced, the accuracy and the effectiveness of the target image obtained through retrieval are improved, and the retrieval requirements of a user in a multi-modal
image retrieval scene are met.