The invention provides a text
image retrieval model training method,
system and device and a storage medium, and is applied to the technical field of medical
image retrieval, and the method comprises the following steps: carrying out multi-level cross-
modal alignment relation extraction on a to-be-processed video;
gaussian noise addition is carried out on continuous picture frames in the video to be processed, and the picture frames after
Gaussian noise addition are
cut into a plurality of image blocks; training an image
encoder by taking the image block feature corresponding to the current image block, the
time sequence feature of the current image block corresponding to the previous frame of image block and the space-time position coding feature corresponding to the current image block as input features; positive and negative samples corresponding to the multi-level cross-
modal alignment relationship are coded, generated text coding features and image coding features are mapped to the same
semantic space, and parameters of a text and
image retrieval model are optimized, so that the
recall rate of the retrieval model is increased, the accuracy of a
retrieval result is ensured, and the retrieval efficiency is improved. And reliable video
information support is provided for medical decision and
medical research.