The application discloses a text-video cross-
modal event element extraction method, collects video data and video introduction text data thereof, respectively labels event types and corresponding event argument roles of the text and video data, wherein the event argument role represents entities playing different roles in events, the event types and the event argument roles corresponding to the event types of the text data and the video data are pre-labeled to be consistent; multi-
modal event reference resolution is carried out to realize co-reference event
pairing between any "text-video" data, that is, text and video referring to the same event are matched to form a text-video co-reference
event pair; the matched "text-video" data is converted into corresponding feature vectors, wherein text tokenization and text embedding are performed on the text data to be converted into a word vector form; ResNet
algorithm is directly used on the video data to obtain global level event element features to construct a video global
feature vector; Fast-R-CNN is used on the video data to identify local objects, ResNet
algorithm is used to obtain local level time elements to construct a video local
feature vector; the text word vector and the video global
feature vector and the local feature vector are unified in vector dimension through a full connection layer to construct a text-video shared vector space; the text word vector and the video global feature vector and the local feature vector are input into a
Transformer encoder, then ONEIE
algorithm is adopted to extract event element information of the
text mode, and T5-base algorithm is adopted to extract event element information of the video mode. The application can more accurately capture the correlation between the internal of multi-
modal, and improves the extraction accuracy.