The invention relates to the technical field of
Internet television services, in particular to an OTT visual
feature extraction system and method based on multi-
modal agent driving, and the method comprises the steps: capturing a screen real-time video
stream of equipment;
processing the target advertisement image and the real-time video
stream, and extracting double-flow heterogeneous visual features, including global content
perception features and local geometric structure features, through a multi-
modal visual perception model; executing a hierarchical matching
algorithm, calculating and screening out candidate frames by using global content
perception features, matching in the candidate frames by using local geometric structure features to establish a corresponding relation set containing all matched initial key points, performing
spatial clustering on the set to separate out advertisement instances, and obtaining bounding boxes of the instances through
geometric transformation calculation; and according to the bounding box, performing highlight display on the area where the target advertisement is located on the original video frame to generate a visual broadcast monitoring result. According to the invention, through multi-mode intelligent body driving, OTT advertisement visual
feature extraction and broadcast monitoring are realized.