Embodiments of the present application disclose
image detection methods, model
training methods and devices, equipment, storage media and products. The
image detection method comprises: obtaining a to-be-detected image and a query text, on the one hand, performing open-vocabulary
object detection on the to-be-detected
image based on the query text to obtain a set of object position indication information, and on the other hand, extracting global features of the to-be-detected image; then, according to the obtained global features and the set of object position indication information, local features of the to-be-detected image are extracted, and according to the global features, the local features and the query text, a target detection result of the query text in the to-be-detected image is obtained. It can be seen that by decoupling low-level positioning (obtaining the set of object position indication information) and high-level understanding (extracting global features), and performing open-vocabulary
object detection on the to-be-detected image in the low-level positioning process, the object position indication information in the set of object position indication information can not be limited to fixed categories, thereby realizing target detection of the image.