The application discloses a kind of accessory identification method, the method first obtains accessory image and extracts text information, by inquiring preset accessory
electronic catalog, fast matching accessory with identifiable text, improve identification efficiency.When there is no matching first candidate accessory or the number of matches is more than 2, extract image
feature vector from accessory image, filter multiple second candidate accessories containing image and text information, cover no text,
text recognition invalid or matching not unique scene, solve the problem of narrow application range of single text extraction technology.Secondly, input the accessory image and the second candidate accessory information into the trained visual-
language model, and use its graphic-text joint understanding ability to accurately distinguish similar accessories, reduce the mismatch rate, and make up for the defects of insufficient accuracy of single technology.Overall, through progressive multi-
modal collaborative identification, the accuracy, consistency and efficiency of accessory identification are significantly improved, ensuring smoothness of accessory transaction and maintenance, and optimizing the user experience.