The invention discloses a webpage element similarity detection method based on adaptive clustering, which comprises the following steps: S1,
input field XPath processing: receiving an
XPath path of at least one target field as input, and verifying the format validity of the
XPath path; s2, element path extraction: positioning a corresponding webpage
HTML (
Hypertext Markup Language) element according to a verified XPath path, and extracting a complete
label path of the
HTML element from a node of the
HTML element to a webpage root node; s3, calculating the similarity; s4, distribution analysis: constructing a symmetric
similarity matrix for the comprehensive similarity of all the element pairs obtained in the step S3, and performing statistical distribution analysis on comprehensive similarity values in the
similarity matrix by adopting
kernel density estimation; s5, self-adaptive clustering: performing multi-target clustering on the HTML elements based on the distribution
peak value and the salient interval identified in the step S4; s6, generating an XPath (X Path); and S7, outputting: outputting the optimized general
XPath expression. The method has the advantages of high adaptability, high accuracy, good robustness, high efficiency, high universality and the like.