The invention discloses a
web crawler defense method based on an http protocol version, and particularly relates to the technical field of
web crawler defense. The method comprises the following steps: performing
data acquisition and clustering on network requests entering a
server to form an original request sample
data set, extracting an HTTP protocol version number and TLS negotiation information based on the original request sample
data set, comparing a UA identifier with version consistency, calculating version consistency parameters, and determining the version consistency of the
server according to the version consistency parameters. The method comprises the following steps: extracting communication characteristics of third-party service
callback, mobile terminal sharing and
payment gateway requests in combination with
server service logs, establishing an HTTP protocol version
distribution model, constructing a joint identification matrix of HTTP protocol versions and service characteristics, calculating version
abnormality scores, performing dynamic classification to identify suspicious crawler requests, generating abnormal request identification data, and sending the abnormal request identification data to a server; and finally, triggering a multi-layer interception
mechanism based on the white
list and
gray level defense. According to the invention, the accuracy of
web crawler detection and the flexibility of defense are effectively improved.