Request identification method, device, and apparatus, and storage medium

CN116910331BActive Publication Date: 2026-09-29CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211600556.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2026-09-29
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

因此,相关技术中对爬虫行为的检测方法,至少存在漏识别的问题

Benefits of technology

[0020]本申请实施例提供一种请求的识别方法、请求的识别装置、电子设备及计算机可读存储介质,该方法包括:获取用户设备发起的至少一个访问请求,并确定每个访问请求的主题;基于主题对至少一个访问请求进行聚类处理,得到至少一个主题簇;基于各个主题簇中的访问请求的数量,确定用户设备发起的访问请求所表征的行为类型;其中,行为类型用于表征用户设备的访问行为是否为自动抓取万维网信息的程序或者脚本的行为。也就是说,本申请对用户设备发起的访问请求的内容进行追踪和分析,通过分析访问请求的资源内容是否主题相关来识别爬虫请求,实现爬虫请求的精准识别,解决了相关技术中对爬虫行为的检测方法至少存在漏识别的问题;并对识别到的爬虫请求进行封禁和限制,保障了网站系统中数据的安全,减少爬虫请求对网站系统的服务器的攻击,降低网络带宽的消耗。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116910331B_ABST
    Figure CN116910331B_ABST
Patent Text Reader

Abstract

The application discloses a request identification method, which comprises the following steps: obtaining at least one access request initiated by a user equipment, and determining the subject of each access request; performing clustering processing on the at least one access request based on the subject, and obtaining at least one subject cluster; determining the behavior type represented by the access request initiated by the user equipment based on the number of access requests in each subject cluster; wherein the behavior type is used to represent whether the access behavior of the user equipment is the behavior of a program or script for automatically grabbing information on the World Wide Web. The application also discloses a request identification device, an electronic equipment and a computer readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of communications, and particularly to a method for identifying a request, a device for identifying a request, an electronic device, and a computer-readable storage medium. Background Technology

[0002] A web crawler is a program or script that automatically retrieves information from the World Wide Web according to specific rules. It is widely used in data mining, public opinion analysis, search engines, and other business fields. Web crawlers typically start by crawling a list of seed pages, then iterate through the details page links in the requests to obtain the details page responses and extract the target information. Currently, some illegal web crawlers exist that use program requests to obtain core data or sensitive information from website systems in bulk, posing a security risk of information leakage. Therefore, website systems need to have the ability to identify web crawler requests.

[0003] In related technologies, the following three methods are commonly used to detect web crawler behavior: First, traffic detection, which involves statistically analyzing the traffic range of the Internet Protocol (IP) address where the web crawler request originates. When the traffic exceeds a threshold, it is identified as a web crawler request. Second, frequency detection, which involves statistically analyzing the request frequency of the account used by the web crawler request. When the request frequency exceeds a threshold, it is identified as a web crawler request. Third, request header detection, which involves detecting and verifying the request header data of the web crawler request, such as fields like User-Agent (UA), Cookies, and Referer. When these fields are missing or abnormal, it is identified as a web crawler request.

[0004] However, with the development of web crawling technology, more and more crawler programs can impersonate a near-real user, bypassing traffic and frequency detection by making crawling requests at low, random intervals. Furthermore, they can bypass browser / request header field detection by disguising headers and simulating browsers. Therefore, existing methods for detecting crawler behavior suffer from at least a problem of missed detection. Summary of the Invention

[0005] This application provides a request identification method, a request identification device, an electronic device, and a computer-readable storage medium.

[0006] The technical solution of this application is implemented as follows:

[0007] In a first aspect, embodiments of this application provide a method for identifying a request, the method comprising:

[0008] Obtain at least one access request initiated by the user equipment, and determine the subject of each access request;

[0009] Based on the topic, the at least one access request is clustered to obtain at least one topic cluster;

[0010] Based on the number of access requests in each of the aforementioned topic clusters, the behavior type represented by the access request initiated by the user equipment is determined; wherein, the behavior type is used to characterize whether the access behavior of the user equipment is the behavior of a program or script that automatically crawls World Wide Web information.

[0011] Secondly, embodiments of this application provide a request identification device, wherein the information processing device includes:

[0012] The acquisition module is used to acquire at least one access request initiated by the user device;

[0013] A processing module is used to determine the subject of each access request;

[0014] The processing module is further configured to perform clustering processing on the at least one access request based on the topic to obtain at least one topic cluster;

[0015] The processing module is further configured to determine the behavior type represented by the access request initiated by the user equipment based on the number of access requests in each of the topic clusters; wherein, the behavior type is used to characterize whether the access behavior of the user equipment is the behavior of a program or script that automatically crawls World Wide Web information.

[0016] Thirdly, an electronic device provided in this application includes: a processor, a memory, and a communication bus;

[0017] Memory, used to store executable instructions;

[0018] A processor for executing executable instructions stored in the memory to implement the steps of the identification method requested above.

[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing executable instructions, wherein the computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the identification method requested above.

[0020] This application provides a request identification method, a request identification device, an electronic device, and a computer-readable storage medium. The method includes: acquiring at least one access request initiated by a user device and determining the topic of each access request; clustering the at least one access request based on the topic to obtain at least one topic cluster; and determining the behavior type represented by the access request initiated by the user device based on the number of access requests in each topic cluster. The behavior type is used to characterize whether the user device's access behavior is the behavior of a program or script that automatically crawls information from the World Wide Web. In other words, this application tracks and analyzes the content of access requests initiated by user devices, identifies crawler requests by analyzing whether the resource content of the access request is topic-related, achieves accurate identification of crawler requests, and solves the problem of at least missed identification in related technologies for detecting crawler behavior. Furthermore, it blocks and restricts identified crawler requests, ensuring the security of data in the website system, reducing attacks on the website system's servers by crawler requests, and reducing network bandwidth consumption. Attached Figure Description

[0021] Figure 1 A flowchart illustrating the request identification method provided for embodiments of this application. Figure 1 ;

[0022] Figure 2 A flowchart illustrating the request identification method provided for embodiments of this application. Figure 2 ;

[0023] Figure 3 A schematic diagram of the structure of a request identification device provided for an embodiment of this application;

[0024] Figure 4 This is a schematic diagram of the structure of an electronic device provided as an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other.

[0026] Embodiments of this application provide a request identification method applied to an electronic device, with reference to... Figure 1 As shown, the method includes the following steps:

[0027] Step 101: Obtain at least one access request initiated by the user device and determine the subject of each access request.

[0028] In this embodiment of the application, the electronic device acquires all access requests initiated by the user device within a target duration. All access requests can be one or multiple. That is, if the user device initiates only one access request within the target duration, then step 101 determines the topic of that single access request, and the number of determined topics is also one; if the user device initiates multiple access requests within the target duration, then step 101 determines the topic of each of the multiple access requests, and the number of determined access request topics is also multiple; that is, the number of access request topics corresponds one-to-one with the number of access requests.

[0029] In this embodiment of the application, the subject of the access request may be determined by the user equipment or sent to the user equipment by other devices associated with the user equipment.

[0030] In this application embodiment, the user equipment (UE) includes mobile terminal devices such as mobile phones, tablets, laptops, personal digital assistants (PDAs), cameras, and wearable devices, as well as fixed terminal devices such as desktop computers.

[0031] In this embodiment of the application, the access request includes a request initiated by the user equipment to access the server of the website system.

[0032] In this embodiment, the access request may be generated by the user device after the user clicks on a corresponding component in the access interface presented by the user device's client. Alternatively, the access request may be generated by the electronic device after the user opens a target link address and clicks on a corresponding component in the presented search interface. Of course, the electronic device may also generate the above-mentioned access request in other ways, and this application does not specifically limit this.

[0033] In this embodiment, the subject of an access request can be understood as the topic / type of the data accessed in the access request, or the access topic of the access request; for example, if all the accessed data in access request A is related to topic B, then the subject of access request A is topic B. Or, if access request B is used to access topic C, then the subject of access request B is topic C. The subject of an access request can be obtained by performing semantic analysis on the accessed data.

[0034] In this embodiment, each topic of the access request is assigned a unique topic identifier / tag. This step can determine the topic identifier / tag for each access request initiated by the user equipment. The identifier includes color identifiers, graphic identifiers, text identifiers, numerical identifiers, location identifiers, etc. Topics of the same type use the same type of identifier / tag.

[0035] Step 102: Cluster at least one access request based on the topic to obtain at least one topic cluster.

[0036] In this embodiment of the application, all access requests initiated by the user equipment are clustered according to the topic of each access request to obtain at least one topic cluster; wherein, the topics of the access requests in each topic cluster are the same or similar.

[0037] In this embodiment of the application, the number of topic clusters can be one or more. That is, if the number of access requests is one, after clustering, one topic cluster is obtained; if the number of access requests is multiple, after clustering, one topic cluster or multiple topic clusters can be obtained.

[0038] In this embodiment, the clustering process can be implemented using different algorithms. According to the clustering algorithm, all access requests initiated by the user equipment can be divided into at least one cluster. Access requests within the same cluster share the same attributes or characteristics, i.e., the topics of the access requests are the same or similar. Here, the clustering function in this application can be implemented by a software system, a hardware device, or a combination of both.

[0039] In this application embodiment, the clustering algorithms include, but are not limited to, k-means clustering (K-means), density-based spatial clustering of applications with noise (DBSCAN), agglomerative hierarchical clustering (AHC), statistical information grid (STING), and expectation-maximization (EM) algorithms.

[0040] Step 103: Based on the number of access requests in each topic cluster, determine the behavior type represented by the access request initiated by the user device.

[0041] Among them, the behavior type is used to characterize whether the user device's access behavior is the behavior of a program or script that automatically crawls information from the World Wide Web; the behavior type includes the type corresponding to normal user access behavior and the type corresponding to abnormal crawler access behavior.

[0042] In this embodiment of the application, if the behavior type is the type corresponding to normal user access behavior, that is, the user device's access behavior is not the behavior of a program or script that automatically crawls information from the World Wide Web. If the behavior type is the type corresponding to abnormal crawler access behavior, that is, the user device's access behavior is the behavior of a program or script that automatically crawls information from the World Wide Web, i.e., a web crawler.

[0043] In this embodiment of the application, determining the behavior type represented by the access request initiated by the user equipment based on the number of access requests in each topic cluster can be achieved through the following steps:

[0044] First, calculate the number of access requests included in each topic cluster within each topic cluster;

[0045] Then, a target topic cluster is selected from each topic cluster based on the number of access requests in each topic cluster. The target topic cluster must have a number of access requests greater than a threshold.

[0046] Next, calculate the total number of access requests across all target subject clusters.

[0047] Finally, the ratio of the total number of access requests in the target topic cluster to the total number of all access requests initiated by the user device is compared; and the behavior type represented by the access requests initiated by the user device is determined based on the ratio.

[0048] Alternatively, based on the number of access requests in each topic cluster, the behavior type represented by the access request initiated by the user device can be determined through the following steps:

[0049] By directly determining the relationship between the number of access requests in each topic cluster and the number threshold, the behavior type represented by the access request initiated by the user device can be determined.

[0050] Here, comparing the ratio of the total number of access requests in the target topic cluster to the total number of access requests initiated by the user device is used to determine whether the access requests initiated by the user device are highly concentrated on certain topics. Determining the relationship between the number of access requests in each topic cluster and a threshold number is used to determine whether the access requests initiated by the user device are all on discrete topics.

[0051] It's important to note that when users browse and search for web resources normally, their behavior tends to concentrate on certain themes within a given timeframe. Malicious web crawlers, on the other hand, typically perform indiscriminate crawling of target websites, resulting in scattered and disjointed access requests. Therefore, we can identify web crawler requests by analyzing whether the content of the requested resources is thematically relevant.

[0052] Embodiments of this application provide a method for identifying requests. The method includes: acquiring at least one access request initiated by a user device and determining the topic of each access request; clustering the at least one access request based on the topic to obtain at least one topic cluster; and determining the behavior type represented by the access request initiated by the user device based on the number of access requests in each topic cluster. The behavior type is used to characterize whether the user device's access behavior is the behavior of a program or script that automatically crawls information from the World Wide Web. In other words, this application tracks and analyzes the content of access requests initiated by user devices, identifies crawler requests by analyzing whether the resource content of the access request is topic-related, achieves accurate identification of crawler requests, and solves the problem of at least missed identification in related technologies for detecting crawler behavior. Furthermore, it blocks and restricts identified crawler requests, ensuring the security of data in the website system, reducing attacks on the website system's servers by crawler requests, and reducing network bandwidth consumption.

[0053] Furthermore, the request identification method provided in this application can be compared with the crawler behavior-based detection method in related technologies, and can be judged from both the request content and crawler behavior, thereby effectively improving the accuracy of crawler request identification.

[0054] In this embodiment of the application, determining the topic of each access request in step 101 can be achieved through the following steps:

[0055] Step A1: Filter the content summary text of each access request to obtain a list of sentences including multiple keywords.

[0056] In this application, filtering the content summary text of the access request includes any one or more of the following: word segmentation, error correction (such as correcting erroneous words in the content summary text), noise reduction (such as removing meaningless letters, symbols, etc.), and (for example, removing stop words and detecting the part of speech of each word). Stop words may include function words with low content-indicating meaning, which are generally difficult to indicate the semantics of the text, such as words like "a," "these," and "of," which are difficult to indicate the semantics of the text.

[0057] In some embodiments, the content summary text of each access request is filtered to obtain a list of sentences including multiple keywords. This can be understood as the electronic device filtering out words or phrases that are not keywords, extracting multiple keywords, and obtaining a list of sentences including multiple keywords. Here, keywords can be words selected from the name, summary, and body of the content summary text that are of substantial significance in expressing the central content of the content summary text.

[0058] For example, electronic devices can extract keywords from documents using the term frequency–inverse document frequency (TF-IDF) algorithm based on Natural Language Processing (NLP) technology.

[0059] Step A2: Determine the similarity between any two sentences in the sentence list.

[0060] In this embodiment of the application, the electronic device, upon obtaining a list of sentences, then obtains multiple vectors corresponding to the list of sentences; and then calculates the similarity between any two vectors.

[0061] Step A3: Calculate the weight coefficient of each sentence in the sentence list based on similarity.

[0062] In this embodiment of the application, the weight coefficient of each sentence in the sentence list can be calculated based on similarity in the following way: based on the similarity between any two vectors, the weight coefficient corresponding to the vector of each sentence is iteratively adjusted.

[0063] Step A4: Select the topic corresponding to the sentence whose weight coefficient meets the coefficient filtering condition as the topic of the access request.

[0064] In some embodiments, in order to provide a more concise summary of the content of the access request initiated by the user equipment, this application may employ relevant extraction algorithms to extract the topic sentences of the summary information of the request content.

[0065] For example, taking the TextRank algorithm as an example, the TextRank algorithm is used to extract topic sentences from the summary information of the requested content, including the following steps:

[0066] First, the content summary text of the access request initiated by the user device is processed by sentence segmentation and word segmentation. Then, word filtering is performed through stop word filtering and part-of-speech filtering to obtain a list of sentences composed of words.

[0067] Next, the similarity between sentences in the sentence list is calculated. Here, the electronic device can calculate the similarity of the preprocessed sentences based on the amount of overlapping information between them, using the following formula:

[0068]

[0069] Among them, Similarity(S) i ,S j ) is used to characterize the calculation sentence S i With sentence S j Similarity between them; S i and S j Used to represent two sentences, sentence S i Including N i Term Sentence S j Including M i Term w k Used to characterize sentence S i And sentence S j One term in the text.

[0070] Next, sentence weights are calculated. A graph model is built using sentences in the content summary text as nodes and the similarity between sentences as edges. The weights of each node are iteratively calculated using the TextRank algorithm until convergence. For any sentence node V... i Weight WS(V) i The calculation formula (2) is as follows:

[0071]

[0072] Where d represents the damping coefficient (0 ≤ d ≤ 1), it is typically taken as 0.85 to ensure that the weight value of each node is greater than 0. In(V i ) indicates pointing to node V i The set of all nodes, Out(V) j ) indicates pointing to node V j The set of all nodes of V. w(ij) represents node V. i and node V j The weights of the edges between them. It should be noted that w(ij) = Similarity(S) i ,S j ).

[0073] Finally, extract the topic sentence.

[0074] For example, all sentences after weight convergence are sorted, and the sentence with the highest score is selected as the topic sentence. Of course, sentences with scores within a preset score range can also be selected as topic sentences.

[0075] In this embodiment of the application, the step 102 of clustering at least one access request based on a topic to obtain at least one topic cluster can be achieved through the following steps:

[0076] Step B1: Based on the reception time of each access request, determine each sliding window corresponding to each access request.

[0077] In this embodiment of the application, in order to prevent crawler requests with low speed and random time intervals, the application analyzes the access requests initiated by the user equipment in the form of a sliding window; that is, after receiving the access request initiated by the user equipment, the sliding window of each access request is determined based on the reception time of each access request.

[0078] The sliding window in this application can be understood as a container used to cache access data for access requests for a certain period of time.

[0079] Step B2: Based on the topic of the access request in each sliding window, cluster the access requests in all sliding windows to obtain at least one topic cluster.

[0080] This application uses access requests within each sliding window as the unit of analysis. First, based on the topic of each access request in each sliding window, the topics are clustered to obtain at least one topic cluster corresponding to each sliding window. Then, based on the number of access requests in the at least one topic cluster corresponding to each sliding window, the behavior type represented by the access requests in the sliding window is determined. Furthermore, by summarizing and analyzing the behavior types represented by access requests in all sliding windows of the website system, the behavior type represented by access requests initiated by user devices is obtained. This improves the accuracy of identifying the behavior type represented by the request.

[0081] In this embodiment, if the number of at least one access request is one, then the number of sliding windows corresponding to that access request is also one. Therefore, clustering one access request within that sliding window yields one topic cluster. If the number of at least one access request is multiple, then those multiple access requests can be accommodated in multiple sliding windows or in a single sliding window; that is, the number of access requests in each sliding window can be one or multiple. Furthermore, if the number of access requests in a sliding window is one, then clustering those access requests yields one topic cluster; if the number of access requests in a sliding window is multiple, then clustering those access requests yields one or multiple topic clusters.

[0082] Furthermore, determining each sliding window corresponding to each access request based on the reception time of each access request in step B1 can be achieved through the following steps:

[0083] Step B11: Obtain the reception time of the Nth access request and the length of the Mth sliding window.

[0084] Wherein, the Mth sliding window is the sliding window corresponding to the (N-1)th access request; N is a positive integer greater than 2 and less than or equal to the number of all access requests sent by the user equipment; M is a positive integer greater than or equal to 1 and less than or equal to N.

[0085] Step B12: Based on the reception time of the Nth access request and the length of the Mth sliding window, determine the target sliding window corresponding to the Nth access request.

[0086] In this embodiment of the application, if the reception time of the Nth access request is within a preset duration and the length of the Mth sliding window is greater than or equal to the preset length, the target sliding window is determined to be the Mth sliding window; if the reception time of the Nth access request is not within the preset duration, or the length of the Mth sliding window is less than the preset length, the target sliding window is determined to be the (M+1)th sliding window.

[0087] Here, the preset length refers to the sum of the length of the access requests already received in the Mth sliding window and the length of the Nth access request. The preset duration, or the time limit for the sliding window, is the maximum time range that the website system pre-sets for a sliding window. That is, the sliding window can only accommodate access requests received within this time range. Access requests exceeding this time range will be directed to the next sliding window.

[0088] In this embodiment, after receiving an access request initiated by the user equipment, if it is determined that the access request is not the first or second access request initiated by the user equipment, it is first determined whether the time of the access request is within the limited time of the sliding window corresponding to the previous access request, i.e., within a preset duration, and whether the remaining length of the sliding window corresponding to the previous access request is greater than or equal to the length of the access request. If the access request is within the limited time of the sliding window corresponding to the previous access request, and the remaining length of the sliding window corresponding to the previous access request is greater than or equal to the length of the access request, the sliding window of the access request is determined to be the sliding window corresponding to the previous access request. If the access request is not within the limited time of the sliding window corresponding to the previous access request, or the remaining length of the sliding window corresponding to the previous access request is less than the length of the access request, then the sliding window of the access request is used as the first request of the next sliding window of the sliding window corresponding to the previous access request. Here, the length of the sliding window corresponding to the previous access request includes the remaining length of the sliding window corresponding to the previous access request and the length of the existing access requests in the sliding window.

[0089] In some embodiments, the sliding window may include A access requests; here, A is a positive integer, greater than 0, and less than or equal to the number of access requests that the default length of the sliding window can accommodate. A access requests may be the number of access requests that can be obtained within a preset time period.

[0090] In some embodiments, if the reception time of the access request indicates that the access request is the first access request sent by the user equipment, the sliding window of the first access request is determined to be the first sliding window; if the reception time of the access request indicates that the access request is the second access request sent by the user equipment, it is determined whether the reception time of the second access request is within a preset duration; if the reception time of the second access request is within the preset duration, the sliding window of the second access request is determined to be the first sliding window; if the reception time of the second access request is not within the preset duration, the sliding window of the second access request is determined to be the second sliding window.

[0091] In some embodiments, this application sets the default length of the sliding window to l, the actual length to l', the step size to p (1≤p<l), the marker variable to stamp, the time limit to t, and the total number of requests within time t to n. The sliding window construction algorithm steps are as follows:

[0092] Step C1: Obtain a request initiated by the user device and mark it as a stamp, which serves as the start of sliding window A and the first request of window A.

[0093] Step C2: Continue to receive subsequent requests initiated by the user equipment. If the time of the subsequent request is within the time limit t of the previous stamp, put the request into the sliding window A and execute step C3; otherwise, execute step C4.

[0094] Step C3: Determine whether the length of the sliding window A is greater than the default length l. If it is greater, proceed to step C5; otherwise, proceed to step C2.

[0095] Step C4: Determine if the number of requests in the sliding window A is greater than p. If it is, proceed to step C5; otherwise, proceed to step C6.

[0096] Step C5: Mark stamp = stamp + p, then execute step C1.

[0097] Step C6: Mark stamp = n, then execute step C1.

[0098] It should be noted that the above loop process will stop when the user device stops initiating requests.

[0099] In this embodiment of the application, step B2, which involves clustering the access requests in all sliding windows based on the topics of the access requests in each sliding window to obtain at least one topic cluster, can be achieved through the following steps:

[0100] Step B21: Determine the subject of the access request located at the cluster center.

[0101] Step B22: Calculate the similarity between the topic of the access request located at the cluster center and the topics of each access request in the sliding window.

[0102] Step B23: Assign access requests to topics with similarity greater than or equal to the similarity threshold to the same topic cluster.

[0103] At least one topic cluster includes the same topic cluster.

[0104] In this embodiment of the application, the electronic device performs topic sentence clustering on the requests within the window based on the content topic sentence of the access request initiated by the user device and the sliding window, thereby obtaining at least one topic cluster.

[0105] For example, this application uses the affinity propagation (AP) algorithm based on the information transfer mechanism for clustering. The clustering steps include the following:

[0106] Step D1: Calculate the pairwise similarity of the topic sentences of each access request in the sliding window to obtain a similarity matrix.

[0107] For example, a sliding window A contains n access requests. The similarity of the topic sentences of the n access requests is calculated pairwise to obtain an n×n similarity matrix S. n×n :

[0108]

[0109] Among them, S 11 S 1n S n1 and S nn This indicates the similarity between the topic sentences of two access requests.

[0110] Step D2: Construct the initial attraction matrix r and belonging matrix a.

[0111] The attraction matrix r and the belonging matrix a are initialized to 0.

[0112] Step D3: Iteratively update the attraction matrix r and the belonging matrix a according to formulas (4) and (5).

[0113]

[0114]

[0115] Where r(i,k) represents the suitability of the topic sentence k of the access request as the cluster center of the topic sentence i of the access request; a(i,k) represents the suitability of the topic sentence i of the access request choosing the topic sentence k of the access request as its cluster center; and S(i,k) represents the similarity between the topic sentence k of the access request and the topic sentence i of the access request.

[0116] Step D3: If the number of iterations reaches the preset value, or the cluster centers do not change with the iteration calculation, then stop the iteration and execute step D4; otherwise, repeat step D2.

[0117] Step D4: Add r(i,k) and a(i,k) together, and select the point with the largest value in each row as the cluster center.

[0118] Step D5: Classify the topic sentences based on the cluster centers. Calculate the distance from the topic sentence of each access request to the cluster center. If the topic sentence of the access request is less than or equal to the preset maximum distance, then classify the topic sentence of the access request into the cluster containing the cluster center with the smallest distance. If the distance from the topic sentence of the access request to the cluster center is greater than the preset maximum distance, then the topic sentence of the access request is discrete and forms a separate cluster.

[0119] In this embodiment of the application, the distance from the topic sentence of each access request to the cluster center includes, but is not limited to, Euclidean distance, Mahalanobis distance, and Hamming distance.

[0120] In this embodiment of the application, if the number of topic clusters after clustering is small and the number of requests in the largest cluster accounts for a high proportion of the total number of requests in the window, then the behavior type represented by the access request initiated by the user device is the type corresponding to the normal access behavior of the user; otherwise, it indicates that the topic of the request content in the window is relatively discrete, and the behavior type represented by the access request initiated by the user device is the type corresponding to the abnormal access behavior of the crawler.

[0121] In some embodiments, determining the type of behavior represented by the access request initiated by the user equipment can be achieved through steps E1 to E3, or through steps E1 to E2 and E4:

[0122] Step E1: Select target topic clusters from at least one topic cluster whose number is greater than a threshold.

[0123] Step E2: Calculate the ratio of the first number of access requests in the target topic cluster to the second number of access requests in the sliding window.

[0124] Step E3: If the ratio is greater than or equal to the preset ratio, determine that the behavior type belongs to normal user access behavior.

[0125] Step E4: If the ratio is less than the preset ratio, determine that the behavior type belongs to abnormal crawler access behavior.

[0126] In this embodiment, the electronic device can sort topic clusters in descending order based on the number of access requests within each cluster, and then select the top N topic clusters. The number of access requests for the top N topic clusters must exceed a threshold. Alternatively, the device can directly select target topic clusters whose number of access requests exceeds a threshold based solely on the number of access requests within each cluster, without needing to sort them.

[0127] Furthermore, the percentage of requests from the top N topic clusters within the entire sliding window is calculated. Where s(Top1) represents the number of access requests included in the topic cluster with the largest number in descending order; s(Top1) represents the number of access requests included in the topic cluster with the second largest number in descending order; s(Top2) represents the number of access requests included in the topic cluster with the Nth largest number in descending order. S represents the total number of requests included in the sliding window. When the topic of the user's request within the sliding window is relevant, it indicates normal user behavior; when If the condition is met, the request is considered a web crawler request. Here, the P-value and the number of requests for the selected Top N topic clusters can be obtained through experiments based on the actual content of the website system, and used to adjust the strictness of the web crawler request behavior judgment.

[0128] Figure 2 This is a schematic diagram of a request identification process provided in an embodiment of this application.

[0129] Step 201: Extract the topic sentence of the access request from the content summary text in the access request initiated by the user device.

[0130] Step 202: Based on user identification information, construct a sliding window for access requests corresponding to each user device.

[0131] Step 203: Based on topic relevance, perform clustering calculations on the topic sentences of the requested content within the sliding window to obtain topic clusters.

[0132] Step 204: Analyze the proportion of the number of requests in each topic cluster within the entire sliding window, and determine the crawler request behavior based on the proportion.

[0133] Embodiments of this application provide a request identification device, which can be applied to... Figure 1 In a corresponding embodiment of the request identification method, referring to Figure 3 As shown, the identification device 3 for the request includes:

[0134] The acquisition module 302 is used to acquire at least one access request initiated by the user equipment;

[0135] Processing module 301 is used to determine the subject of each access request;

[0136] Processing module 301 is used to cluster at least one access request based on a topic to obtain at least one topic cluster;

[0137] The processing module 301 is used to determine the behavior type represented by the access request initiated by the user equipment based on the number of access requests in each topic cluster; wherein, the behavior type is used to characterize whether the access behavior of the user equipment is the behavior of a program or script that automatically crawls World Wide Web information.

[0138] In other embodiments of this application, the processing module 301 is used to determine the sliding window corresponding to each access request based on the reception time of each access request;

[0139] The processing module 301 is used to cluster the access requests in all sliding windows based on the topic of the access request in each sliding window to obtain at least one topic cluster.

[0140] In other embodiments of this application, the processing module 301 is used to obtain the reception time of the Nth access request and the length of the Mth sliding window; wherein, the Mth sliding window is the sliding window corresponding to the (N-1)th access request; N is a positive integer greater than 2 and less than or equal to the number of all access requests sent by the user equipment; M is a positive integer greater than or equal to 1 and less than or equal to N;

[0141] Processing module 301 is used to determine the target sliding window corresponding to the Nth access request based on the reception time of the Nth access request and the length of the Mth sliding window.

[0142] In other embodiments of this application, the processing module 301 is used to determine the target sliding window as the Mth sliding window if the reception time of the Nth access request is within a preset duration and the length of the Mth sliding window is greater than or equal to the preset length.

[0143] The processing module 301 is used to determine the target sliding window as the (M+1)th sliding window if the reception time of the Nth access request is not within the preset duration, or the length of the Mth sliding window is less than the preset length.

[0144] In other embodiments of this application, the processing module 301 is configured to determine the sliding window of the first access request as the first sliding window if the reception time of the access request indicates that the access request is the first access request sent by the user equipment.

[0145] The processing module 301 is used to determine whether the receiving time of the second access request is within a preset duration if the receiving time of the access request indicates that the access request is a second access request sent by the user equipment.

[0146] The processing module 301 is used to determine the sliding window of the second access request as the first sliding window if the reception time of the second access request is within a preset time.

[0147] The processing module 301 is used to determine the sliding window of the second access request as the second sliding window if the reception time of the second access request is not within the preset time.

[0148] In other embodiments of this application, the processing module 301 is used to filter the content summary text of each access request to obtain a list of sentences including multiple keywords;

[0149] Processing module 301 is used to determine the similarity between any two sentences in the sentence list;

[0150] Processing module 301 is used to calculate the weight coefficient of each sentence in the sentence list based on similarity;

[0151] The processing module 301 is used to select the topic corresponding to the sentence whose weight coefficient meets the coefficient filtering condition as the topic of the access request.

[0152] In other embodiments of this application, processing module 301 is used to determine the topic of the access request located at the cluster center;

[0153] Processing module 301 is used to calculate the similarity between the topic of the access request located at the cluster center and the topics of each access request in the sliding window;

[0154] The processing module 301 is used to group access requests corresponding to topics with a similarity greater than or equal to a similarity threshold into the same topic cluster; wherein at least one topic cluster includes the same topic cluster.

[0155] In other embodiments of this application, the processing module 301 is used to filter out target topic clusters from at least one topic cluster whose number is greater than a number threshold;

[0156] Processing module 301 calculates the ratio of the first number of access requests in the target topic cluster to the second number of access requests in the sliding window;

[0157] Processing module 301 determines that the behavior type belongs to normal user access behavior if the ratio is greater than or equal to the preset ratio.

[0158] Processing module 301 determines that the behavior type belongs to abnormal crawler access behavior if the ratio is less than the preset ratio.

[0159] It should be noted that the specific implementation process of the steps executed by the processing module 301 in this embodiment can be referred to Figure 1 The implementation process of the request identification method provided in the corresponding embodiment will not be described in detail here.

[0160] Embodiments of this application provide an electronic device that can be applied to... Figure 1 In a corresponding embodiment of the request identification method, referring to Figure 4 As shown, the electronic device 4 ( Figure 4 Electronic devices 4 and Figure 3 The identification device 3 corresponding to the request includes: a processor 401, a memory 402, and a communication bus 403, wherein:

[0161] The communication bus 403 is used to realize the communication connection between the processor 401 and the memory 402.

[0162] The processor 401 is used to execute the request identification program stored in the memory 402 to perform the following steps:

[0163] Obtain at least one access request initiated by the user device and determine the subject of each access request;

[0164] Cluster at least one access request based on a topic to obtain at least one topic cluster;

[0165] Based on the number of access requests in each topic cluster, the behavior type represented by the access request initiated by the user device is determined; whereby the behavior type is used to characterize whether the access behavior of the user device is the behavior of a program or script that automatically crawls information from the World Wide Web.

[0166] In other embodiments of this application, processor 401 is used to execute a request identification program stored in memory 402 to perform the following steps:

[0167] Based on the reception time of each access request, determine the corresponding sliding window for each access request;

[0168] Based on the topic of the access request in each sliding window, the access requests in all sliding windows are clustered to obtain at least one topic cluster.

[0169] In other embodiments of this application, processor 401 is used to execute a request identification program stored in memory 402 to perform the following steps:

[0170] Obtain the reception time of the Nth access request and the length of the Mth sliding window; where the Mth sliding window is the sliding window corresponding to the (N-1)th access request; N is a positive integer greater than 2 and less than or equal to the number of all access requests sent by the user equipment; M is a positive integer greater than or equal to 1 and less than or equal to N;

[0171] Based on the reception time of the Nth access request and the length of the Mth sliding window, determine the target sliding window corresponding to the Nth access request.

[0172] In other embodiments of this application, processor 401 is used to execute a request identification program stored in memory 402 to perform the following steps:

[0173] If the reception time of the Nth access request is within the preset duration, and the length of the Mth sliding window is greater than or equal to the preset length, the target sliding window is determined to be the Mth sliding window;

[0174] If the reception time of the Nth access request is not within the preset duration, or the length of the Mth sliding window is less than the preset length, the target sliding window is determined to be the (M+1)th sliding window.

[0175] In other embodiments of this application, processor 401 is used to execute a request identification program stored in memory 402 to perform the following steps:

[0176] If the reception time of the access request indicates that the access request is the first access request sent by the user equipment, then the sliding window for the first access request is determined as the first sliding window;

[0177] If the reception time of the access request indicates that the access request is the second access request sent by the user equipment, determine whether the reception time of the second access request is within the preset duration;

[0178] If the receiving time of the second access request is within the preset duration, the sliding window of the second access request is determined to be the first sliding window;

[0179] If the receiving time of the second access request is not within the preset duration, the sliding window for the second access request is determined to be the second sliding window.

[0180] In other embodiments of this application, processor 401 is used to execute a request identification program stored in memory 402 to perform the following steps:

[0181] The content summary text of each access request is filtered to obtain a list of sentences including multiple keywords;

[0182] Determine the similarity between any two sentences in the sentence list;

[0183] Based on similarity, calculate the weight coefficient of each sentence in the sentence list;

[0184] The topic corresponding to the sentence whose weight coefficient meets the coefficient filtering criteria is selected as the topic of the access request.

[0185] In other embodiments of this application, processor 401 is used to execute a request identification program stored in memory 402 to perform the following steps:

[0186] Determine the subject of the access request located at the cluster center;

[0187] Calculate the similarity between the topic of the access request located at the cluster center and the topics of each access request in the sliding window;

[0188] Access requests for topics with a similarity greater than or equal to a similarity threshold are grouped into the same topic cluster; wherein at least one topic cluster includes the same topic cluster.

[0189] In other embodiments of this application, processor 401 is used to execute a request identification program stored in memory 402 to perform the following steps:

[0190] From at least one topic cluster, select target topic clusters whose number exceeds a certain threshold;

[0191] Calculate the ratio of the first number of access requests in the target topic cluster to the second number of access requests in the sliding window;

[0192] If the ratio is greater than or equal to the preset ratio, the behavior type is determined to be normal user access behavior.

[0193] If the ratio is less than the preset ratio, the behavior type is determined to be abnormal crawler access behavior.

[0194] The method provided in this application embodiment can be directly embodied as a combination of software modules executed by processor 401. The software modules can be located in a storage medium, which is located in memory 402. Processor 401 reads the executable instructions included in the software modules in memory 402 and combines them with necessary hardware to complete the method provided in this application embodiment.

[0195] As an example, processor 401 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0196] It should be noted that the specific implementation process of the steps executed by the processor in this embodiment can be referred to Figure 1 The implementation process of the request identification method provided in the corresponding embodiment will not be described in detail here.

[0197] Embodiments of this application provide a computer-readable storage medium storing one or more programs that can be executed by one or more processors to perform, as follows: Figure 1 The implementation process of the request identification method provided in the corresponding embodiment will not be described in detail here.

[0198] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0199] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0200] It should be understood that the terms "an embodiment," "an embodiment," "an embodiment of this application," "the foregoing embodiment," "some embodiments," or "some implementations" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, the phrases "an embodiment," "an embodiment," "an embodiment of this application," "the foregoing embodiment," "some embodiments," or "some implementations" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0201] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0202] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0203] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0204] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0205] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0206] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0207] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0208] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0209] It is worth noting that the accompanying drawings in this application are only for illustrating the schematic positions of various devices on the terminal device and do not represent their actual positions in the terminal device. The actual positions of each device or area may be changed or shifted according to the actual situation (e.g., the structure of the terminal device). Furthermore, the proportions of different parts in the terminal device in the drawings do not represent the actual proportions.

[0210] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for identifying a request, characterized in that, The method includes: At least one access request initiated by a user device is obtained, and the topic of each access request is determined; the topic of the access request is obtained by semantic analysis of the access data in the access request. Based on the topic, the at least one access request is clustered to obtain at least one topic cluster; Based on the number of access requests in each of the aforementioned topic clusters, the behavior type represented by the access request initiated by the user equipment is determined; wherein, the behavior type is used to characterize whether the access behavior of the user equipment is the behavior of a program or script that automatically crawls World Wide Web information; The clustering process based on the topic to obtain at least one topic cluster includes: Based on user identification information, a sliding window for access requests corresponding to each user device is constructed; based on the reception time of each access request, each sliding window corresponding to each access request is determined; based on the topic of the access request in each sliding window, the access requests in all sliding windows are clustered to obtain at least one topic cluster. The step of determining each sliding window corresponding to each access request based on the reception time of each access request includes: The receiving time of the Nth access request and the length of the Mth sliding window are obtained; wherein, the Mth sliding window is the sliding window corresponding to the (N-1)th access request; N is a positive integer greater than 2 and less than or equal to the number of all access requests sent by the user equipment; M is a positive integer greater than or equal to 1 and less than or equal to N; based on the receiving time of the Nth access request and the length of the Mth sliding window, the target sliding window corresponding to the Nth access request is determined; The determination of the behavior type represented by the access request initiated by the user equipment based on the number of access requests in each of the topic clusters includes: From the at least one topic cluster, select the target topic clusters whose number is greater than the number threshold; Calculate the ratio of a first number of access requests in the target topic cluster to a second number of access requests in the sliding window; If the ratio is greater than or equal to a preset ratio, the behavior type is determined to be normal user access behavior; If the ratio is less than the preset ratio, the behavior type is determined to be abnormal crawler access behavior.

2. The method according to claim 1, characterized in that, The step of determining the target sliding window corresponding to the Nth access request based on the reception time of the Nth access request and the length of the Mth sliding window includes: If the reception time of the Nth access request is within a preset duration, and the length of the Mth sliding window is greater than or equal to the preset length, then the target sliding window is determined to be the Mth sliding window. If the reception time of the Nth access request is not within the preset duration, or the length of the Mth sliding window is less than the preset length, the target sliding window is determined to be the (M+1)th sliding window.

3. The method according to claim 1, characterized in that, The step of determining each sliding window corresponding to each access request based on the reception time of each access request further includes: If the reception time of the access request indicates that the access request is the first access request sent by the user equipment, then the sliding window of the first access request is determined as the first sliding window; If the reception time of the access request indicates that the access request is a second access request sent by the user equipment, determine whether the reception time of the second access request is within a preset duration; If the receiving time of the second access request is within the preset duration, the sliding window of the second access request is determined to be the first sliding window; If the receiving time of the second access request is not within the preset duration, the sliding window of the second access request is determined to be the second sliding window.

4. The method according to claim 1, characterized in that, Determining the topic of each access request includes: The content summary text of each access request is filtered to obtain a list of sentences including multiple keywords; Determine the similarity between any two sentences in the sentence list; Based on the similarity, the weight coefficient of each sentence in the sentence list is calculated; The topic corresponding to the sentence whose weight coefficient meets the coefficient filtering condition is selected as the topic of the access request.

5. The method according to claim 1, characterized in that, The process of clustering access requests across all sliding windows based on the topics of access requests in each sliding window to obtain at least one topic cluster includes: Determine the subject of the access request located at the cluster center; Calculate the similarity between the topic of the access request located at the cluster center and the topics of each access request in the sliding window; Access requests corresponding to topics with a similarity greater than or equal to a similarity threshold are grouped into the same topic cluster; wherein, the at least one topic cluster includes the same topic cluster.

6. A request identification device, characterized in that, The device includes: The acquisition module is used to acquire at least one access request initiated by the user device; The processing module is used to determine the topic of each access request; the topic of the access request is obtained by semantic analysis of the access data in the access request; The processing module is further configured to perform clustering processing on the at least one access request based on the topic to obtain at least one topic cluster; The processing module is further configured to determine the behavior type represented by the access request initiated by the user equipment based on the number of access requests in each of the topic clusters; wherein, the behavior type is used to characterize whether the access behavior of the user equipment is the behavior of a program or script that automatically crawls World Wide Web information; The processing module is further configured to: construct a sliding window for access requests corresponding to each user device based on user identification information; determine each sliding window corresponding to each access request based on the reception time of each access request; and perform clustering processing on the access requests in all sliding windows based on the topic of the access requests in each sliding window to obtain the at least one topic cluster. The processing module is further configured to obtain the reception time of the Nth access request and the length of the Mth sliding window; wherein the Mth sliding window is the sliding window corresponding to the (N-1)th access request; N is a positive integer greater than 2 and less than or equal to the number of all access requests sent by the user equipment; M is a positive integer greater than or equal to 1 and less than or equal to N; and based on the reception time of the Nth access request and the length of the Mth sliding window, determine the target sliding window corresponding to the Nth access request; The processing module is further configured to: filter out target topic clusters from the at least one topic cluster whose number is greater than a number threshold; calculate the ratio of a first number of access requests in the target topic cluster to a second number of access requests in the sliding window; if the ratio is greater than or equal to a preset ratio, determine that the behavior type belongs to normal user access behavior; if the ratio is less than the preset ratio, determine that the behavior type belongs to abnormal crawler access behavior.

7. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the identification method of any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs that can be executed by one or more processors to implement the method for identifying the request as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text theme mining method based on intra-sentence association graph

    CN104298709A

  • Abnormal access behavior detection method and device and electronic equipment

    CN113535823A