Access data processing method and device, equipment and storage medium

Judging crawler users through behavioral analysis models solves the problem that traditional methods are difficult to identify crawler users, realizes filtering of crawler access and protecting system resources, and improves data processing efficiency and security.

CN120354425APending Publication Date: 2025-07-22BEIJING XIAKEHUI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510399583.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional methods are difficult to accurately identify and intercept the abnormal behavior of crawler users, resulting in security risks and malicious use of resources, affecting the access experience of normal users.

Method used

By obtaining the pending access data, input it into the behavioral analysis model of the logistic regression model trained based on multiple historical access data features and label information, it is determined whether it is a crawler user operation, and refuses to process the data when it is determined to be a crawler user.

Benefits of technology

Effectively filter crawler access, reduce system burden, improve data processing efficiency, protect system resources, and ensure the security and stability of normal user access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354425A_ABST
    Figure CN120354425A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an access data processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring to-be-processed access data; and inputting the to-be-processed access data into a behavior analysis model to obtain target label information corresponding to the to-be-processed access data, the behavior analysis model being obtained by training a logistic regression model based on feature data in the multiple pieces of historical access data and label information corresponding to the multiple pieces of historical access data, for feature data and corresponding label information in any historical access data, the feature data is features related to the crawler user, and the label information is whether the crawler user is available or not; and if the target label information indicates that the to-be-processed access data is the crawler user operation, determining that the to-be-processed access data is not processed. The method is used for avoiding potential safety hazards caused by malicious access, malicious occupation of resources, influence on access experience of normal users and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, and in particular, to a method, apparatus, device, and storage medium for processing access data. Background Art

[0002] With the rapid development of Internet technologies, the number of users and the volume of access to websites and online services have increased explosively. This trend has brought huge opportunities to Internet companies, but at the same time, it has also brought unprecedented security threats and management challenges.

[0003] As the number of users increases, the complexity and diversity of user behaviors have also increased, and traditional analysis methods are no longer sufficient to meet the requirements. Especially for high-traffic users, such as crawler users, their behavior characteristics are often significantly different from those of normal users, but traditional defense means are often difficult to accurately identify and intercept these abnormal behaviors.

[0004] Therefore, how to solve the security risks brought by malicious access, the malicious occupation of resources, and the impact on the access experience of normal users has become a technical problem to be solved urgently. Summary of the Invention

[0005] Embodiments of this application provide a method, apparatus, device, and storage medium for processing access data, which can avoid security risks brought by malicious access, the malicious occupation of resources, and the impact on the access experience of normal users.

[0006] In a first aspect, embodiments of this application provide a method for processing access data, including:

[0007] Obtain access data to be processed;

[0008] Input the access data to be processed into a behavior analysis model to obtain target label information corresponding to the access data to be processed. The behavior analysis model is obtained by training a logistic regression model based on feature data in multiple historical access data and label information corresponding to the multiple historical access data. For the feature data and corresponding label information in any historical access data, the feature data is a feature related to a crawler user, and the label information is whether it is a crawler user;

[0009] If the target label information indicates that the access data to be processed is an operation by a crawler user, determine that the access data to be processed is not processed.

[0010] In a possible implementation manner, the step of inputting the access data to be processed into a behavior analysis model to obtain target label information corresponding to the access data to be processed includes:

[0011] Verify the to-be-processed access data according to the verification conditions to obtain the verification result corresponding to the to-be-processed access data. The verification conditions are determined based on at least one verification item, and the at least one verification item includes: the number of accesses to a preset Internet Protocol (IP) address within a preset time duration is less than the access count threshold, the user identity identifier (ID) is not in the preset user ID blacklist, and the user ID is in the preset user ID whitelist. The preset IP address is the user IP address and / or the accessed IP address;

[0012] If the verification result indicates that the verification conditions are passed, input the to-be-processed access data into the behavior analysis model to obtain the target tag information.

[0013] In a possible implementation manner, before verifying the to-be-processed access data according to the verification conditions to obtain the verification result corresponding to the to-be-processed access data, the method further includes:

[0014] Perform parsing processing on the to-be-processed access data to obtain at least one target metric item, where the at least one target metric item includes: the target user ID and the target IP address;

[0015] Put the at least one target metric item into the statistical data of the data pool, where the data pool is used to store data obtained by parsing the real-time access data acquired from the gateway and performing multi-dimensional statistics based on time periods and at least one metric item.

[0016] In a possible implementation manner, verifying the to-be-processed access data according to the verification conditions to obtain the verification result corresponding to the to-be-processed access data includes:

[0017] In response to a configuration request from the user, perform a configuration operation on at least one verification item in the verification conditions to obtain updated verification conditions;

[0018] Verify the to-be-processed access data according to the updated verification conditions to obtain the verification result corresponding to the to-be-processed access data.

[0019] In a possible implementation manner, inputting the to-be-processed access data into the behavior analysis model to obtain the target tag information corresponding to the to-be-processed access data includes:

[0020] Extract the features related to crawler users in the to-be-processed access data as target feature data;

[0021] Input the target feature data into the behavior analysis model to obtain the target tag information.

[0022] In a possible implementation manner, the method further includes:

[0023] If the target label information indicates that the access data to be processed is a non-spider user operation, obtain the resource usage information of the cluster corresponding to the access data to be processed;

[0024] Process the access data to be processed according to the resource usage information.

[0025] In a second aspect, an embodiment of the present application provides a processing device for access data, including:

[0026] An obtaining module, configured to obtain access data to be processed;

[0027] A processing module, configured to input the access data to be processed into a behavior analysis model to obtain target label information corresponding to the access data to be processed. The behavior analysis model is trained based on feature data in multiple historical access data and label information corresponding to the multiple historical access data. For the feature data and corresponding label information in any historical access data, the feature data is a feature related to a spider user, and the label information is whether it is a spider user;

[0028] A determining module, configured to determine that the access data to be processed is not processed if the target label information indicates that the access data to be processed is a spider user operation.

[0029] In a possible implementation manner, the processing module is specifically configured to:

[0030] Verify the access data to be processed according to verification conditions to obtain a verification result corresponding to the access data to be processed. The verification conditions are determined based on at least one verification item. The at least one verification item includes: the number of accesses to a preset Internet Protocol (IP) address within a preset time period is less than an access count threshold, the user identity identifier (ID) is not in a preset user ID blacklist, and the user ID is in a preset user ID whitelist. The preset IP address is the user IP address and / or the accessed IP address;

[0031] If the verification result indicates that the verification conditions are passed, input the access data to be processed into the behavior analysis model to obtain the target label information.

[0032] In a possible implementation manner, before verifying the access data to be processed according to the verification conditions to obtain a verification result corresponding to the access data to be processed, the processing module is further configured to:

[0033] Perform parsing processing on the access data to be processed to obtain at least one target metric item. The at least one target metric item includes: a target user ID and a target IP address;

[0034] Put the at least one target metric item into the statistical data of the data pool, where the data pool is used to store data obtained by parsing the real-time access data acquired from the gateway and performing multi-dimensional statistics based on time periods and at least one metric item.

[0035] In a possible implementation manner, the processing module verifies the to-be-processed access data according to the verification conditions to obtain a verification result corresponding to the to-be-processed access data, and specifically is used for:

[0036] In response to a user's configuration request, perform a configuration operation on at least one verification item in the verification conditions to obtain updated verification conditions;

[0037] Verify the to-be-processed access data according to the updated verification conditions to obtain a verification result corresponding to the to-be-processed access data.

[0038] In a possible implementation manner, the processing module is specifically used for:

[0039] Extract features related to crawler users in the to-be-processed access data as target feature data;

[0040] Input the target feature data into the behavior analysis model to obtain the target label information.

[0041] In a possible implementation manner, the determination module is further used for:

[0042] If the target label information indicates that the to-be-processed access data is a non-crawler user operation, obtain resource usage information of the cluster corresponding to the to-be-processed access data;

[0043] Process the to-be-processed access data according to the resource usage information.

[0044] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0045] The memory stores computer execution instructions;

[0046] The processor executes the computer execution instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0048] Fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which when executed by a processor, implements the above first aspect and / or various possible implementation manners of the first aspect.

[0049] The method, apparatus, device, and storage medium for processing access data provided by the embodiments of the present application obtain the access data to be processed; input the access data to be processed into a behavior analysis model to obtain target label information corresponding to the access data to be processed. The behavior analysis model is trained based on the feature data in multiple historical access data and the label information corresponding to the multiple historical access data. For the feature data and the corresponding label information in any historical access data, the feature data is the feature related to the crawler user, and the label information is whether it is a crawler user; if the target label information indicates that the access data to be processed is an operation by a crawler user, it is determined that the access data to be processed will not be processed. This technical solution can accurately generate target label information for the access data to be processed by obtaining the access data to be processed and inputting it into a behavior analysis model trained by the crawler user-related feature data in multiple historical access data and the corresponding label information of whether it is a crawler user, so as to determine whether it is an operation by a crawler user. If it is determined to be an operation by a crawler user, that is, it is determined that the access data to be processed will not be processed. This process effectively filters out malicious or invalid access data that may be generated by crawler users, reduces the data processing burden of the system, improves the pertinence and efficiency of data processing, avoids the invalid processing of crawler access data, and at the same time protects system resources, prevents crawler users from maliciously accessing and obtaining data from the system, maintains the security and stability of the system, and ensures that the system can focus on processing valid access data generated by normal users, providing a strong guarantee for the normal operation of the system and the accuracy of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0051] Figure 1 It is a schematic diagram of the architecture for processing access data provided by the embodiments of the present application;

[0052] Figure 2 It is a schematic flowchart of the method for processing access data provided by the embodiments of the present application Figure 1 ;

[0053] Figure 3 It is a schematic flowchart of the training process of the behavior analysis model provided by the embodiments of the present application;

[0054] Figure 4 It is a schematic flowchart of the method for processing access data provided by the embodiments of the present application Figure 2;

[0055] Figure 5 Schematic flowchart of the method for processing access data provided by the embodiments of the present application Figure 3 ;

[0056] Figure 6 Schematic structural diagram of the device for processing access data provided by the embodiments of the present application;

[0057] Figure 7 Schematic structural diagram of the electronic device provided by the embodiments of the present application.

[0058] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be provided hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0059] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0060] Technical terms:

[0061] Flink: As an open-source stream processing framework, it has the advantages of high performance, low latency, high throughput, and strong fault tolerance, and is very suitable for processing large-scale and high-concurrency real-time data streams. By leveraging the real-time processing capabilities of the Flink framework, the system can achieve real-time collection and processing of access data at the gateway layer, and record key metrics such as PageView (PV), Daily Active User (DAU), Unique Visitor (UV), etc. Technical background:

[0063] With the rapid development of Internet technology, the number of users and access volume of websites and online services have shown explosive growth. This trend has brought huge business opportunities to Internet enterprises, but also brought unprecedented security threats and management challenges. Especially in a high-traffic environment, the complexity and diversity of user behaviors make traditional static defense strategies appear inadequate and difficult to effectively cope with malicious attacks and abnormal accesses. Therefore, developing a system that can monitor in real time, dynamically configure access rules, and efficiently process large amounts of access data has become an urgent problem for current Internet enterprises.

[0064] In the current Internet environment, the analysis and management of user behavior are crucial for maintaining website security and enhancing the user experience. However, with the increase in the number of users, the complexity and diversity of user behavior have also increased, making it difficult for traditional analysis methods to meet the requirements. Especially for high-traffic users, such as crawler users, their behavior characteristics are often significantly different from those of normal users, but traditional defense measures often struggle to accurately identify and intercept these abnormal behaviors. This not only poses security risks to the website but also may lead to malicious occupation of resources, affecting the access experience of normal users.

[0065] In addition, traditional static defense strategies, such as fixed access restrictions and Internet Protocol (IP) address bans, although can resist malicious attacks to a certain extent, their response speed and flexibility are limited. Facing the rapidly changing network environment and evolving attack methods, static strategies often struggle to adjust and optimize in a timely manner, thus being unable to effectively cope with new threats.

[0066] Based on the above technical problems, the inventor's technical concept is as follows: If, when obtaining streaming access data, it is possible to accurately predict this access data to obtain the access process of a crawler user, and thus reject the processing of the access process of a crawler user, it is possible to avoid security risks brought by malicious access, malicious occupation of resources, and even affecting the access experience of normal users. And to increase the accuracy of prediction, by refining and analyzing the behavior characteristics such as the access path and stay time of high-traffic users or crawler users, the logistic regression model is trained for subsequent prediction implementation.

[0067] Furthermore, to improve the accuracy of prediction, some strategies can also be configured to filter before prediction. For example, limit the number of accesses of the same user within a certain time, limit the access frequency of the same IP address, etc., to preliminarily screen and intercept abnormal access behaviors. At the same time, it also supports further verification of suspected abnormal users through verification means such as verification codes and sliders to improve the accuracy of interception and the user experience.

[0068] Figure 1 The architecture schematic diagram for the processing of access data provided by the embodiments of this application is as Figure 1 shown. This architecture includes: a gateway, Kafka, Flink, a traffic statistics unit, a dynamic rule configuration unit, and a user behavior analysis unit.

[0069] In a possible implementation, when a user triggers an access request, the gateway obtains the access data to be processed corresponding to the access request, and then sequentially passes through Kafka and Flink, and then is processed by the traffic statistics unit, the dynamic rule configuration unit, and the user behavior analysis unit.

[0070] Among them, in the traffic statistics unit, based on the access data to be processed, various indicators are statistically analyzed according to the configuration, such as PV, DAU, UV, monthly active users, user login times, etc.; in the dynamic rule configuration unit (verification conditions), the dimensions of traffic statistics can be configured, such as configuring blacklists, whitelists, IP-based traffic limiting policies, and user ID-based traffic limiting policies; in the user behavior analysis unit, based on the access data to be processed, the interests and hobbies of users are analyzed, crawler users are identified, user portraits are established, and user access paths are established.

[0071] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0072] Figure 2 Flow diagram of the method for processing access data provided by the embodiments of the present application Figure 1 , such as Figure 2 shown, the method includes:

[0073] Step 21, obtain the access data to be processed;

[0074] In this step, when a user makes a data access request (such as business query, data query, web page browsing, etc.), the gateway obtains the access data corresponding to the request, so that the device obtains the access data from the gateway, denoted as the access data to be processed.

[0075] Step 22, input the access data to be processed into the behavior analysis model to obtain the target label information corresponding to the access data to be processed;

[0076] Among them, the behavior analysis model is trained based on the feature data in multiple historical access data and the label information corresponding to the multiple historical access data. For the feature data and the corresponding label information in any historical access data, the feature data is the feature related to crawler users, and the label information is whether it is a crawler user.

[0077] In this step, in order to avoid wasting and occupying system resources, after obtaining the access data to be processed, the access data to be accessed can be predicted and identified to determine whether it is generated by crawler users. At this time, the access data to be processed is input into the trained behavior analysis model. If the access data to be processed is generated by crawler users, it should contain the same features as the feature data in the historical access data. Therefore, the behavior analysis model can perform effective analysis and identification.

[0078] In a possible implementation, the ML library of Flink (such as Flink-ML) can be used or other machine learning frameworks (such as TensorFlow, PyTorch, etc.) can be integrated. Multiple access data to be processed are input into the machine learning model through the DataStream API of Flink to achieve real-time or batch user behavior analysis. At the same time, it can also be displayed in the visualization interface, and the analysis results (i.e., multiple target label information) are presented in the form of charts, reports, etc.

[0079] Optionally, a possible implementation of step 22 can be:

[0080] Step 1: Extract the features related to crawler users in the access data to be processed as target feature data;

[0081] In this implementation, the features related to crawler users are extracted from the access data to be processed. For example: date, Uniform Resource Locator (URL), user ID, request method, timestamp, last request time (calculated based on the timestamp and can be determined from historical data), etc. are used as target feature data.

[0082] Before this implementation, the access data to be processed can be cleaned to remove data with format errors, missing key information, or obvious anomalies.

[0083] Step 2: Input the target feature data into the behavior analysis model to obtain target label information.

[0084] In this implementation, the target feature data is input into the behavior analysis model, and the target label information corresponding to the target feature data, that is, whether it is a crawler user, is output.

[0085] Exemplarily, Figure 3 is a schematic diagram of the training process of the behavior analysis model provided by the embodiment of the present application, as Figure 3 shown, including: data preparation, model training, and model evaluation.

[0086] In the data preparation stage: collect data (multiple historical access data, including: date, URL, user ID, device type, longitude and latitude, client IP, token, User Agent (UA) identifier, request method, timestamp, request parameters, etc.), data cleaning (remove data records with incorrect formats, missing key information, or obvious anomalies in the data. For example, if the UID of a certain record is negative, it can be determined as abnormal data and filtered out), feature extraction (feature data: extract features related to determining whether a user is a crawler from the cleaned data. For example: date, URL, user ID, request method, timestamp, last request time (calculated based on the timestamp), etc. as features).

[0087] In the model training stage: input data (a part of the feature data in multiple historical access data and a part of the label information corresponding to multiple historical access data), train the model, use the sigmoid function for features, the cross-entropy loss function, and the gradient descent algorithm to update parameters, obtain the model parameters (weights and bias terms), and generate a model.

[0088] In the model evaluation stage: prepare a test data set (a part of the feature data in multiple historical access data and another part of the label information corresponding to multiple historical access data), use the cross-validation method, perform prediction evaluation, output the model accuracy, and output a behavior analysis model after meeting the standard.

[0089] Step 23: If the target label information indicates that the access data to be processed is an operation by a crawler user, determine that the access data to be processed will not be processed.

[0090] In this step, if the target label information indicates that the access data to be processed is an operation by a crawler user, then it is considered that the access data to be processed is abnormal, and no subsequent response operation is performed on the access data to be processed.

[0091] Furthermore, if the target label information indicates that the access data to be processed is not an operation by a crawler user, obtain the resource usage information of the cluster corresponding to the access data to be processed; process the access data to be processed according to the resource usage information.

[0092] In this implementation, when the target label information indicates that the access data to be processed is not an operation by a crawler user, it indicates that this access is an operation behavior of a normal user.

[0093] At this time, obtain the resource usage information of the cluster corresponding to the access data to be processed. This cluster is a set of computing resources for processing data, and the resource usage information may include metrics such as CPU usage rate, memory usage, storage resource occupancy, network bandwidth utilization rate, etc. These information reflect the current operating state of the cluster and the remaining resources;

[0094] After that, the to-be-processed access data is processed according to the obtained resource usage information. For example, if the CPU usage rate of the cluster is low, the memory is sufficient, and there is remaining network bandwidth, the device can immediately and efficiently process the to-be-processed access data to meet the user's request; if the resources are relatively tight, the device can process the to-be-processed access data according to certain strategies (such as queuing and waiting, adjusting the processing priority, etc.) to avoid processing failures or affecting other ongoing tasks due to insufficient resources.

[0095] In the above way, while ensuring that normal user access requests are reasonably processed, the processing method can be flexibly adjusted according to the cluster resource status, improving the overall operation efficiency and stability, and realizing the reasonable allocation and utilization of resources.

[0096] In a possible implementation, the APIs of the ResourceManager and JobManager of Flink are used to obtain the resource usage information of the cluster in real time. Through a custom Scheduler or by using the default scheduling policy of Flink, the dynamic scheduling and resource allocation of tasks (i.e., to-be-processed access data) are performed. At the same time, a resource warning and alarm mechanism can be implemented. When the resource usage rate exceeds the preset threshold, an alarm message is sent to the administrator's terminal in a timely manner to enable the administrator to intervene.

[0097] In addition, the performance of the stream processing task can be optimized by adjusting the configuration parameters of Flink (such as parallelism, taskmanager.memory.process.size, etc.). The Savepoint and Checkpoint functions of Flink can also be used to perform fault recovery of tasks and ensure data consistency.

[0098] The processing method for access data provided by the embodiment of the present application includes obtaining the access data to be processed; inputting the access data to be processed into a behavior analysis model to obtain the target label information corresponding to the access data to be processed. The behavior analysis model is trained based on the feature data in multiple historical access data and the label information corresponding to the multiple historical access data. For the feature data and the corresponding label information in any historical access data, the feature data is the feature related to the crawler user, and the label information is whether it is a crawler user; if the target label information indicates that the access data to be processed is an operation by a crawler user, it is determined that the access data to be processed will not be processed. This technical solution can accurately generate target label information for the access data to be processed by obtaining the access data to be processed and inputting it into a behavior analysis model trained by the feature data related to the crawler user in multiple historical access data and the corresponding label information indicating whether it is a crawler user, so as to determine whether it is an operation by a crawler user. If it is determined to be an operation by a crawler user, that is, it is determined that the access data to be processed will not be processed. This process effectively filters out malicious or invalid access data that may be generated by crawler users, reduces the data processing burden of the system, improves the pertinence and efficiency of data processing, avoids the invalid processing of crawler access data, and at the same time protects system resources, prevents crawler users from maliciously accessing and obtaining data from the system, maintains the security and stability of the system, and ensures that the system can focus on processing valid access data generated by normal users, providing a strong guarantee for the normal operation of the system and the accuracy of data.

[0099] Based on the above embodiment, Figure 4 is a schematic flowchart of the processing method for access data provided by the embodiment of the present application Figure 2 , as Figure 4 shown, step 22 can be implemented as follows:

[0100] Step 41: Verify the access data to be processed according to the verification conditions to obtain the verification result corresponding to the access data to be processed;

[0101] Among them, the verification conditions are determined based on at least one verification item. The at least one verification item includes: the number of accesses to a preset IP address within a preset time period is less than the access count threshold, the user identity ID is not in the preset user ID blacklist, and the user ID is in the preset user ID whitelist. The preset IP address is the user IP address or / and the accessed IP address;

[0102] In this step, after obtaining the access data to be processed transmitted by the gateway, the access data to be processed is parsed to obtain the attribute information of the verification items involved in the verification conditions, so as to use the verification conditions to judge the attribute information one by one to obtain the verification result.

[0103] Exemplarily, the verification of the access data to be processed can be as follows:

[0104] 1) The number of accesses to a preset IP address (taking the IP address to be accessed as an example) within a preset time period is less than the access count threshold, that is, the IP address to be accessed involved is parsed from the access data to be processed (i.e., the attribute information). If this IP address is the preset IP address, the number of accesses to this IP address within the preset time period is obtained from the data pool. If it is less than the access count threshold, the verification result is passed; if it is not less than the access count threshold, the verification result is not passed;

[0105] 2) The user identity ID is not in the preset user ID blacklist, that is, the user ID involved is parsed from the access data to be processed (i.e., the attribute information), and it is determined whether this user ID is in the preset user ID blacklist. If it is not in, the verification result is passed; if it is in, the verification result is not passed

[0106] 3) The user ID is in the preset user ID white list, that is, the user ID involved is parsed from the access data to be processed (i.e., the attribute information), and it is determined whether this user ID is in the preset user ID white list. If it is in, the verification result is passed; if it is not in, the verification result is not passed.

[0107] In a possible implementation, the DataStream API of Flink can be used, combined with the message queue corresponding to the access data to be processed, to load and update the verification conditions in real time. At the same time, by defining a custom ProcessorFunction or SinkFunction, the verification conditions are embedded into the stream processing task, so as to realize the verification of the access data to be processed in the message queue.

[0108] Optionally, a possible implementation of step 41 can be:

[0109] Step 1: In response to the user's configuration request, perform a configuration operation on at least one verification item in the verification conditions to obtain the updated verification conditions;

[0110] In this implementation, the verification conditions can be configured to adjust the verification items in the verification conditions to meet different judgment requirements.

[0111] For example, if the configuration request is to delete a certain ID in the blacklist, then this ID in the preset blacklist is deleted.

[0112] Step 2: Verify the access data to be processed according to the updated verification conditions to obtain the verification result corresponding to the access data to be processed.

[0113] In this implementation, after obtaining the updated verification conditions, when the gateway acquires the access data to be processed, the updated verification conditions are used to verify the access data to be processed, and a verification result is obtained.

[0114] Before step 41, the following can also be executed:

[0115] Step 1: Parse and process the access data to be processed to obtain at least one target metric item. The at least one target metric item includes: target user ID, target IP address;

[0116] In this implementation, at least one target metric item in the access data to be processed is extracted. For each target metric item, it can be basic information for real-time calculation and recording of key metrics such as PV, DAU, UV, monthly active users, and user login times.

[0117] Step 2: Put the at least one target metric item into the statistical data of the data pool. The data pool is used to store data obtained through multi-dimensional statistics of the real-time access data acquired by the gateway based on time periods and at least one metric item.

[0118] In this implementation, after obtaining at least one target metric item, the corresponding metric item is found in the data pool, and a +1 operation is performed on the corresponding metric item to update the statistical data of the corresponding metric item.

[0119] That is, according to different configurations, users can view statistical data such as PV, DAU, UV, and monthly active users for different time periods.

[0120] In a possible implementation, Apache Flink is used as the real-time data processing framework. Utilizing its powerful state management and fault tolerance capabilities, the traffic data at the gateway layer is collected and processed in real time. That is, multiple operators (such as Map, Filter, Reduce, etc.) are defined to process the data stream, and intermediate results are maintained through state storage (such as ValueState, ListState, etc.), thus achieving efficient real-time traffic statistics.

[0121] Step 42: If the verification result indicates that the verification conditions are passed, input the access data to be processed into the behavior analysis model to obtain target label information.

[0122] Correspondingly, if the verification result indicates that the verification conditions are not passed, it is directly determined that the access data to be processed is not processed, further improving the efficiency of abnormal interception.

[0123] The access data processing method provided by the embodiments of the present application performs a verification on the to-be-processed access data according to the verification conditions to obtain a verification result corresponding to the to-be-processed access data. Among them, the verification conditions are determined based on at least one verification item, and the at least one verification item includes: the number of accesses to a preset Internet Protocol (IP) address within a preset time duration is less than the access count threshold, the user identity identifier (ID) is not in the preset user ID blacklist, and the user ID is in the preset user ID white list. The preset IP address is the user IP address and / or the accessed IP address. If the verification result indicates that the verification conditions are passed, the to-be-processed access data is input into the behavior analysis model to obtain the target label information. In this technical solution, the to-be-processed access data is preliminarily screened through at least one verification item to avoid waste of resources such as directly performing behavior analysis, etc.

[0124] Based on the above embodiments, Figure 5 is a flowchart illustration of the access data processing method provided by the embodiments of the present application Figure 3 , as Figure 5 shown, an example of the access data processing method is as follows:

[0125] Step 51: Obtain the to-be-processed access data (request);

[0126] Step 52: Configure the verification conditions;

[0127] Step 53: Verify the to-be-processed access data; if passed, execute 54; if not passed, execute 56;

[0128] Step 54: Is it a normal user? If yes, execute 55 and 60; if no, execute 56;

[0129] Step 55: Allow the request;

[0130] Step 56: Reject the request;

[0131] And:

[0132] After step 51, the following steps are also executed:

[0133] Step 57: Kafka, Flink;

[0134] Step 58: Start counting user behavior;

[0135] Step 59: Judge using the user behavior model;

[0136] Step 60: Record whether it is an abnormal user.

[0137] The access data processing method provided by the embodiments of the present application has the same implementation principle and technical effects as those of the above embodiments, and will not be elaborated here.

[0138] Figure 6 The structural schematic diagram of the access data processing device provided by the embodiment of the present application is as follows Figure 6 As shown, the access data processing device provided by this embodiment includes:

[0139] An acquisition module 61, configured to acquire access data to be processed;

[0140] A processing module 62, configured to input the access data to be processed into a behavior analysis model to obtain target label information corresponding to the access data to be processed. The behavior analysis model is trained based on feature data in multiple historical access data and label information corresponding to the multiple historical access data. For the feature data and corresponding label information in any historical access data, the feature data is features related to crawler users, and the label information is whether it is a crawler user;

[0141] A determination module 63, configured to determine that the access data to be processed is not processed if the target label information indicates that the access data to be processed is an operation by a crawler user.

[0142] In a possible implementation manner, the processing module 62 is specifically configured to:

[0143] Verify the access data to be processed according to verification conditions to obtain a verification result corresponding to the access data to be processed. The verification conditions are determined based on at least one verification item. The at least one verification item includes: the number of accesses to a preset Internet protocol IP address within a preset time duration is less than an access count threshold, the user identity identifier ID is not in a preset user ID blacklist, and the user ID is in a preset user ID whitelist. The preset IP address is the user IP address or / and the accessed IP address;

[0144] If the verification result indicates passing the verification conditions, input the access data to be processed into the behavior analysis model to obtain target label information.

[0145] In a possible implementation manner, before verifying the access data to be processed according to verification conditions to obtain a verification result corresponding to the access data to be processed, the processing module 62 is further configured to:

[0146] Perform parsing processing on the access data to be processed to obtain at least one target metric item. The at least one target metric item includes: a target user ID, a target IP address;

[0147] Put the at least one target metric item into the statistical data of the data pool. The data pool is used to store data obtained by parsing real-time access data acquired by the gateway and statistically analyzing in multiple dimensions of time period and at least one metric item.

[0148] In a possible implementation, the processing module 62 verifies the access data to be processed according to the verification conditions, and obtains a verification result corresponding to the access data to be processed. Specifically, it is used for:

[0149] In response to a user's configuration request, perform a configuration operation on at least one verification item in the verification conditions to obtain updated verification conditions;

[0150] Verify the access data to be processed according to the updated verification conditions, and obtain a verification result corresponding to the access data to be processed.

[0151] In a possible implementation, the processing module 62 is specifically used for:

[0152] Extract the features related to the crawler user in the access data to be processed as target feature data;

[0153] Input the target feature data into the behavior analysis model to obtain target label information.

[0154] In a possible implementation, the determination module 63 is further used for:

[0155] If the target label information indicates that the access data to be processed is a non-crawler user operation, obtain the resource usage information of the cluster corresponding to the access data to be processed;

[0156] Process the access data to be processed according to the resource usage information.

[0157] The access data processing device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, so details are not described here in this embodiment.

[0158] Figure 7 This is a schematic structural diagram of the electronic device provided in the embodiment of the present application. As Figure 7 shown, the electronic device provided in this embodiment includes: at least one processor 71 and a memory 72. Optionally, the electronic device further includes a communication component 73. Among them, the processor 71, the memory 72, and the communication component 73 are connected through a bus 74.

[0159] In the specific implementation process, at least one processor 71 executes the computer execution instructions stored in the memory 72, so that at least one processor 71 executes the above method.

[0160] The specific implementation process of the processor 71 can refer to the above method embodiment, and its implementation principle and technical effect are similar, so details are not described here in this embodiment.

[0161] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or implemented by the combination of the hardware and software modules in the processor.

[0162] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0163] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0164] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0165] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0166] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0167] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0168] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the couplings or direct couplings or communication connections shown or discussed between each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0169] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0170] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0171] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0172] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks or optical discs that can store program codes.

[0173] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present invention. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for processing data access, characterized in that, Including: Obtain access data to be processed; Input the access data to be processed into a behavior analysis model to obtain target label information corresponding to the access data to be processed. The behavior analysis model is trained based on feature data in multiple historical access data and label information corresponding to the multiple historical access data. For the feature data and corresponding label information in any historical access data, the feature data is features related to crawler users, and the label information is whether it is a crawler user; If the target label information indicates that the access data to be processed is an operation by a crawler user, determine that the access data to be processed is not processed.

2. The method according to claim 1, characterized in that, The step of inputting the access data to be processed into a behavior analysis model to obtain target label information corresponding to the access data to be processed includes: Verify the access data to be processed according to verification conditions to obtain a verification result corresponding to the access data to be processed. The verification conditions are determined based on at least one verification item. The at least one verification item includes: the number of accesses to a preset Internet Protocol (IP) address within a preset time period is less than an access count threshold, the user identity identifier (ID) is not in a preset user ID blacklist, and the user ID is in a preset user ID whitelist. The preset IP address is the user IP address or / and the accessed IP address; If the verification result indicates that the verification conditions are passed, input the access data to be processed into the behavior analysis model to obtain the target label information.

3. The method according to claim 2, wherein Before verifying the access data to be processed according to the verification conditions to obtain a verification result corresponding to the access data to be processed, the method further includes: Parse the access data to be processed to obtain at least one target metric item. The at least one target metric item includes: a target user ID and a target IP address; Put the at least one target metric item into the statistical data of a data pool. The data pool is used to store data obtained by parsing real-time access data acquired by a gateway and performing multi-dimensional statistics based on time periods and at least one metric item.

4. The method according to claim 2, characterized in that, The step of verifying the access data to be processed according to the verification conditions to obtain a verification result corresponding to the access data to be processed includes: In response to a user's configuration request, perform a configuration operation on at least one verification item in the verification conditions to obtain updated verification conditions; Verify the access data to be processed according to the updated verification conditions to obtain a verification result corresponding to the access data to be processed.

5. The method according to any one of claims 1-4, characterized in that, The step of inputting the access data to be processed into a behavior analysis model to obtain target label information corresponding to the access data to be processed includes: Extract features related to crawler users in the access data to be processed as target feature data; Input the target feature data into the behavior analysis model to obtain the target label information.

6. The method according to any one of claims 1 to 4, characterized in that The method further includes: If the target label information indicates that the access data to be processed is a non-crawler user operation, obtain resource usage information of the cluster corresponding to the access data to be processed; Process the to-be-processed access data according to the resource usage information.

7. A processing device for accessing data, characterized in that, Including: An acquisition module, configured to acquire to-be-processed access data; A processing module, configured to input the to-be-processed access data into a behavior analysis model to obtain target label information corresponding to the to-be-processed access data. The behavior analysis model is obtained by training a logistic regression model based on feature data in a plurality of historical access data and label information corresponding to the plurality of historical access data. For the feature data and corresponding label information in any historical access data, the feature data is a feature related to a crawler user, and the label information is whether it is a crawler user; A determination module, configured to determine that the to-be-processed access data is not processed if the target label information indicates that the to-be-processed access data is an operation by a crawler user.

8. An electronic device, characterized in that, Including: A memory and a processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1-6.

10. A computer program product, characterized in that, Including a computer program, which when executed by a processor implements the method according to any one of claims 1-6.