A data processing method and apparatus

By using a target language processing model to compare the target user's end sequence and cloud sequence, web crawler users can be identified and distinguished. This solves the problem of insufficient accuracy and coverage in malicious crawler identification, and achieves more efficient crawler detection and information protection.

CN116827591BActive Publication Date: 2026-07-21RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2023-04-25
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify malicious web crawlers, leading to the leakage of user privacy or corporate information, and their detection accuracy and coverage are inadequate.

Method used

By acquiring the target user's end sequence and cloud sequence, and using a pre-trained target language processing model, such as BERT, the matching degree between the end sequence and cloud sequence is determined to distinguish between web crawler users and normal users.

Benefits of technology

It improves the accuracy and coverage of web crawler detection, enabling timely identification and intervention of web crawler behavior, and protecting the information security of users and enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116827591B_ABST
    Figure CN116827591B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a data processing method and device, the embodiment of the application inputs the end sequence and the cloud sequence corresponding to the target user into the target language processing model obtained by pre-training, determines whether the end sequence and the cloud sequence match, thereby judging whether the access operation of the target user is forged or tampered, and further determining whether the target user is a network crawler user. The embodiment of the application detects the end sequence and the cloud sequence corresponding to the access operation of the target user one by one, thereby improving the accuracy and coverage of the crawler detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and more specifically, to a data processing method and apparatus. Background Technology

[0002] A web crawler (or simply crawler) is a program or script that automatically retrieves information from the internet according to certain rules. Generally, web crawlers help people find the web pages they need, providing positive assistance. For example, search engines can use web crawlers to retrieve web page information, thereby returning richer search results to users.

[0003] However, in some cases, malicious web crawlers emerge, posing two main threats. Firstly, they consume server bandwidth, crowding out legitimate users' traffic and increasing bandwidth costs. More seriously, they steal user information, leading to the leakage or misuse of user privacy or corporate information assets. For example, some criminals may trick users into granting authorization to crawl their credit scores or credit limits in third-party payment applications. Therefore, effectively identifying web crawlers and taking appropriate measures to stop them, thereby preventing the leakage of user or corporate information, becomes crucial. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a data processing method and apparatus to improve the accuracy and coverage of crawler detection.

[0005] Firstly, a data processing method is provided, the method comprising:

[0006] Obtain the terminal sequence and cloud sequence corresponding to the target user. The terminal sequence is the action sequence of the target user making a request from the terminal device to the cloud. The cloud sequence is the action sequence of the cloud responding to the target user's request and returning corresponding data.

[0007] The end sequence and the cloud sequence are input into a pre-trained target language processing model to determine the matching degree between the end sequence and the cloud sequence.

[0008] The target user category is determined based on the matching degree, and the category includes web crawler users and normal users.

[0009] In a second aspect, a data processing apparatus is provided, the apparatus comprising:

[0010] The acquisition module is used to acquire the terminal sequence and cloud sequence corresponding to the target user. The terminal sequence is the action sequence in which the target user makes a request from the terminal device to the cloud, and the cloud sequence is the action sequence in which the cloud responds to the target user's request and returns corresponding data.

[0011] The first determining module is used to input the end sequence and the cloud sequence into a pre-trained target language processing model to determine the matching degree between the end sequence and the cloud sequence;

[0012] The second determining module is used to determine the target user category based on the matching degree, the category including web crawler users and normal users.

[0013] Thirdly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and the computer program, when executed by a processor, implements the method described in the first aspect.

[0014] Fourthly, an electronic device is provided, including a memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect.

[0015] This invention, through its embodiments, inputs the endpoint sequence and cloud sequence corresponding to the target user into a pre-trained target language processing model to determine whether the endpoint sequence and cloud sequence match. This determines whether the target user's access operation has been forged or tampered with, and thus whether the target user is a web crawler. This invention improves the accuracy and coverage of crawler detection by comparing and detecting the endpoint sequence and cloud sequence corresponding to the target user's access operation one by one. Attached Figure Description

[0016] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0017] Figure 1 This is a schematic diagram of the data processing system according to an embodiment of the present invention;

[0018] Figure 2 This is a flowchart of a data processing method according to an embodiment of the present invention;

[0019] Figure 3 This is a schematic diagram of the target language processing model according to an embodiment of the present invention;

[0020] Figure 4 This is a schematic diagram of the structure of the Transformer according to an embodiment of the present invention;

[0021] Figure 5 This is a flowchart of the method for training a target language processing model according to an embodiment of the present invention;

[0022] Figure 6 This is one of the schematic diagrams illustrating the determination of the model training dataset in an embodiment of the present invention;

[0023] Figure 7 This is the second schematic diagram illustrating the determination of the model training dataset in an embodiment of the present invention;

[0024] Figure 8 This is a visual diagram illustrating the corresponding vectors of the normal user terminal sequence and the cloud sequence in an embodiment of the present invention.

[0025] Figure 9 This is a visual diagram illustrating the corresponding vectors of the web crawler user-end sequence and the cloud sequence in an embodiment of the present invention.

[0026] Figure 10 This is a data flow diagram of data processing according to an embodiment of the present invention;

[0027] Figure 11 This is a schematic diagram of a data processing device according to an embodiment of the present invention;

[0028] Figure 12 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0029] The present invention is described below based on embodiments, but the invention is not limited to these embodiments. In the detailed description of the invention below, certain specific details are described in detail. Those skilled in the art will fully understand the invention even without these details. To avoid obscuring the essence of the invention, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0030] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0031] Unless the context explicitly requires it, words such as "including" or "contains" in the instruction manual should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0032] In the description of this invention, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0033] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only under the premise of legality (e.g., obtaining the consent of the personal information subject, or being necessary for the performance of a contract, or complying with laws and administrative regulations). Furthermore, processing will only be conducted within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0034] Figure 1 This is a schematic diagram of a data processing system according to an embodiment of the present invention. Figure 1 As shown, the data processing system includes a server 11 and multiple terminals 12. The server 11 can be a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The terminal 12 is a terminal device held by the target user. Specifically, a terminal device refers to a general-purpose terminal capable of running web browsers, applications, or mini-programs required for web crawler detection, such as a mobile phone, computer, or tablet computer. In some embodiments, the terminal 12 can also be a dedicated terminal with an application program embedded in an application-specific integrated circuit (ASIC). The open platform provides third-party developers with an Application Programming Interface (API), enabling them to develop their own mini-programs based on the open platform. These mini-programs can be programs developed based on applications within the platform and used to perform corresponding operations.

[0035] Web crawlers, also known as web spiders, web robots, or web chasers, are programs or scripts that automatically retrieve information from the internet according to certain rules. Most web crawlers follow a process of "sending a request - obtaining the page - parsing the page - extracting and storing the content," simulating the process of a normal user obtaining information using a browser, application, or app.

[0036] During the data acquisition phase, server 11 completes data acquisition by interacting with terminal 12. In this process, by embedding tracking points on server 11 and terminal 12, the behavioral sequences of the target user when using the relevant web browser, application, or mini-program requiring web crawler detection can be obtained. Tracking is a common data acquisition method, mainly involving the capture, processing, and transmission of target user behavior or events, along with related technologies and implementation processes. Tracking points are the source of data, and the collected data can be used to analyze the usage of the target application or user behavior habits. Tracking trigger operations can include clicks, exposures (i.e., the target user entering a page or refreshing a page), and page dwell time recording (i.e., recording the time the target user spends on a page). The following explanation uses an application as an example; the application currently being used by the target user and requiring web crawler detection is the target application. Specifically, when the target user is a normal user, terminal 12 determines the tracking trigger operations generated by the target user when using the target application through the human-computer interaction interface and records the corresponding behavioral sequence, i.e., the terminal sequence, in the terminal log data. Simultaneously, the request corresponding to the event triggered by the event tracking is sent to server 11. Server 11 responds to the request by returning the corresponding data to terminal 12 and storing the corresponding behavior sequence, i.e., the cloud sequence, on the server. When the target user is a web crawler, the web crawler simulates a normal user by calling the target application's Application Programming Interface (API) to directly send a request to the server to obtain information. Therefore, no terminal sequence is generated on the terminal device. However, when the web crawler sends a request to server 11 and server 11 responds to the request by returning the corresponding data to terminal 12, a cloud sequence is generated and stored on server 11.

[0037] During the data processing phase, server 11 retrieves terminal log data from terminal 12 and searches for terminal sequences corresponding to each cloud sequence within the terminal log data. If a terminal sequence corresponding to a cloud sequence exists, the target user is a normal user; otherwise, the target user is a web crawler user. Specifically, server 11 retrieves terminal log data from terminal 12 and inputs the target user's terminal sequence and the cloud sequence stored on server 11 into a pre-trained target language processing model. The terminal sequence and cloud sequence are compared to determine the matching degree between them, thereby determining whether the target user's access operation has been forged or tampered with. If the matching degree is higher than a predetermined matching degree threshold, it indicates that the target user's access operation is a genuine operation by the target user, and the target user is a normal user. If the matching degree is not higher than the predetermined threshold, it indicates that the target user's access operation is not a genuine operation by the target user, but rather a web crawler simulating a normal user using the target application, and the target user is a web crawler user.

[0038] In one possible implementation, server 11 can automatically perform web crawler detection according to a preset detection period, such as 1 hour, 6 hours, or 12 hours. When the detection period is reached, terminal 12 proactively reports terminal log data to server 11. Upon receiving the terminal log data, server 11 begins web crawler detection. Alternatively, upon reaching the detection period, server 11 sends a terminal log retrieval command to terminal 12, and terminal 12 responds by sending terminal log data to server 11.

[0039] In one possible implementation, when a cloud sequence is detected, the server 11 can immediately send a terminal log retrieval instruction to the terminal 12 to obtain the corresponding terminal sequence for web crawler detection.

[0040] In one possible implementation, the server can also determine when to perform web crawler detection based on actual conditions and needs. That is, the web browser, application, or mini-program corresponding to the target page sends a web crawler detection instruction to the server 11 based on actual conditions and needs. After receiving the web crawler detection instruction, the server 11 begins to perform web crawler detection.

[0041] This invention improves the accuracy and coverage of web crawler detection by inputting the endpoint sequence and cloud sequence corresponding to the target user into a pre-trained target language processing model to determine whether the endpoint sequence and cloud sequence match, thereby judging whether the target user's access operation has been forged or tampered with, and whether the target user is a web crawler. This invention improves the accuracy and coverage of web crawler detection by comparing and detecting the endpoint sequence and cloud sequence corresponding to the target user's access operation one by one.

[0042] Figure 2 This is a flowchart of a data processing method according to an embodiment of the present invention, such as... Figure 2 As shown, the data processing method includes the following steps:

[0043] In step S201, the terminal sequence and cloud sequence corresponding to the target user are obtained. The terminal sequence is the action sequence of the target user making a request from the terminal device to the cloud, and the cloud sequence is the action sequence of the cloud responding to the target user's request and returning corresponding data.

[0044] The target user refers to the user of the web browser, application, or mini-program that needs to be detected for web crawling. The following explanation uses an application as an example; the application that needs to be detected for web crawling is the target application.

[0045] The target user's corresponding terminal and cloud sequences can be determined from the terminal log data and cloud log data corresponding to the target user. Specifically, the terminal log data can be obtained by embedding tracking points on the terminal, and the cloud log data can be obtained by embedding tracking points on the cloud. The user behavior sequences obtained through tracking points are recorded in the corresponding log data. Terminal tracking points can record most of the target user's behaviors, including behaviors that require sending requests to the server and behaviors that do not. For example, recording the duration of the target user's stay on the target page does not require sending requests to the server, but the terminal log data still records the behavioral information corresponding to the page stay duration. Therefore, terminal tracking points can collect more comprehensive target user behavior data. However, when performing web crawler detection, terminal log data needs to be reported to the server, which may result in reporting delays due to poor network connectivity. Server tracking points can collect target user behavior data in real time, with high accuracy and no reporting delays. It can also more accurately collect target user behavior data for the target application's key business logic.

[0046] Since some user actions do not require requests to the server, there is no unique correspondence between the endpoint sequence and the cloud sequence for the target user. Generally, the ratio of endpoint sequence to cloud sequence is mostly concentrated in the range of 2-8. Besides the fact that some user actions do not need to be reported to the server, another reason for the lack of a unique correspondence between endpoint and cloud sequences may be that one endpoint action of the target user corresponds to multiple server requests. That is, one endpoint sequence can correspond to multiple cloud sequences. For example, if the endpoint action is to confirm payment, the corresponding server requests may include requesting a payment message from a third-party payment server, receiving a payment message, and sending a payment message to the terminal device. Other possibilities include packet loss or delays in endpoint log data reporting to the server, and incomplete endpoint tracking.

[0047] Specifically, when a target user uses the target application through a terminal device, the relevant controls on the current page of the target application are triggered, and a corresponding request is sent to the server. The user's behavior is recorded locally on the terminal device, resulting in a corresponding behavior sequence, or endpoint sequence, which is stored in the terminal log data. The server responds to the received request by sending corresponding data to the terminal device, recording the behavior on the server, and obtaining a corresponding behavior sequence, or cloud sequence, which is stored in the cloud log data. When web crawler detection is required, the server retrieves the terminal log data from the terminal and extracts the corresponding endpoint and cloud sequences from both the terminal log data and the cloud log data.

[0048] In one possible implementation, obtaining the target user's endpoint sequence and cloud sequence can specifically involve acquiring the target user's endpoint log data and cloud log data within a preset historical time period. Then, based on the endpoint log data and the cloud log data, the target user's endpoint sequence and cloud sequence are determined. In other words, when analyzing the target user's behavior, the analysis can be performed on behavior within a preset historical time period. This preset historical time period can be a pre-set detection cycle, such as 1 hour, 6 hours, or 12 hours. Alternatively, it can be determined based on instructions sent by the target application's operator, such as instructions requiring a historical time period of one week, 15 days, or one month for detection.

[0049] In this embodiment, web crawler detection based on endpoint sequences and cloud sequences within a preset historical time period can periodically detect users of the target application, promptly identifying web crawler users for timely intervention. In one possible implementation, obtaining the endpoint and cloud sequences corresponding to the target user can specifically involve acquiring the target user's endpoint log data and cloud log data, respectively. Then, based on the endpoint and cloud log data, multiple operation batches are determined, each batch including multiple endpoint and cloud sequences for the target user. Here, an operation batch refers to a batch processing script, which is a sequence of actions corresponding to the target object processed in batches. An operation batch includes a set of consecutive operations; for example, when the target application is an e-commerce application, "purchase," "confirm payment," and "enter password" for a product constitute a set of operation sequences.

[0050] In this implementation, web crawler detection is performed on an operation batch basis, ensuring comprehensive detection of target user behavior. Furthermore, the correlation between actions within each operation batch helps the target language processing model recognize contextual information, thereby improving the model's accuracy.

[0051] After obtaining the target user's endpoint sequence and corresponding cloud sequence through the above implementation method, further data analysis can be performed on the endpoint sequence and corresponding cloud sequence.

[0052] In step S202, the end sequence and the cloud sequence are input into a pre-trained target language processing model to determine the matching degree between the end sequence and the cloud sequence.

[0053] Since each element in the end sequence and cloud sequence can be considered a word or character in natural language, and the sequence (end sequence or cloud sequence) can be viewed as a sentence as a whole, the matching of end sequences and cloud sequences can be handled using a language processing model. Due to the different forms of sequences, there are two forms used to represent sequences: word sequences and template sequences.

[0054] The target language processing model can be a model based on BERT (Bidirectional Encoder Representations from Transformers). BERT is a novel bidirectional pre-trained language model that uses bidirectional transformer encoding. This encoding method employs an attention mechanism during model construction and considers contextual features bidirectionally when predicting words. Therefore, it can more accurately obtain the relationships between words. BERT's advancement is based on two points: First, it uses novel pre-training tasks such as MLM (Masked Language Model) and NSP (Next Sentence Prediction). Second, it has ample data and computational power to support BERT's training intensity. In contrast, traditional state-of-the-art (SOTA) models such as Word2Vec (Word to Vector Model) or ELMO (Embedding from Language Models) generate pre-training methods that use unidirectional training from left to right or shallow bidirectional training, and cannot achieve the bidirectionality of the BERT model.

[0055] The most direct way to train a target language processing model is to use word sequences as input. However, log data is massive, with each sequence consisting of multiple word sequences. When represented by word sequences, the sequence length becomes inflated, sometimes several times or even tens of times longer than the log sequence, and the lengths are inconsistent. For example, the content length of each log sequence is approximately 3 to 95. Therefore, a sequence of length 3 might correspond to word sequences of length 9 to 285. Furthermore, the lengths of individual word sequences may not be consistent. To ensure consistent input sequence lengths, the BERT model uses a maximum value (MAXLEN) parameter to guarantee that each input data point has a consistent length. When the length of the input data exceeds MAXLEN, truncation occurs, removing the portion exceeding MAXLEN. When the length of the input data is less than MAXLEN, padding is performed. Processing occurs for both lengths exceeding and falling below MAXLEN, increasing computational complexity and time consumption. The longer MAXLEN is, the greater the computational power and memory required by the model, and the longer the training time. Therefore, as an optional implementation, the maximum length of the MAXLEN parameter is set to no more than 512 to reduce computational complexity and time consumption. Compared to word sequences, template sequences have the same length as each sequence in the log data, and all template sequences have a consistent sequence length. At the same sequence length, the template sequence is shorter than the word sequence. With the same processing length in the BERT model, the template sequence can represent significantly more log content than the word sequence. Compared to word sequences, the length of the template sequence depends on a fixed window size; therefore, as an optional implementation, the MAXLEN parameter for the template sequence is set to a fixed window size. In this case, the model can completely capture all log data in the template sequence. This way, the model will not need to perform additional processing on each sequence in the log data, such as truncation or padding.

[0056] The BERT model framework consists of a pre-trained model and a fine-tuning model. During pre-training, the model is trained on unlabeled data through various pre-training tasks. Because the BERT model uses a bidirectional transformer architecture for encoding, the weight of each word is transmitted to the context words, and these words have a certain correlation with each other. After pre-training, the BERT model outputs a vector of the input sequence. To perform the next fine-tuning operation, the pre-trained parameters are used to initialize the BERT fine-tuning model, and then all parameters are fine-tuned using labeled data from downstream tasks. In other words, when a sequence is input into the BERT model, the BERT model outputs a vector corresponding to the sequence, which is then classified according to a trained linear classifier.

[0057] Figure 3This is a schematic diagram of the target language processing model according to an embodiment of the present invention. The target language processing model is a deep language learning model obtained by fine-tuning the BERT model. Figure 3 As shown, the target language processing model may include an input layer 310, a BERT model 320, a fully connected layer 330, and an output layer 340. This embodiment of the invention fine-tunes the parameters of the BERT model to enable the target language processing model to better perform comparative analysis of end-to-end sequences and cloud sequences.

[0058] The input layer 310 may include a word segmentation tool to segment the input end sequences or cloud sequences. The word segmentation tool can directly segment the end sequences and cloud sequences, or it can first parse the end sequences and cloud sequences to obtain the corresponding actions or events for each end sequence and cloud sequence before segmentation. Different word segmentation methods correspond to different BERT models 320. If sequence parsing is required, the input layer can also include a parsing tool, such as the Drain parsing tool. Drain parses each end sequence and cloud sequence in real time in a streaming manner. Furthermore, Drain uses a fixed-depth parse tree to speed up the parsing process. Each end sequence and cloud sequence is mapped to a corresponding action or event. Other parsing methods can also be used to replace the Drain parsing tool, such as clustering-based methods and heuristic-based methods. Figure 3 As shown, in input layer 310, the CLS flag is placed at the beginning of the first sequence to indicate the processing type. The SEP flag is used to separate two input sequences. For example, if the input sequences are A and B, the SEP flag is added between sequences A and B. If only one sequence A is input, the SEP flag is added at the end of A. A1, A2...A 30 These are the word segments obtained after segmenting the end sequence or cloud sequence.

[0059] After inputting the obtained word segments into the BERT model 320, each segment is first converted into a corresponding word vector. Word vectors are typically high-dimensional vectors, such as 768-dimensional vectors. Here, E represents the embedded vector. After conversion to word vectors, the word vectors are processed through multiple Transformers. Each Transformer operates based on the left and right contexts from all layers. Figure 4 This is a schematic diagram of the Transformer (converter) according to an embodiment of the present invention. Figure 4As shown, the Transformer is an encoder-decoder structure, formed by stacking several encoders and decoders. In the diagram, one Trm corresponds to one Transformer. The Transformer is used to transform the input word vectors into feature vectors. Generally, encoder 410 contains 6 sub-encoders, and similarly, decoder 420 contains 6 sub-decoders. Each encoder has the same structure but different parameters. The input to each sub-encoder is the output of the previous sub-encoder, and the encoder's output can be set to a vector with the same dimensions as the input vector. The input to each sub-decoder includes not only the output of its previous decoder but also the output of the entire encoding part. After processing by the Transformer, the feature vectors corresponding to the word vectors are obtained. In the diagram, T represents the feature vector obtained after processing by the BERT model, and C is a linear classifier that outputs the classification result.

[0060] Finally, BERT essentially operates on a self-supervised learning method based on massive amounts of training data. It provides a model for transfer learning to other tasks. When performing different language information processing tasks, BERT can be jointly trained by adding output layers with different functions according to the task objective. In other words, BERT can be fine-tuned for specific tasks and used as a feature extractor before further processing of language vectors. In this embodiment, the target language processing model adds a fully connected layer 330 after the BERT model 320, projecting the vectors generated by the BERT model 320 onto a higher-dimensional vector. Each dimension corresponds to a unique word score. Then, the scores of the entire end sequence and cloud sequence are converted into matching probabilities, i.e., matching degrees. Finally, the matching degree between the end sequence and the corresponding cloud sequence is output through the output layer 340.

[0061] The target language processing model in this embodiment of the invention uses the BERT model as its backbone, stacking multiple Transformer encoders together. The Transformer, based on the well-known multi-head attention module, can more thoroughly capture the bidirectional relationships between different sentences, resulting in more accurate matching between the end-to-end sequence and the cloud sequence. The BERT in the target language processing model can also be replaced with deep learning models such as ALBERT (Lightweight Bidirectional Encoder Transformer, ALite BERT) or RoBERT (Robustly Optimized Bidirectional Encoder Transformer).

[0062] Once the structure of the target language processing model is determined, pre-training can begin. Before training, all parameters in the target language processing model are initialized, and during training, these parameters are fine-tuned according to the training objectives. This allows the trained target language processing model to more accurately determine the matching degree between end sequences and cloud sequences. Specifically, after representing the end sequence and cloud sequence corresponding to the same operation of the target user as corresponding vectors, the target language processing model needs to make the distance between the end sequence vector and the cloud sequence vector as close as possible, and the distance between the end sequence vector corresponding to this operation and the cloud sequence vector corresponding to other operations as far as possible.

[0063] Figure 5 This is a flowchart of a method for training a target language processing model according to an embodiment of the present invention. Figure 5 As shown, the method for training the target language processing model includes the following steps:

[0064] In step S501, the terminal sequence and cloud sequence corresponding to multiple users are obtained.

[0065] It is worth noting that the terminal sequence and cloud sequence corresponding to each user will be processed under the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of the contract, or complying with the provisions of laws and administrative regulations, etc.), and will only be processed within the scope of the regulations or agreements.

[0066] In step S502, the model training dataset is determined based on the terminal sequence and cloud sequence corresponding to the multiple users.

[0067] In one possible implementation, BERT can be pre-trained using the NSP task. NSP allows BERT to deduce relationships between sequences by predicting whether consecutive sequences are coherent. Therefore, when constructing the model training dataset, cloud sequences and corresponding end sequences of the same user can be placed adjacently, and then sorted with sequences from different users. That is, the sorted sequence would be: cloud sequence of user 1, end sequence of user 1, cloud sequence of user 2, end sequence of user 3, and so on.

[0068] Figure 6 This is one of the schematic diagrams illustrating the determination of the model training dataset in an embodiment of the present invention. After obtaining the end sequences and cloud sequences corresponding to multiple users, the order of the cloud sequences and / or the end sequences in an operation batch can be shuffled according to a preset proportion. This preset proportion can be set according to actual conditions and needs, such as 15%, 30%, or 45%. In response to a cloud sequence matching the corresponding end sequence, the label of the cloud sequence and the end sequence is set to 0; in response to a cloud sequence not matching the corresponding end sequence, the label of the cloud sequence and the end sequence is set to 1, thereby determining the model training dataset.

[0069] like Figure 6 As shown, the cloud sequences corresponding to users 1, 2, 3, and 4 are C1, C2, C3, and C4, respectively, and their corresponding end sequences are E1, E2, E3, and E4, respectively. The order of some end sequences corresponding to users is shuffled. For example, the order of the end sequences corresponding to users 2 and 3 is shuffled. After shuffling, the cloud sequence C1 corresponding to user 1 matches the end sequence E1, so the label for user 1 is set to 0. The cloud sequence C2 corresponding to user 2 does not match the shuffled end sequence E3, so the label for user 2 is set to 1. The cloud sequence C3 corresponding to user 3 does not match the shuffled end sequence E2, so the label for user 3 is set to 1. The cloud sequence C4 corresponding to user 4 matches the end sequence E4, so the label for user 4 is set to 0. Performing the above operation on all acquired users determines the labeled training data, and all training data constitute the model training dataset.

[0070] In one possible implementation, the SimCSE (Simple Contrastive Learning of Sentence Embeddings) model can be used to determine the matching degree between end sequences and cloud sequences. Strictly speaking, SimCSE is not a new model, but a framework based on contrastive learning. SimCSE can be combined with the BERT model for unsupervised language contrastive learning. SimCSE can use natural language inference for contrastive learning training, and its task is to determine the relationship between two sentences. The possible relationships are entailment, contradiction, or neutrality. SimCSE needs to construct positive and negative samples for contrastive learning. Therefore, a set of end sequences and cloud sequences with entailment can be naturally used as positive samples. If other sets of end sequences and cloud sequences in the same batch are used as negative samples, the model training dataset can be successfully constructed.

[0071] Figure 7 This is the second schematic diagram illustrating the determination of the model training dataset in an embodiment of the present invention. After obtaining the end sequences and cloud sequences corresponding to multiple users, the end sequences and corresponding cloud sequences of the same user are used as positive samples, and the end sequences and cloud sequences belonging to different users are used as negative samples, thereby determining the model training dataset.

[0072] like Figure 7As shown, the cloud sequences corresponding to users 1, 2, 3, and 4 are C1, C2, C3, and C4, respectively, and their corresponding end sequences are E1, E2, E3, and E4, respectively. Taking user 1 as an example, user 1's cloud sequence C1 and its corresponding end sequence E1 are set as positive samples, and C1 and the end sequences E2, E3, and E4 of other users are set as negative samples, respectively. This allows us to determine the model training dataset with positive and negative samples.

[0073] In step S503, the target language processing model is trained based on the model training dataset.

[0074] Specifically, in target language processing models, cosine similarity can be used to represent the distance between corresponding vectors of end sequences and corresponding vectors of cloud sequences in the vector space. Therefore, the following loss function can be used when training a target language processing model using a model training dataset:

[0075]

[0076] Where i represents the order of vector features, l i Let i represent the loss function corresponding to the i-th set of vector features. A set of vector features includes a vector feature corresponding to an end sequence and a vector feature corresponding to a cloud sequence. The vector features of the cloud sequence, The sim() function refers to cosine similarity, which is the vector feature of the end sequence. Vector features and The cosine similarity is given by γ, where γ is the temperature parameter. The temperature parameter determines the degree of attention given to difficult negative samples in contrastive learning; the larger the temperature coefficient, the greater the attention given to difficult negative samples.

[0077] After the target language processing model is continuously converged by the above loss function, the end sequence and cloud sequence corresponding to the target user are input into the target language processing model to determine the matching degree between the end sequence and the cloud sequence.

[0078] Specifically, the matching degree is used to characterize whether a set of end sequences and cloud sequences of the input target language processing model correspond to the same operation of the target user. That is, whether the distance between the feature vectors corresponding to the end sequences and the feature vectors corresponding to the cloud sequences in the vector space is close enough. The closer the distance, the greater the matching degree, and the greater the probability that the end sequences and cloud sequences correspond to the same operation. The farther the distance, the smaller the matching degree, and the less likely that the end sequences and cloud sequences correspond to the same operation.

[0079] In step S203, the target user category is determined based on the matching degree, and the category includes web crawler users and normal users.

[0080] Here, a web crawler user refers to a user whose account is used by a web crawler to access the target application, while a normal user refers to a user whose account is not used by a web crawler to access the target application. A matching threshold can be preset. If the matching degree between the target user's sequence and the cloud sequence is higher than the matching degree threshold, the target user is determined to be a normal user; otherwise, the target user is determined to be a web crawler user.

[0081] After the end sequence and cloud sequence are input into the target language processing model, each sequence will be transformed into a corresponding multi-dimensional vector, such as 768-dimensional. To make the matching degree more intuitive, the multi-dimensional vector can be transformed into a two-dimensional or three-dimensional vector.

[0082] Figure 8 This is a visual diagram illustrating the corresponding vectors of normal user-end sequences and cloud sequences in an embodiment of the present invention. For example... Figure 8 As shown, after dimensionality reduction of the vectors corresponding to the normal user terminal sequence and the cloud sequence, it can be seen that the distance between vector 801 corresponding to the normal user terminal sequence and vector 802 corresponding to the cloud sequence is very small.

[0083] Figure 9 This is a visual diagram illustrating the corresponding vectors of the web crawler user-end sequence and the cloud sequence in an embodiment of the present invention. For example... Figure 9 As shown, after dimensionality reduction of the vectors corresponding to the web crawler user sequence and the cloud sequence, it can be seen that the distance between the vector 901 corresponding to the web crawler user sequence and the vector 902 corresponding to the cloud sequence is very large.

[0084] like Figure 8 and Figure 9 As shown, in this embodiment of the invention, inputting the target user's end sequence and cloud sequence into the target language processing model can make the vector distance between the end sequence and cloud sequence of a normal user significantly smaller than that between the end sequence and cloud sequence of a web crawler user.

[0085] This invention, through its embodiments, inputs the endpoint sequence and cloud sequence corresponding to the target user into a pre-trained target language processing model to determine whether the endpoint sequence and cloud sequence match. This determines whether the target user's access operation has been forged or tampered with, and thus whether the target user is a web crawler. This invention improves the accuracy and coverage of crawler detection by comparing and detecting the endpoint sequence and cloud sequence corresponding to the target user's access operation one by one.

[0086] Figure 10 This is a data flow diagram of data processing according to an embodiment of the present invention, such as... Figure 10 As shown, the data processing flow is as follows:

[0087] In step S1001, the obtained end sequence and cloud sequence are input into the word segmentation tool to obtain the word segments corresponding to each sequence.

[0088] Specifically, the system receives terminal log data uploaded from terminal devices and extracts terminal sequences from the terminal log data. Then, it extracts cloud sequences from cloud log data. The extracted terminal sequences and cloud sequences are then input into a word segmentation tool to obtain the corresponding word segments.

[0089] In step S1002, the word segmentation is input into the vector encoder to determine the word vector corresponding to each word segmentation.

[0090] Specifically, the word vector corresponding to each word segment consists of three parts: token embeddings, position embeddings, and segment embeddings. The token embeddings represent the vector corresponding to that word. The position embeddings represent the positional information of that word. The segment embeddings indicate which sequence the word belongs to.

[0091] In step S1003, the obtained word vectors are input into the Transformer to obtain the corresponding feature vectors.

[0092] Specifically, a Transformer consists of an encoder and a decoder. Generally, the encoder contains multiple sub-encoders, and similarly, the decoder contains multiple sub-decoders. Each encoder has the same structure but different parameters. The input to each sub-encoder is the output of the previous sub-encoder, and the encoder's output can be set to a vector with the same dimensions as the input vector. The input to each sub-decoder includes not only the output of its predecessor but also the output of the entire encoding part. After processing by the Transformer, the feature vectors corresponding to the word vectors are obtained.

[0093] In step S1004, the feature vector is input into the fully connected layer to obtain the matching degree between the end sequence and the cloud sequence.

[0094] Specifically, in the fully connected layer, the feature vectors of the end sequence and the cloud sequence are projected onto a higher-dimensional vector, with each dimension corresponding to a unique word score. Then, the scores of the entire end sequence and cloud sequence are converted into matching probabilities, i.e., matching degrees. Finally, the target user's classification is determined based on the matching degree. For example, a matching degree threshold is pre-set. If the matching degree between the target user's end sequence and cloud sequence is higher than this threshold, the target user is identified as a normal user; otherwise, the target user is identified as a web crawler user.

[0095] It is worth noting that step S1001 is not mandatory. It can be determined based on the limitations of different BERT models on the input data or the purpose of use. Alternatively, the end sequence and cloud sequence can be directly input into the vector encoder to obtain the vector corresponding to the entire sequence.

[0096] This invention addresses this issue by vectorizing the end sequence and cloud sequence corresponding to the target user into word vector groups, using a Transformer to extract features from each word vector group, and finally projecting each feature vector into a higher-dimensional vector using a fully connected layer. The distance between the corresponding vectors of the end sequence and cloud sequence in the vector space is then obtained, thereby determining the matching degree between the end sequence and cloud sequence. This invention improves the accuracy and coverage of crawler detection by comparing and detecting the end sequence and cloud sequence corresponding to the target user's access operation one by one.

[0097] Figure 11 This is a schematic diagram of a data processing apparatus according to an embodiment of the present invention. Figure 11 As shown, the data processing apparatus of this embodiment includes:

[0098] The acquisition module 1101 is used to acquire the terminal sequence and cloud sequence corresponding to the target user. The terminal sequence is the action sequence of the target user making a request from the terminal device to the cloud, and the cloud sequence is the action sequence of the cloud responding to the target user's request and returning corresponding data.

[0099] The first determining module 1102 is used to input the end sequence and the cloud sequence into a pre-trained target language processing model to determine the matching degree between the end sequence and the cloud sequence;

[0100] The second determining module 1103 is used to determine the target user category based on the matching degree, the category including web crawler users and normal users.

[0101] This invention, through its embodiments, inputs the endpoint sequence and cloud sequence corresponding to the target user into a pre-trained target language processing model to determine whether the endpoint sequence and cloud sequence match. This determines whether the target user's access operation has been forged or tampered with, and thus whether the target user is a web crawler. This invention improves the accuracy and coverage of crawler detection by comparing and detecting the endpoint sequence and cloud sequence corresponding to the target user's access operation one by one.

[0102] Figure 12 This is a schematic diagram of an electronic device according to an embodiment of the present invention. (For example...) Figure 12 As shown, Figure 12The illustrated electronic device is a general address lookup device, comprising a general computer hardware architecture, including at least a processor 1201 and a memory 1202. The processor 1201 and memory 1202 are connected via a bus 1203. The memory 1202 is adapted to store instructions or programs executable by the processor 1201. The processor 1201 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 1201 executes the instructions stored in the memory 1202 to perform the method flow of the embodiments described above or to implement the system module architecture described above, thereby processing data and controlling other devices. The bus 1203 connects the aforementioned components together, and also connects the aforementioned components to a display controller 1204, a display device, and an input / output (I / O) device 1205. The input / output (I / O) device 1205 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 1205 is connected to the system via the input / output (I / O) controller 1206.

[0103] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus (devices), or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0104] This application is described with reference to flowchart illustrations of methods, apparatus (devices), and computer program products according to embodiments of this application. It should be understood that each step in the flowchart can be implemented by computer program instructions.

[0105] These computer program instructions may be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction means, the implementation process of which is described in the instruction means. Figure 1 The function specified in one or more processes.

[0106] These computer program instructions may also be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce instructions for implementing processes. Figure 1 A device for a function specified in one or more processes.

[0107] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.

[0108] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program specifying the relevant hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the terminal sequence and cloud sequence corresponding to the target user. The terminal sequence is the action sequence of the target user making a request from the terminal device to the cloud. The cloud sequence is the action sequence of the cloud responding to the target user's request and returning corresponding data. The end sequence and the cloud sequence are input into a pre-trained target language processing model to determine the matching degree between the end sequence and the cloud sequence. The target user category is determined based on the matching degree, and the category includes web crawler users and normal users; The step of obtaining the terminal sequence and cloud sequence corresponding to the target user includes: Acquire the target user's terminal log data and cloud log data respectively; Based on the terminal log data and the cloud log data, multiple operation batches are determined. Each operation batch includes multiple terminal sequences and cloud sequences of the target user. Each operation batch includes a set of consecutive operations. Data processing is performed on the basis of the operation batches to determine the target user classification.

2. The method according to claim 1, characterized in that, The acquisition of the terminal sequence and cloud sequence corresponding to the target user includes: Acquire the target user's terminal log data and cloud log data within a preset historical time period, respectively; Based on the terminal log data and the cloud log data, the terminal sequence and cloud sequence of the target user are determined.

3. The method according to claim 1 or 2, characterized in that, The terminal log data is obtained through terminal tracking, and the cloud log data is obtained through cloud tracking.

4. The method according to claim 1, characterized in that, The target language processing model is trained using the following method: Obtain the terminal sequence and cloud sequence for multiple users; Based on the terminal sequences and cloud sequences corresponding to the multiple users, the model training dataset is determined; The target language processing model is trained based on the model training dataset.

5. The method according to claim 4, characterized in that, The step of determining the model training dataset based on the terminal sequences and cloud sequences corresponding to the multiple users includes: The training dataset for the model is determined by using the terminal sequence and corresponding cloud sequence of the same user as positive samples, and the terminal sequences and cloud sequences belonging to different users as negative samples.

6. The method according to claim 4, characterized in that, The step of determining the model training dataset based on the terminal sequences and cloud sequences corresponding to the multiple users includes: Disorder the cloud sequence and / or the terminal sequence according to a preset ratio; In response to the cloud sequence matching the corresponding end sequence, the labels of the cloud sequence and the end sequence are set to 0; In response to a mismatch between the cloud sequence and the corresponding end sequence, the labels of the cloud sequence and the end sequence are set to 1; Determine the training dataset for the model.

7. A data processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire the terminal sequence and cloud sequence corresponding to the target user. The terminal sequence is the action sequence in which the target user makes a request from the terminal device to the cloud, and the cloud sequence is the action sequence in which the cloud responds to the target user's request and returns corresponding data. The first determining module is used to input the end sequence and the cloud sequence into a pre-trained target language processing model to determine the matching degree between the end sequence and the cloud sequence; The second determining module is used to determine the target user category based on the matching degree, the category including web crawler users and normal users; The acquisition module is specifically used to acquire the terminal log data and cloud log data of the target user respectively. Based on the terminal log data and the cloud log data, multiple sets of operation batches are determined. Each set of operation batches includes multiple sets of terminal sequences and cloud sequences of the target user. Each set of operation batches includes a set of continuous operations. Data processing is performed on the basis of the operation batches to determine the classification of the target user.

8. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-6.