E-commerce anti-spider counter-crawling method, device and equipment based on GNN learning

By constructing the initial heterogeneous graph and multi-level graph convolution operation of graph neural networks, combining cross-level fusion attention and Transformer modules, the graph neural network is updated in real time and the reward function is designed, which solves the problem of insufficient perception of e-commerce anti-crawler environment in the existing technology, real-time perception and dynamic optimization of e-commerce website environment, and improves the adaptability and stability of crawler systems.

CN120086426BActive Publication Date: 2025-07-22XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510577955.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-22
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

When the existing network crawling technology faces the multi-level and multi-dimensional anti-crawling mechanism of e-commerce platforms, it is difficult to achieve real-time perception and strategy adjustment of environmental changes, resulting in insufficient adaptability and intelligence.

Method used

The initial heterogeneous graph of the graph neural network is constructed, and multi-scale embedded vectors are generated through multi-level parallel graph convolution operations. Combined with cross-level fusion attention and hierarchical gating mechanisms, the Transformer module is introduced for global interactive modeling, the graph neural network is updated in real time, and the reward function is used for crawling and evaluation, so as to realize real-time perception and dynamic optimization of the e-commerce website environment.

Benefits of technology

It improves the adaptability and stability of the crawler system in complex environments, enhances the anti-blocking capability, and is suitable for intelligent e-commerce data collection and market analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086426B_ABST
    Figure CN120086426B_ABST
Patent Text Reader

Abstract

The anti-crawler countermeasure crawling method, device and equipment for e-commerce based on GNN learning provided by the present invention relate to the technical field of data crawling. According to the historical data of the e-commerce website, the present invention constructs an initial heterogeneous graph, and uses multi-level parallel graph convolution operations to generate multi-scale embedding vectors of nodes and edges, constructing a multi-level graph neural network; introducing cross-level fusion attention and hierarchical gating mechanisms to perform intra-layer self-attention processing on each layer of the multi-level graph neural network and dynamically adjust the weight allocation of different levels; then performing cross-layer fusion of the multi-scale embedding vectors, and introducing a Transformer module to obtain the final embedding vector; updating the multi-level graph neural network by real-time obtaining website feedback information and generating the current state space embedding vector; finally, combining the reward function to evaluate and analyze the crawling results of the e-commerce website so as to adjust the website anti-crawler strategy in real time. The present invention can realize the real-time perception and dynamic optimization of the environmental state of the e-commerce website.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and data crawling, and more particularly, to an anti-crawler counter-crawling method, device, and equipment for e-commerce based on GNN learning. Background Art

[0002] With the rapid development of e-commerce, a large amount of product information, user reviews, price data, etc. have become important bases for enterprises' market analysis, formulation of competitive strategies, and optimization of user experience. Therefore, various web crawler technologies have emerged to automatically obtain a large amount of data from e-commerce platforms.

[0003] However, e-commerce platforms have deployed multi-level and multi-dimensional anti-crawler mechanisms to protect their own data resources and prevent malicious crawling. These mechanisms include:

[0004] (1) Request frequency limit: By detecting the access frequency, high-frequency requests within a short period are restricted, often responded with HTTP status code 429 (Too Many Requests) or 403 (Forbidden). (2) IP address banning: IP addresses of abnormal access sources are banned, and a blacklist mechanism is used to block high-frequency or abnormal requests. (3) Behavior detection and CAPTCHA verification: By analyzing user behavior characteristics (such as mouse trajectory, page scrolling, etc.) and CAPTCHA verification, etc., to distinguish crawlers from normal users. (4) Dynamic content loading and encrypted transmission: By means of Ajax asynchronous loading, Token verification, etc., the crawling difficulty is increased.

[0005] Existing web crawler technologies usually adopt static rules or simple reinforcement learning algorithms for countering, and it is difficult to adapt to complex and changeable anti-crawler strategies. In the static rule method, the crawler system lacks adaptability and intelligence. In the existing reinforcement learning frameworks, the update of the environmental state often lags behind, and it is difficult to achieve real-time perception of environmental changes and policy adjustment.

[0006] In view of this, the applicant proposes this application. Summary of the Invention

[0007] The present invention aims to provide an anti-crawler counter-crawling method, device, and storage medium for e-commerce based on GNN learning, so as to solve the problems that the environmental modeling in the field of e-commerce anti-crawlers by existing reinforcement learning methods is not accurate enough, and it is difficult to achieve real-time perception of environmental changes and policy adjustment.

[0008] To solve the above technical problems, the present invention is realized through the following technical solutions:

[0009] An anti-crawler counter-crawling method for e-commerce based on GNN learning, comprising:

[0010] S1. Construct an initial heterogeneous graph of a graph neural network based on the historical log data of an e-commerce website;

[0011] S2. Adopt multi-level parallel graph convolution operations based on the initial heterogeneous graph to generate multi-scale embedding vectors of nodes and edges, and construct a multi-level graph neural network;

[0012] S3. Introduce cross-level fusion attention and hierarchical gating mechanisms to perform intra-layer self-attention processing on each layer of the multi-level graph neural network and dynamically adjust the weight distribution between different levels to obtain multi-scale self-attention embedding vectors;

[0013] S4. Perform cross-layer fusion on the multi-scale self-attention embedding vectors and introduce a Transformer module for global interaction modeling to obtain final embedding vectors;

[0014] S5. Real-time obtain website feedback information to update the multi-level graph neural network, and generate a current state space embedding vector in combination with the final embedding vector;

[0015] S6. Crawl and evaluate e-commerce website information according to the current state space embedding vector in combination with a reward function to adjust the website anti-crawling strategy in real time.

[0016] Preferably, the nodes of the initial heterogeneous graph include website page nodes, IP address nodes, and request type nodes;

[0017] The edges of the initial heterogeneous graph include access relationship edges, warning trigger edges, and ban record edges;

[0018] Among them, the access relationship edge means that the relationship between nodes is a request type; the warning trigger edge means that the relationship between nodes is that one party's access to the other party triggers a warning; the ban record edge means that between two nodes, at least one party is in a banned state.

[0019] Preferably, the S3 is specifically:

[0020] After constructing a multi-level graph neural network, based on the embedding representations of different levels, use the standard self-attention mechanism to model the internal information of each layer, and introduce a hierarchical importance scoring function to calculate the importance weights of each layer to dynamically adjust the contributions of each layer. The expression is:

[0021] ;

[0022] ;

[0023] Among them, represents the embedding representation of the th layer; represents The embedded representation after standard self-attention, i.e., the self-attention embedding vector of the th layer; represents the standard self-attention function;

[0024] represents the importance weight of the th layer; represents the th layer of the self-attention embedding vector; M represents the total number of layers of the graph neural network; represents the importance scoring function.

[0025] Preferably, the importance scoring function uses a fully connected layer plus a non-linear activation function to process the input vector, and the expression is:

[0026] ;

[0027] where and are trainable parameters; represents the transpose; represents taking the average of the nodes; is the activation function.

[0028] Preferably, the S4 is specifically:

[0029] Performing cross-layer fusion on the multi-scale self-attention embedding vector, and the expression is:

[0030] ;

[0031] where is the representation after cross-layer fusion; M represents the total number of layers of the graph neural network; represents the importance weight of the th layer; represents the th layer of the self-attention embedding vector;

[0032] Introducing a Transformer module to perform global interaction modeling on the representation after cross-layer fusion to obtain the final embedding vector, and the expression is:

[0033] ;

[0034] ;

[0035] where is the fused representation after passing through the Transformer module, i.e., the final embedding vector; and are the final embedding vectors of the nodes and edges respectively; Represents the operation of the Transformer module, including multi-head cross-attention and a feed-forward neural network, to automatically identify and strengthen the nodes and edges that play a key role in multi-scale information.

[0036] Preferably, it further includes introducing a time decay weighting and incremental update mechanism to update the multi-level graph neural network, specifically:

[0037] First, introduce a time decay function , and the expression is:

[0038] ;

[0039] Among them, is the time decay function; T is the current global time; λ is the decay rate; t is the current moment;

[0040] Update the features of each node in the local area of the multi-level graph neural network in an incremental update manner according to the time decay function, and the expression is:

[0041] ;

[0042] Among them, represents the node feature at the current moment t; represents the node feature at the previous moment; represents the increment of the node feature at the current moment;

[0043] Then, use the local graph update function LocalUpdate to perform local update on the multi-level graph neural network, and the formula is:

[0044] ;

[0045] Among them, represents the multi-level graph neural network of the local area L at the current moment t; represents all the feature changes and new graph neural network structure information caused by the event in the local area L;

[0046] After local update, use a fusion function to splice and fuse the local update result with the non-updated area to generate the globally latest graph , and the expression is:

[0047] ;

[0048] Among them, represents the splicing and fusion operation.

[0049] Preferably, the reward function is:

[0050] ;

[0051] Wherein, represents the reward for the current state; represents the current state, obtained according to the current state space embedding vector; represents the action taken in the current state; , , , , are set weights;

[0052] represents the reward or punishment given according to the status returned by the website page, rewarding for normal return and punishing for errors; represents the differential reward given according to the importance of the website page content, the higher the value of the page content, the higher the reward; represents the punitive constraint on the request frequency, giving corresponding punishments for both too fast and too slow request frequencies; represents the reward or punishment based on the status of the proxy website IP; represents rewarding behaviors that conform to human characteristics and punishing crawling behaviors with obvious mechanicalness.

[0053] Preferably, it further includes dynamically adjusting the weight parameters of the reward function based on the current state reward and environmental feedback information, and the expression is:

[0054] ;

[0055] Wherein, is the set of weight parameters at the current time t; represents the dynamically adjusted learning rate; represents the derivative symbol; represents the reward for the current state; represents the current state, obtained according to the current state space embedding vector; represents the action taken in the current state.

[0056] The present invention also provides an anti-crawler counter-crawling device based on GNN learning, including:

[0057] An initial heterogeneous graph unit, configured to construct an initial heterogeneous graph of a graph neural network according to historical log data of an e-commerce website;

[0058] A multi-level graph neural network unit, configured to generate multi-scale embedding vectors of nodes and edges by using multi-level parallel graph convolution operations according to the initial heterogeneous graph, and construct a multi-level graph neural network;

[0059] An attention adjustment unit, which is used to introduce cross-level fusion attention and a hierarchical gating mechanism, perform intra-layer self-attention processing on each layer of the multi-level graph neural network, and dynamically adjust the weight distribution between different levels to obtain multi-scale self-attention embedding vectors;

[0060] A final embedding vector unit, which is used to perform cross-layer fusion on the multi-scale self-attention embedding vectors, and introduce a Transformer module for global interaction modeling to obtain final embedding vectors;

[0061] A current state generation unit, which is used to obtain website feedback information in real time to update the multi-level graph neural network, and generate a current state space embedding vector in combination with the final embedding vector;

[0062] An evaluation and decision-making unit, which is used to crawl and evaluate e-commerce website information according to the current state space embedding vector in combination with a reward function, so as to adjust the website anti-crawling strategy in real time.

[0063] The present invention also provides an e-commerce anti-crawler anti-crawling evaluation device based on GNN learning, including a processor and a memory. The memory stores a computer program, and the computer program can be executed by the processor to implement an e-commerce anti-crawler anti-crawling method based on GNN learning as described above.

[0064] The present invention also provides a computer-readable storage medium. A computer-readable instruction is stored on the computer-readable storage medium. When the computer-readable instruction is executed by a processor of a device where the computer-readable storage medium is located, an e-commerce anti-crawler anti-crawling method based on GNN learning as described above is implemented.

[0065] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0066] The present invention constructs a graph neural network model to depict the multi-dimensional feature relationships (such as page relationships, IP addresses, request types, etc.) in the e-commerce website anti-crawler environment, and combines an adaptive multi-level graph neural network with a reward function reinforcement learning strategy to realize real-time perception and dynamic optimization of the e-commerce website environment state.

[0067] The present invention not only pays attention to the attention weights inside each layer of the graph neural network, but also designs a cross-level fusion module, so that the information between different levels can be complementary to strengthen the representation of key features.

[0068] Based on a dynamic hierarchical gating mechanism and combined with a hierarchical importance scoring function, the present invention adaptively adjusts the contribution of each layer of the graph neural network to the overall representation, so as to more sensitively capture the dynamic changes of key nodes and relationships in the environment.

[0069] The entire process of the present invention forms a formal and trainable representation learning framework through the combination of intra-layer self-attention, importance scoring, weighted fusion, and cross-layer Transformer modules, providing a high-quality and multi-scale environmental representation for subsequent decision-making.

[0070] The present invention comprehensively evaluates website feedback information (such as status codes, content value, request rhythm, proxy usage, and user behavior), designs a hierarchical reward function, and improves the policy generalization ability and robustness based on meta-learning and online incremental update mechanisms to optimize and enhance the crawling effect of the model in different e-commerce website environments.

[0071] The present invention significantly improves the adaptability, stability, and anti-blocking ability of the crawler system in complex environments, and is applicable to scenarios such as intelligent collection of e-commerce data and market analysis. Brief Description of the Drawings

[0072] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0073] Figure 1 It is a schematic flowchart of a method for anti-crawler counter-crawling in e-commerce based on GNN learning provided for Embodiment 1.

[0074] Figure 2 It is a schematic diagram of a device for anti-crawler counter-crawling in e-commerce based on GNN learning provided for Embodiment 2.

[0075] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments. Detailed Embodiments

[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention claimed, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0077] Example 1

[0078] Example 1 of the present invention provides an anti - crawler counter - crawling method for e - commerce based on GNN learning, which can be implemented by an anti - crawler counter - crawling device for e - commerce based on GNN learning (hereinafter referred to as the crawling device), and in particular, is executed by one or more processors in the crawling device.

[0079] In this embodiment, the crawling device may be an electronic device equipped with a processor, and the processor has a computer program of the anti - crawler counter - crawling method for e - commerce based on GNN learning and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.

[0080] In this embodiment, a web crawler is a program or script that automatically crawls Internet information according to certain rules. Its core goal is to extract valuable data from web pages for subsequent analysis, storage, or display, and plays a key role in fields such as data collection, information aggregation, search engine construction, and business intelligence analysis. The working principle of the web crawler is as follows:

[0081] (1) Initialize the seed URL set: The crawler first starts from one or more initial URLs (seed URLs), which are usually the home page of a certain website or links to specific pages.

[0082] (2) Send an HTTP request to obtain the web page content: The crawler simulates the behavior of a browser and sends an HTTP request to the server corresponding to these URLs to obtain the source code of the web page in formats such as HTML, XML, JSON, etc.

[0083] (3) Parse the web page content: Use a parsing library (such as BeautifulSoup, lxml, PyQuery, etc. in Python) or regular expressions to extract the required data from the web page source code, such as text, image links, URLs of other web pages, etc.

[0084] (4) Store the data: Store the extracted data in a predetermined format (such as CSV, JSON, database table, etc.) for subsequent processing and analysis.

[0085] Process new URLs: For the new URLs extracted from the current web page, the crawler adds them to the URL queue to be crawled and decides the next URL to crawl according to a certain strategy (such as breadth - first, depth - first, priority queue, agent decision, etc.).

[0086] (5) Loop execution: Repeat the above steps until the URL queue is empty or other termination conditions are met (such as reaching the maximum crawling depth, time limit, etc.).

[0087] As Figure 1 shown, an anti-crawler countermeasure crawling method for e-commerce based on GNN learning includes steps S1 to S6.

[0088] S1. Construct an initial heterogeneous graph of a graph neural network based on the historical log data of the e-commerce website.

[0089] In this embodiment, according to the historical access log records of the e-commerce website, the graph structure is initialized .

[0090] The node set V of the initial heterogeneous graph includes various heterogeneous nodes such as the website home page, product detail page, shopping cart page, IP address, request type, GET request, and POST request, reflecting the diverse characteristics of the e-commerce website crawling environment.

[0091] The edge set E of the initial heterogeneous graph includes access relationship edges, warning trigger edges, and ban record edges, etc., describing the interaction and association relationships between nodes. Among them, the access relationship edge means that the relationship between nodes is the request type; the warning trigger edge means that the relationship between nodes is that one party's access to the other party triggers a warning; the ban record edge means that between two nodes, at least one party is in a banned state.

[0092] As shown in Table 1 of the historical access log records, the node set V of the initialized graph structure includes:

[0093] (1) Website page nodes, such as "Home Page", "Product Detail Page", "Shopping Cart Page";

[0094] (2) IP address nodes, such as : IP "192.168.1.100", : IP "192.168.1.101";

[0095] (3) Request type nodes, such as : GET request, : POST request.

[0096] According to the interaction relationships in the log records, construct the edge set E, mainly including the following categories:

[0097] (1) Access relationship edge: Such as edge : From the IP node from IP node (192.168.1.100) to page node (Home Page), and associate with the request type node, (GET); corresponding time record .

[0098] Edge : from IP node (192.168.1.101) to page node (Home Page), associate with the request type node (GET); corresponding record .

[0099] Edge : from IP node to page node (Product Detail Page), associate with the request type node (GET); corresponding record .

[0100] (2) Warning trigger edge: such as edge : For the time record , IP node (192.168.1.101)'s repeated access to the page node (Home Page) triggers a warning, and a warning trigger edge is established.

[0101] This edge can also be directly established between the IP node and the page node , and marked as "warning trigger" in the edge attribute.

[0102] (3) Banning record edge: If subsequent data indicates that a certain IP is banned, a banning record edge can be established between the corresponding IP node and the "banning status" node. There is no banning record here and it is not shown in the example.

[0103] Table 1: Historical access log records of an e-commerce website

[0104]

[0105] Initialize the graph structure according to the historical data in Table 1, integrate the above nodes and edges to form an initial heterogeneous graph . This graph structure not only completely expresses the interaction relationship between pages, IPs and request types, but also provides a clear data basis for subsequent feature embedding learning and reinforcement learning strategy design using graph neural networks.

[0106] S2. According to the initial heterogeneous graph, adopt multi-level parallel graph convolution operations to generate multi-scale embedding vectors of nodes and edges, and construct a multi-level graph neural network.

[0107] In this embodiment, an adaptive multi-scale graph neural network (Multi-Scale GNN) is used to construct a multi-scale graph neural network, as shown in the following example:

[0108] (1) Construct a first-level graph: The page jump relationship is used to capture the macroscopic page interaction pattern, which is expressed as , where represents the embedding representation of the th layer. For example, represents the embedding representation of the th layer; represents the number of nodes in the th layer; d is the embedding dimension; represents a real number.

[0109] For example, if a user often enters the "Product Detail Page" ( ) from the "Home Page" ( ), then an edge from to is constructed in the first-level graph, and the weight of the edge (hereinafter referred to as the edge weight) can be set to the jump frequency (for example, the weight is 2, indicating 2 jump records).

[0110] Example: Edge : represents from node to node (weight set to 2); edge : represents from node to (weight set to 1).

[0111] (2) Construct a second-level graph: A fine-grained feature graph such as DOM nodes and request types inside the page is used to express the internal structure of the page in detail, which is expressed as .

[0112] For example, the "Home Page" is divided into sub-nodes such as "Header", "Navigation Bar", and "Content Area", and the request type or component call relationship is recorded.

[0113] Example is as follows: For the "Home Page" ( ), sub-nodes (Header), (Content Area) are constructed, and a connection edge of "Header-Content Area" is established inside . At the same time, information such as the GET request ( ) is embedded into this layer to express the influence of the request inside the page.

[0114] (3)Construct a three - level graph: a global environment information graph of IP addresses, proxy pools, user behaviors, etc., comprehensively modeling user behaviors and environmental characteristics, expressed as 。

[0115] For example, if the IP "192.168.1.101" ( ) often triggers warnings, in the three - level graph, a strong connection can be established between it and the "abnormal behavior" node, and at the same time, a proxy pool node (such as "Proxy A") is introduced to associate with this IP.

[0116] The example is as follows: In addition to the original IP node and , a new global node is introduced to represent "high - risk behavior", and an edge is constructed: connecting and , and the edge weight reflects the risk level (for example, the weight is 0.8).

[0117] Through the constructed three - level graph, graph convolution operations are used in each layer to calculate the embedding vectors in parallel:

[0118] For example, to calculate the embedding vector of a node: Assume that after graph convolution, the "home page" node obtains a 128 - dimensional embedding vector , for example: = [0.12, −0.05, 0.33,..., 0.07].

[0119] To calculate the edge embedding vector: Similarly, the access relationship edge (for example: and , request type GET) obtains a 64 - dimensional embedding vector .

[0120] S3. Introduce cross - level fusion attention and hierarchical gating mechanisms to perform intra - layer self - attention processing on each layer of the multi - level graph neural network and dynamically adjust the weight distribution between different levels to obtain multi - scale self - attention embedding vectors.

[0121] Specifically, to fully explore the importance differences between multi - level information, a Transformer module is used to perform attention weighting on the embeddings of each layer. After constructing the multi - level graph neural network, based on the embedding representations of different levels, a standard self - attention mechanism is used to model the internal information of each layer, and a hierarchical importance scoring function is introduced to calculate the scalar importance weights of each layer to dynamically adjust the contributions of each layer. The expression is:

[0122] ;

[0123] ;

[0124] Among them, represents the embedding representation of the -th layer, such as = 1, 2, 3; represents the embedding representation after performing standard self-attention, that is, the self-attention embedding vector of the -th layer; represents the standard self-attention function, which aggregates information for each node with its neighbors within this layer;

[0125] represents the importance weight of the -th layer; represents the -th layer's self-attention embedding vector; M represents the total number of layers of the graph neural network; represents the importance scoring function.

[0126] In this embodiment, the importance scoring function uses a fully connected layer plus a non-linear activation to process the input vector, and the expression is:

[0127] ;

[0128] Among them, and are trainable parameters; represents transpose; represents taking the average of nodes; is an activation function (such as using activation functions such as ReLU, Sigmoid, etc.).

[0129] For example, for a certain website detection, if the abnormal jump behavior of the "product details page" (node ) arouses suspicion, the Transformer mechanism can dynamically increase the embedding weights related to , and at the same time reduce the weights of more regular pages (such as the "home page").

[0130] The original embedding vector of node is , and after being processed by the attention mechanism, a new self-attention embedding vector is generated, and abnormal features may be amplified so that the subsequent model can capture risk information more accurately.

[0131] Of course, other methods can also be used to score the importance of each layer of the graph structure, which is not limited here.

[0132] S4. Cross-layer fuse the multi-scale self-attention embedding vectors and introduce a Transformer module for global interaction modeling to obtain the final embedding vector.

[0133] Specifically, cross-layer fuse the multi-scale embedding vectors after in-layer self-attention processing. The expression is:

[0134] ;

[0135] where is the representation after cross-layer fusion; M represents the total number of layers of the graph neural network; represents the scalar importance weight of the -th layer; represents the embedding representation after standard self-attention, that is, the self-attention embedding vector;

[0136] To capture the complex interactions between layers, introduce a Transformer module to perform global interaction modeling on the representation after cross-layer fusion to obtain the final embedding vector. The expression is:

[0137] ;

[0138] where is the fused representation after passing through the Transformer module, that is, the final embedding vector; represents the Transformer module operation, including multi-head cross-attention and a feed-forward neural network, to automatically identify and strengthen the nodes and edges that play a key role in multi-scale information.

[0139] The updated final embedding vector will be used as the input for subsequent tasks (such as reinforcement learning, anti-crawling policy decision-making). Its formal representation is:

[0140] ;

[0141] where and are the final embedding vectors of nodes and edges respectively.

[0142] According to the final embedding vector obtain the state feature vector , which is used as the input for subsequent reinforcement learning. Specifically, combine the final embedding vectors of nodes and edges after being adjusted by multi-level parallel graph convolution and attention mechanism to form the state feature vector of the reinforcement learning model at time t.

[0143] For example, if the embedding dimension of the processed node is still 128 and the edge embedding is 64, then the final state feature vector It is a 192-dimensional vector and can be directly used as the input of the reinforcement learning model RL for subsequent decision-making and reward adjustment.

[0144] Through the above examples, the present invention realizes steps such as multi-level graph construction, parallel graph convolution embedding, attention mechanism optimization, and reinforcement learning state construction. The embodiments establish graph structures from three perspectives: macroscopic page jumps, page internal details, and the global environment; use the multi-scale graph neural network Multi-Scale GNN to calculate the feature vectors of nodes and edges in each layer simultaneously; dynamically adjust the information weights of each layer through the combination of Transformer-GNN (graph neural network combined with Transformer); then use the processed embedding vectors as the input of the reinforcement learning model RL to assist in realizing real-time anti-crawling decisions for the website. This combined scheme of layering, parallelism, and attention mechanism can capture multi-scale information and key features more accurately, providing a more accurate and dynamic environmental representation for the anti-crawler system.

[0145] The solution of this embodiment not only improves the representation ability of key nodes and relationships but also provides a more flexible and refined dynamic weight adjustment mechanism for the system when facing complex and changeable crawler environments, thus more effectively supporting the optimization and decision-making of anti-crawling strategies.

[0146] S5. Obtain website feedback information in real time to update the multi-level graph neural network, and generate the current state space embedding vector in combination with the final embedding vector.

[0147] In this embodiment, according to the new feedback information, through streaming data processing, local graph updates are dynamically triggered to generate the latest graph representation in real time.

[0148] For example, the real-time feedback information includes HTTP status codes, captcha triggers, ban records, dynamic loading status, etc. Assume that at time t, the received feedback information is:

[0149] HTTP status code: 403 (indicating access prohibited);

[0150] Captcha trigger: True (the captcha verification mechanism is triggered);

[0151] Ban record: True (this IP has been recorded as a violation and banned);

[0152] Dynamic loading status: "loading failed" (the dynamic content of the page fails to load normally);

[0153] Record this feedback as: {403, True, True, "loading failed"}.

[0154] Then, update the graph structure according to the real-time feedback information. In this embodiment, a time decay weighting and incremental update mechanism is introduced to update the multi-level graph neural network, specifically as follows:

[0155] Introduce a time decay function , and the expression is:

[0156] ;

[0157] where T is the current global time; λ is the decay rate, such as setting the default value to 0.85; t is the current moment;

[0158] Update each node feature in the local area of the multi-level graph neural network using an incremental update method according to the time decay function, and the expression is:

[0159] ;

[0160] where, represents the node feature at the current moment t; represents the node feature at the previous moment; represents the increment of the node feature at the current moment;

[0161] Use the local graph update function LocalUpdate to perform local updates on the multi-level graph neural network, and the formula is:

[0162] ;

[0163] where, represents the multi-level graph neural network of the local area L at the current moment t; represents all feature changes and new graph neural network structure information caused by the event in the local area L;

[0164] After local update, use a fusion function to splice and fuse the local update result with the unupdated area to generate the globally latest graph , and the expression is:

[0165] ;

[0166] where, represents the splicing and fusion operation.

[0167] Through the streaming data processing mechanism, the above update operations will be triggered in real time, and only the affected areas in the local graph (such as the local subgraph related to IP "192.168.1.101" and "Home Page") will be updated to ensure that the entire graph can reflect the latest environmental state in real time when new feedback is received.

[0168] This online real-time update mechanism can ensure that the graph neural network always performs feature extraction and subsequent decision-making based on the latest environmental state, providing the anti-crawler system with dynamic and accurate environmental modeling capabilities.

[0169] S6. Crawl and evaluate the e-commerce website information according to the current state space embedding vector combined with the reward function to adjust the anti-crawler decision in real time.

[0170] In this step, according to the real-time updated graph structure , generate the current state space embedding vector , so as to obtain the current state of the reward function .

[0171] In this embodiment, by designing a hierarchical reward mechanism, the crawler behavior can be comprehensively and accurately evaluated.

[0172] The reward function can be set as:

[0173] ;

[0174] Among them, , , , , are set weights, such as set as , , = 0.15, , , and the sum is 1; the weight parameter set .

[0175] represents the reward or punishment given according to the HTTP website page return status, positive reward for normal return, and negative punishment for error return.

[0176] For example: If the HTTP return status code is 200 (normal response), a positive reward is given; if an error status such as 403 or 404 is returned, a negative punishment is given. Example: If a certain request returns 403, then .

[0177] represents the differential reward given according to the importance of the website page content, the higher the value of the page content, the higher the reward.

[0178] For example, for an important product detail page, due to its high commercial value, if it can be obtained normally, a higher reward is given; otherwise, the reward is low or 0. Example: If the request target is a product detail page, but the content acquisition fails due to an exception, then make .

[0179] Represents a punitive constraint on the request frequency. Both too fast and too slow request frequencies trigger a penalty item, and corresponding penalties are given.

[0180] For example, a too fast request frequency (such as when the request interval < 200 ms) may be judged as a crawler behavior, and a too slow frequency (such as when the request interval > 5 s) affects the data update efficiency. If the system detects an abnormal request frequency (for example, extremely high), then .

[0181] Represents rewards or penalties based on the status of the proxy website IP;

[0182] For example, evaluate according to the status of the proxy IP. If the proxy works normally, it is rewarded; if it is abnormal or blocked, it is punished. Example: If the proxy IP used in the request has been blocked, then .

[0183] Represents rewarding behaviors that conform to human characteristics and punishing crawling behaviors with obvious mechanicalness.

[0184] For example, encourage access patterns that imitate normal human behaviors. If the behavior shows strong mechanicalness (such as fixed-time, non-fluctuating request intervals), then a penalty is given; on the contrary, if the behavior has reasonable human randomness, it is rewarded. Example: If it is detected that the request intervals are all fixed times (obviously mechanized), then .

[0185] Thus, the overall reward is calculated , if the overall reward is negative, it means that this request behavior has received a negative reward due to multiple abnormal performances, prompting the system to take adjustment measures in subsequent decisions (such as reducing the request frequency, changing the proxy, imitating an access strategy closer to human behavior).

[0186] In order to make the reward function more in line with the actual environmental feedback, the present invention adopts a dynamic adjustment strategy, and the weight parameters are updated in real time according to historical rewards and environmental feedback. The expression for dynamically adjusting the weight parameters is:

[0187] ;

[0188] Among them, is the set of weight parameters at the current time t; represents the learning rate of dynamic adjustment; represents the derivative symbol; represents the reward of the current state; represents the current state, obtained according to the current state space embedding vector; Indicates the actions taken in the current state.

[0189] Suppose that after several consecutive test operations, the system discovers (Proxy state) plays a key role in determining the overall crawler behavior. At this time, the system may increase the weight value through the gradient ascent strategy.

[0190] For example, if it is calculated that , and the learning rate , then the update formula is:

[0191] , making slightly decrease, further increasing the punishment for proxy state anomalies, so as to prompt the model to pay more attention to the health status of proxy IPs when making decisions.

[0192] This dynamic parameter adjustment mechanism makes the reward function adaptive, capable of automatically optimizing the decision-making process according to the continuously changing environmental feedback, thereby enhancing the robustness and intelligence of the anti-crawler system.

[0193] This embodiment realizes real-time state perception of e-commerce anti-crawlers through state embedding generation, hierarchical reward function design, and dynamic parameter optimization. Using the current latest graph to generate the current state vector provides rich multi-dimensional information for subsequent decision-making. By constructing a comprehensive reward model composed of HTTP status, page content, request frequency, proxy state, and user behavior, each dimension is quantitatively evaluated; based on the gradient adjustment strategy, the reward weights of each item are updated in real time, enabling the system to respond more sensitively to environmental changes and achieve the goal of accurate anti-crawling. This hierarchical reward mechanism can comprehensively and accurately evaluate crawler behavior, providing an effective basis for subsequent strategy adjustment and reinforcement learning decision-making.

[0194] Due to the constantly changing environment (for example, new crawler strategies, IP dynamic changes, page structure adjustments, etc.), it is necessary to periodically train the graph neural network.

[0195] For example, every other day or when detecting a drastic environmental change, the system automatically triggers the retraining of the graph neural network. The updated embedding vector makes the policy network more accurate when processing the latest environment.

[0196] For example, abnormal behaviors that may have been misjudged based on the old embedding can be corrected through the new embedding, improving the overall anti-crawling effect.

[0197] This embodiment can continuously and dynamically update the graph neural network through experience collection, policy update, meta-learning integration, fast fine-tuning and game optimization. In the state of the graph neural network , the agent selects actions according to the current policy , and obtain immediate rewards and new states, forming experience tuples and storing them in the replay pool. Using randomly sampled data from the experience replay pool, update the parameters of the reinforcement learning policy network. When training in multiple environments, the initial parameters of the general policy can be extracted so that the model can quickly adapt to the new environment. Based on the initial parameters, quickly adapt in the new environment, and further optimize the policy through the game mechanism; and periodically retrain the graph neural network to update the final embedding vector to ensure the real-time accuracy of environmental features. This comprehensive method of policy optimization and meta-learning enables the anti-crawler system to quickly respond and dynamically adjust decision-making strategies when facing a changing environment, improving the robustness and intelligence of the system.

[0198] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0199] (1) Cross-level information fusion: The present invention not only focuses on the attention weights within each layer, but also designs a cross-level fusion module, enabling the information between different layers to be complementary and jointly strengthening the representation of key features.

[0200] (2) Dynamic gating mechanism: The present invention can adaptively adjust the contribution of each layer to the overall representation by introducing a hierarchical importance scoring function, thereby more sensitively capturing the dynamic changes of key nodes and relationships in the environment.

[0201] (3) Formalized fusion model: The entire process forms a formalized and trainable representation learning framework through the combination of intra-layer self-attention, importance scoring, weighted fusion, and cross-layer Transformer modules, providing a high-quality and multi-scale environmental representation for subsequent decision-making.

[0202] (4) Adaptive reward mechanism: The present invention comprehensively evaluates the status code, content value, request rhythm, proxy usage, user behavior, etc., designs a hierarchical reward function, and improves the policy generalization ability and robustness based on meta-learning and online incremental update mechanisms, optimizing the crawling effect of e-commerce websites in different environments.

[0203] Embodiment 2

[0204] As Figure 2 shown, the second embodiment of the present invention also provides an e-commerce anti-crawler anti-crawling device based on GNN learning, including:

[0205] An initial heterogeneous graph unit for constructing an initial heterogeneous graph of a graph neural network according to the historical log data of an e-commerce website;

[0206] A multi-level graph neural network unit for generating multi-scale embedding vectors of nodes and edges by using multi-level parallel graph convolution operations according to the initial heterogeneous graph, and constructing a multi-level graph neural network;

[0207] An attention adjustment unit, which is used to introduce cross-level fusion attention and hierarchical gating mechanisms, perform intra-layer self-attention processing on each layer of the multi-level graph neural network, and dynamically adjust the weight distribution between different levels, so as to obtain multi-scale self-attention embedding vectors;

[0208] A final embedding vector unit, which is used to perform cross-layer fusion on the multi-scale self-attention embedding vectors, and introduce a Transformer module for global interaction modeling to obtain final embedding vectors;

[0209] A current state generation unit, which is used to obtain website feedback information in real time to update the multi-level graph neural network, and generate a current state space embedding vector in combination with the final embedding vector;

[0210] An evaluation and decision-making unit, which is used to crawl and evaluate e-commerce website information according to the current state space embedding vector in combination with a reward function, so as to adjust the website anti-crawling strategy in real time.

[0211] Embodiment III

[0212] The third embodiment of the present invention also provides an e-commerce anti-crawler counter-crawling evaluation device based on GNN learning, which includes a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement the e-commerce anti-crawler counter-crawling method based on GNN learning as described above.

[0213] Embodiment IV

[0214] The fourth embodiment of the present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, the e-commerce anti-crawler counter-crawling method based on GNN learning as described above is implemented.

[0215] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the part of the module, program segment, or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0216] In addition, each functional module in various embodiments of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0217] If the described functions are implemented in the form of software function modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0218] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.

[0219] It should be understood that the term "and / or" used herein is merely a description of an association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the sole existence of A, the simultaneous existence of A and B, and the sole existence of B. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0220] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0221] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by "first / second" can be interchanged appropriately so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0222] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An anti-spider counter-crawling method for e-commerce based on GNN learning, characterized in that Including: S1. Construct an initial heterogeneous graph of a graph neural network based on the historical log data of an e-commerce website; S2. Generate multi-scale embedding vectors of nodes and edges by using multi-level parallel graph convolution operations based on the initial heterogeneous graph, and construct a multi-level graph neural network; S3. Introduce cross-level fusion attention and hierarchical gating mechanisms to perform intra-layer self-attention processing on each layer of the multi-level graph neural network and dynamically adjust the weight distribution between different levels to obtain multi-scale self-attention embedding vectors; specifically: After constructing the multi-level graph neural network, based on the embedding representations of different levels, use the standard self-attention mechanism to model the internal information of each layer, and introduce a hierarchical importance scoring function to calculate the importance weights of each layer to dynamically adjust the contributions of each layer. The expression is: ; ; Among them, represents the embedding representation of the -th layer; represents the embedding representation after performing standard self-attention, that is, the self-attention embedding vector of the -th layer; represents the standard self-attention function; Indicates the importance weight of the layer; Indicates the self-attention embedding vector of the layer; M represents the total number of layers of the graph neural network; Indicates the importance scoring function; S4. Perform cross-layer fusion on the multi-scale self-attention embedding vectors, and introduce a Transformer module to perform global interaction modeling to obtain the final embedding vector; S5. Real-time obtain website feedback information to update the multi-level graph neural network, and generate the current state space embedding vector in combination with the final embedding vector; S6. Crawl and evaluate the information of the e-commerce website according to the current state space embedding vector in combination with the reward function to adjust the website anti-crawling strategy in real time; the reward function is: ; Among them, represents the reward for the current state; represents the current state, obtained according to the current state space embedding vector; represents the action taken in the current state; , , , , are the set weights; Indicates the rewards or punishments given according to the return status of the website page. Rewards are given for normal returns, and punishments are given for errors; Indicates the differential rewards given according to the importance of the website page content. The higher the value of the page content, the higher the reward; Indicates the punitive constraints imposed on the request frequency. Appropriate punishments are given for both too fast and too slow request frequencies; Indicates the rewards or punishments given according to the status of the proxy website IP; Indicates rewarding behaviors that conform to human characteristics and punishing crawler behaviors that are obviously mechanical; It also includes dynamically adjusting the weight parameters of the reward function based on the current state reward and environmental feedback information. The expression is: ; Among them, is the set of weight parameters at the current moment t; represents the dynamically adjusted learning rate; represents the derivative symbol; represents the reward of the current state; represents the current state, obtained according to the current state space embedding vector; represents the action taken in the current state.

2. The anti - web - crawler counter - crawling method for e - commerce based on GNN learning according to claim 1, wherein , The nodes of the initial heterogeneous graph include website page nodes, IP address nodes, and request type nodes; The edges of the initial heterogeneous graph include access relationship edges, warning trigger edges, and ban record edges; Among them, the access relationship edge means that the relationship between nodes is the request type; the warning trigger edge means that the relationship between nodes is that one party's access to the other party triggers a warning; the ban record edge means that at least one of the two nodes is in a banned state.

3. The anti - crawler counter - crawling method for e - commerce based on GNN learning according to claim 1, characterized in that , the importance scoring function uses a fully connected layer plus a non-linear activation function to process the input vector, and the expression is: ; Among them, and are trainable parameters; represents transpose; represents taking the average of nodes; is an activation function.

4. The anti - crawler counter - crawling method for e - commerce based on GNN learning according to claim 1, characterized in that , The specific content of S4 is: Perform cross-layer fusion on the multi-scale self-attention embedding vectors. The expression is: ; Among them, is the representation after cross-layer fusion; M represents the total number of layers of the graph neural network; represents the importance weight of the th layer; represents the self-attention embedding vector of the Introduce a Transformer module to perform global interaction modeling on the representation after cross-layer fusion to obtain the final embedding vector. The expression is: ; ; Among them, is the fused representation after the Transformer module, that is, the final embedding vector; , are the final embedding vectors of nodes and edges respectively; represents the operation of the Transformer module, including multi-head cross-attention and a feed-forward neural network, to automatically identify and strengthen the nodes and edges that play key roles in multi-scale information.

5. A method for anti-crawler confrontation and crawling prevention in e-commerce based on GNN learning according to claim 1, characterized in that , It also includes introducing a time decay weighting and incremental update mechanism to update the multi-level graph neural network. Specifically: First, introduce the time decay function , and the expression is: ; Among them, T is the current global time; λ is the decay rate; t is the current moment; Update the features of each node in the local area of the multi-level graph neural network in an incremental update manner according to the time decay function. The expression is: ; Among them, represents the node feature at the current moment t; represents the node feature at the previous moment; represents the increment of the node feature at the current moment; Then, use the local graph update function LocalUpdate to perform local update on the multi-level graph neural network. The formula is: ; Among them, represents the multi-level graph neural network of the local area L at the current time t; represents all the feature changes and new graph neural network structure information caused by the event within the local area L; After local update, a fusion function is used to splice and fuse the local update result with the non-updated area to generate the latest global map , and the expression is: ; Among them, represents a splicing and fusion operation.

6. An anti-crawler device for e-commerce anti-crawler based on GNN learning, characterized in that, Including: An initial heterogeneous graph unit for constructing an initial heterogeneous graph of a graph neural network based on the historical log data of an e-commerce website; A multi-level graph neural network unit for generating multi-scale embedding vectors of nodes and edges by using multi-level parallel graph convolution operations based on the initial heterogeneous graph and constructing a multi-level graph neural network; An attention adjustment unit, which is used to introduce cross-level fusion attention and hierarchical gating mechanism, perform intra-layer self-attention processing on each layer of the multi-level graph neural network, and dynamically adjust the weight distribution between different levels to obtain multi-scale self-attention embedding vectors; specifically: After constructing the multi-level graph neural network, based on the embedding representations of different levels, use the standard self-attention mechanism to model the internal information of each layer, and introduce a hierarchical importance scoring function to calculate the importance weights of each layer to dynamically adjust the contributions of each layer. The expression is: ; ; Among them, represents the embedding representation of the th layer; represents the embedding representation after performing standard self-attention, that is, the self-attention embedding vector of the th layer; represents the standard self-attention function; Indicates the importance weight of the Indicates the self-attention embedding vector of the layer; M represents the total number of layers of the graph neural network; Indicates the importance scoring function; A final embedding vector unit, which is used to perform cross-layer fusion on the multi-scale self-attention embedding vectors, and introduce a Transformer module for global interaction modeling to obtain the final embedding vector; A current state generation unit, which is used to obtain website feedback information in real time to update the multi-level graph neural network, and combine the final embedding vector to generate a current state space embedding vector; An evaluation and decision-making unit, which is used to crawl and evaluate e-commerce website information according to the current state space embedding vector in combination with a reward function to adjust the website anti-crawling strategy in real time; the reward function is: ; Among them, represents the reward of the current state; represents the current state, obtained according to the current state space embedding vector; represents the action taken in the current state; , , , , are the set weights; Indicates the rewards or punishments given according to the return status of the website page. Rewards are given for normal returns, and punishments are given for errors; Indicates the differentiated rewards given according to the importance of the website page content. The higher the value of the page content, the higher the reward; Indicates the punitive constraints on the request frequency. Appropriate punishments are given for both too fast and too slow request frequencies; Indicates the rewards or punishments given according to the status of the proxy website IP; Indicates rewarding behaviors that conform to human characteristics and punishing crawling behaviors with obvious mechanicalness; It also includes dynamically adjusting the weight parameters of the reward function based on the current state reward and environmental feedback information. The expression is: ; Among them, is the set of weight parameters at the current time t; represents the dynamically adjusted learning rate; represents the derivative symbol; represents the reward of the current state; represents the current state, obtained according to the current state space embedding vector; represents the action taken in the current state.

7. An anti-crawler countermeasure crawling device for e-commerce based on GNN learning, characterized in that, It includes a processor and a memory. The memory stores a computer program, and the computer program can be executed by the processor to implement a method for anti-crawling against e-commerce anti-crawlers based on GNN learning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Text information processing method and device, equipment, software program and storage medium

    CN116541517A

  • Clothing trend crawler improvement system based on near-end strategy optimization model

    CN119322879A