Anti-crawling application method and system based on ELK log collection

Through the anti-crawl method based on the ELK log collection system, combining traffic analysis and user portrait recognition crawler behavior, dynamically adjusting the anti-crawl strategy, the problem of high cost and easy to block the anti-crawl strategy of existing airlines is solved, and a lightweight and customized anti-crawl solution is realized, improving system performance and user experience.

CN120358060APending Publication Date: 2025-07-22TRAVELSKY TECHNOLOGY LIMITED
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510539080.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing anti-crawling strategies of airlines have problems such as high cost, complex deployment, easy to be mis-sealed and missealed. Especially small and medium-sized airlines are difficult to afford high-cost WAF solutions, and the existing technology is difficult to implement customized anti-crawling strategies, which affects user experience and system performance.

Method used

The anti-climbing method based on the ELK log collection system is adopted, log cleaning is performed through Logstash, and crawling behavior is identified by combining traffic analysis and user portraits, anti-climbing strategies are dynamically adjusted, and cache and alarm center are deployed at the gateway layer to realize a lightweight and customized anti-climbing strategy.

Benefits of technology

It realizes low-cost and efficient anti-crawling capabilities, reduces misblocking of normal users, improves system performance and user experience, reduces deployment and maintenance difficulties, and adapts to the customized needs of different airlines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358060A_ABST
    Figure CN120358060A_ABST
Patent Text Reader

Abstract

The invention provides an anti-crawling application method and system based on ELK log collection, and the method comprises the steps: 1, a request receiving step: carrying out the log cleaning in Logstash after a flow request forwarded by a gateway is received; 2, a flow analysis step: accurately analyzing the peak period and the valley period of the flow in the current time period according to the number of requests in the whole time period, and dynamically adjusting an anti-crawling strategy; 3, an anti-crawling filtering analysis step: intercepting User-Agent, Referer and cookie fields in the request header, comparing the fields with a standard request header in a database, analyzing whether an access path and a flow of the Referer header conform to reasonable user portrait information in the database, and determining whether all interception or partial interception is performed according to a time period; 4, an open customization step: the request is issued to an interface of the airline to enter different modules of the airline according to the type of the request; and 5, a timed pushing step: storing the identified problem IP in a cache, and regularly pushing crawler behavior IP data to an alarm center. Compared with anti-climbing tools in the market, the anti-climbing application method and system adopted by the invention are more flexible, low in deployment cost and easier to maintain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of network security technology, and particularly relates to an anti-crawler application method and system based on an ELK log collection system. Background Art

[0002] With the booming development of the aviation industry and the informatization of flight information, the demand of the general public for flight information and fare inquiries has increased accordingly when purchasing air tickets. People can not only query information on official websites, but also through third-party travel agents, search engines, and price comparison platforms. This growth in information query demand provides a breeding ground for crawler activities.

[0003] Appropriate crawler behavior promotes the sales of airlines. However, during peak hours, a large number of crawler behaviors will cause slow response of the airline's official website and poor user experience. There are even malicious crawler behaviors such as sending a large number of query requests or being used by criminals for ticket hoarding, ticket brushing, etc. These behaviors will not only significantly increase the checking ratio of airlines and thus the ticket sales cost, but also leak the airline's fare, flight schedule, seat inventory, etc., which are its core business assets. Once these data are easily crawled and misused by third parties, it may lead to market price chaos and affect the revenue management and pricing strategy of airlines.

[0004] The existing anti-crawler strategies of airlines and their disadvantages are as follows:

[0005] 1. Dynamic verification code: The existing dynamic verification code technology is an effective anti-crawler measure. However, attackers may try to bypass the verification code by means of OCR recognition, AI technology, etc. At the same time, if there are vulnerabilities in the implementation of the verification code, such as a predictable generation algorithm, it may also be targeted for attack. It cannot solely undertake the anti-crawler task and is suitable for combination with other security measures.

[0006] 2. Deploying a WAF firewall: The WAF is not specifically designed for anti-crawling, but it can set rules to limit the number of requests from a single IP address or IP segment within a specific time period, which can play a role in blocking or slowing down the crawling speed of crawlers. However, the WAF has obvious defects in anti-crawling:

[0007] The WAF can customize rules and monitoring strategies, but it cannot implement customized anti-crawler strategies for the additional needs of airlines.

[0008] In terms of the formulation of rules and strategies, it is easy to cause mis-blocking and missed-blocking situations. For high-end WAF solutions, the purchase, deployment, maintenance, and possible customized development all require high costs. In terms of servers, the cost of dedicated servers for WAF is too high, and most small and medium-sized airlines can only afford the cost of shared servers. The same set of WAF resources (such as processing power and bandwidth) will be allocated to multiple websites or applications, resulting in unstable performance and bandwidth during peak periods, affecting the user experience.

[0009] Common anti-crawling strategies for airlines also include IP address restrictions and rate control, login authentication and session management, etc. The functions are scattered. Small airlines will embed the functions into applications or gateways, resulting in increased gateway pressure, affecting the business processing speed, and increasing the complexity and coupling degree of the code. Summary of the Invention

[0010] In view of the deficiencies of the existing technology, the present application provides a processing method and system based on multi-dimensional fare search. The technical problems to be solved by the present application are realized through the following technical solutions.

[0011] The first aspect of the present application provides a processing method based on multi-dimensional fare search. The above method includes:

[0012] Step 1: Receive request step. After receiving the traffic request forwarded by the gateway, enter the log cleaning module of the ELK system and perform log cleaning in Logstash;

[0013] Step 2: Traffic analysis step. Accurately analyze the peak and trough periods of the current time period according to the number of requests in the whole time period, dynamically adjust the anti-crawling strategy, count and analyze the number of requests in different time periods, and identify the traffic patterns of each day, week, and month;

[0014] Step 3: Anti-crawling filtering analysis step. Intercept the User-Agent, Referer, and cookie fields in the request header, compare them with the standard request headers in the database, analyze whether the access path and process of the Referer header conform to the reasonable user profile information in the database, and determine whether to intercept all or part according to the time period;

[0015] Step 4: Open customization step. The requests sent to the airline's interface will enter different modules of the airline according to the type of the request;

[0016] Step 5: Timed push step. Store the identified problem IPs in the cache and push them to the alarm center for crawler behavior IP data at regular intervals.

[0017] Combined with the first aspect, the above log cleaning in Logstash further includes:

[0018] In Logstash, the log data is parsed and filtered to extract key information, which includes the source IP address, the User-Agent information in the request header, and the airline interface accessed by the request.

[0019] Combined with the first aspect, the above anti-crawler filtering and analysis steps further include:

[0020] Extract the IP in the host, detect high-frequency accessed IPs, set gradient-level thresholds according to different degrees of crawler behavior, assign corresponding tags to each threshold level, and according to the user's requirements, once a certain threshold is reached, even if it meets the specifications, the access will be blocked.

[0021] Combined with the first aspect, the above anti-crawler filtering and analysis steps further include:

[0022] Dynamically adjust the interception strategy of the request header according to different time periods. During the traffic peak period, intercept all IPs that do not meet the request header conditions, and use low-threshold tags according to the user's requirements; during the traffic trough period, appropriately relax the strategy to reduce the possibility of misblocking normal users; use Elasticsearch Watcher to notify the operation and maintenance team to intervene when abnormal traffic or policy failure is detected.

[0023] Combined with the first aspect, the above open customization steps further include:

[0024] The query module has a fast interface response and low resource consumption, which has a certain positive impact from an operational perspective and has relatively low anti-crawler requirements; the booking module has a relatively complex underlying booking interface and high resource consumption. The booking interface is also one of the targets of malicious crawlers. Customize the following rules for the booking interface: If the same IP repeatedly calls the booking interface of the same flight, then attach a first-level anti-crawler tag and store it in the cache.

[0025] Combined with the first aspect, the above regular push steps further include: In steps three and four, different levels of danger tags will be added to the suspected crawler IPs. Use the tags as the cache KEY, and each KEY corresponds to multiple suspected IPs as the value; the KEY with the highest level of danger will be immediately pushed to the alarm center; the level of danger decreases in a gradient, and the lower the danger, the slower the push frequency.

[0026] Combined with the first aspect, it further includes: Push the suspected crawler IPs to the alarm center through the cache. Different tags represent different levels of urgency, allowing the alarm center to determine the push frequency and method according to the different tag levels. The alarm center pushes the highest-level tag once a minute. In addition to pushing to the gateway, it also sends emails and mobile phone numbers to the configured operation and maintenance team.

[0027] The second aspect of the present application provides an anti-crawler application system based on ELK log collection. The above system includes:

[0028] A request receiving module, after receiving a traffic request forwarded from the gateway, enters the log cleaning module of the ELK system and performs log cleaning in Logstash;

[0029] A traffic analysis module, which accurately analyzes the traffic peak and trough periods in the current period according to the number of requests in the whole time period, dynamically adjusts the anti-crawler strategy, counts and analyzes the number of requests in different time periods, and identifies the traffic patterns of each day, week, and month;

[0030] An anti-crawler filtering and analysis module, which intercepts the User-Agent, Referer, and cookie fields in the request header, compares them with the standard request headers in the database, analyzes whether the access path and process of the Referer header conform to the reasonable user profile information in the database, and determines whether to intercept all or part of them according to the time period;

[0031] An open customization module, and the requests sent to the airline's interface will enter different modules of the airline according to the type of the requests;

[0032] A timing push module, which stores the identified problem IPs in the cache and pushes them to the alarm center for crawler behavior IP data at regular intervals.

[0033] Compared with the prior art, the present application has the following advantages:

[0034] Lightweight: The ELK log system, cache, etc. adopted in the present application have low requirements for CPU and memory.

[0035] Customizable: It is more flexible than anti-crawler tools such as WAF firewalls on the market. It can make customizable anti-crawler rules and can also be dynamically adjusted through traffic analysis.

[0036] Easy to maintain: The airline's maintenance team does not need to input too many instructions, etc. It only needs to customize rules on the visualization interface and receive alarms on the gateway, which is simple and easy to use.

[0037] Low cost: The deployment of a set of ELK log systems can be separated by indexes and sliced by role access control to divide tenants, which can be provided for multiple airlines to use together, reducing the deployment cost and being easier to maintain.

[0038] Other features and advantages of the present application will be described in the following specification, and part of them will become obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures pointed out in the specification, claims, and drawings. Brief Description of the Drawings

[0039] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0040] Figure 1 It shows a module diagram of an anti-crawling application system based on ELK log collection;

[0041] Figure 2 It shows a flowchart of an anti-crawling application method based on ELK log collection;

[0042] Figure 3 It is a schematic structural diagram of a device in an embodiment of the present application. Detailed implementation manners

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0044] Before elaborating on the present application in detail, the terms and abbreviations used in the technical solutions of the present application will be explained to better understand the present application.

[0045] Check ratio: The ratio of the number of flight inquiries received by an airline's website or reservation system to the number of actual bookings completed within a certain period. It is an important indicator for measuring the efficiency of an airline reservation system and the cost-effectiveness of online inquiries.

[0046] OCR: Optical Character Recognition, which aims to extract text from paper documents, image files, or photos and convert it into an editable and searchable digital text format.

[0047] ELK: A lightweight log collection and processing tool.

[0048] WAF: Web Application Firewall, which is mainly used to protect Web applications from various attacks, such as SQL injection, Cross-Site Scripting (XSS), Cross-Site Request Forgery (CSRF), etc.

[0049] Elasticsearch: It collects and analyzes log data from various sources, and can be used for data exploration and report generation in applications, as well as monitoring the performance metrics of the system and displaying data in real time. It is part of the Elastic Stack (also known as the ELK Stack).

[0050] Logstash: It collects, parses, and transports log files and is part of the Elastic Stack (also known as the ELK Stack).

[0051] The present application will be further elaborated in detail below with reference to the accompanying drawings and specific embodiments.

[0052] As Figure 1 shown, in order to enhance the anti-crawler ability and overall performance of the system, the anti-crawler application is deployed below the gateway, and efficient request processing is achieved through the following architecture:

[0053] First of all, all external requests will first pass through the gateway. The main responsibility of the gateway is to achieve load balancing and evenly distribute requests to different anti-crawler application instances to ensure that the system can still operate stably under high concurrency. At the same time, the airline gateway has the function of blocking IPs, receiving and processing suspected crawler IPs pushed by the alarm center. The specific processing operations depend on the airline gateway's situation.

[0054] Then, the requests processed by the gateway load balancing will be forwarded to the anti-crawler application. The core task of the anti-crawler application is to efficiently and strictly filter and verify these requests, identify and block malicious crawler behaviors. The anti-crawler application dynamically determines the anti-crawler rules through real-time traffic analysis (here the anti-crawler rules are conventional anti-crawler rules, mainly receiving suspected crawler IPs pushed by the alarm center, and each rule has a threshold label), and then performs anti-crawler filtering according to the customized rules for the airline interface. Finally, the suspected crawler IPs are stored in the cache with the label as the KEY, and then the alarm center pulls the data and feeds it back to the airline gateway, and the airline gateway decides whether to block or throttle.

[0055] After passing through the filtering of the anti-crawler application, legitimate requests will be further forwarded to the backend service for actual business processing. The backend service is responsible for completing the specific business logic of the requests to ensure the normal operation of the system functions.

[0056] To achieve comprehensive monitoring and log management of the entire system, the anti-crawler application introduces the ELK (Elasticsearch, Logstash, Kibana) technology stack. First, Logstash is responsible for collecting log data from the gateway and backend services and performing preliminary cleaning and filtering. Then, these log data will be stored in Elasticsearch for quick retrieval and analysis. Finally, through Kibana, the anti-crawler application can view and analyze these log data in real time to identify potential problems and improve the anti-crawler strategy.

[0057] Through the above architecture design, the anti-crawler application not only realizes effective defense against malicious crawlers but also ensures the high availability and scalability of the system. The processing process of each request has been closely monitored and managed, thus ensuring the security and efficiency of the system.

[0058] As Figure 2 shown, the requests forwarded by the anti-crawler application gateway are mainly executed through the following steps:

[0059] Step 1: After receiving the traffic request forwarded from the gateway, it enters the log cleaning module of the ELK system and performs detailed log cleaning work in Logstash. Specifically, the Logstash module will parse and filter the log data, extract key information from it, including but not limited to the source IP address, User-Agent information in the request header, and the airline interface accessed by the request, etc. This process involves structuring the original log data, extracting fields crucial for subsequent analysis and processing, and writing them into Elasticsearch for the anti-crawler application to obtain.

[0060] Step 2: Enter the traffic analysis module. In order to accurately analyze the peak and trough periods of the current time period based on the number of requests throughout the day and dynamically adjust the anti-crawler strategy. Statistically analyze the number of requests in different time periods to identify daily, weekly, and monthly traffic patterns. A real-time traffic analysis report can be generated by presetting and scheduling tasks to summarize the traffic peak of the previous day. According to the traffic analysis results, dynamically adjust the anti-crawler strategy. At the same time, Elasticsearch Watcher can be used to notify the operation and maintenance team to intervene in a timely manner when abnormal traffic or policy failure is detected.

[0061] During the traffic peak period, adopt more stringent crawler detection and blocking measures; during the traffic trough period, the policy can be appropriately relaxed to reduce the possibility of mis-blocking normal users.

[0062] Step 3: Enter the regular anti-crawler filtering module, intercept fields such as User-Agent, Referer, and cookie in the request header, and compare them with the standard request headers in the database. Analyze the access path of the Referer header to see if the process conforms to the reasonable user profile in the database. Control all interceptions or partial interceptions according to the time period. For example, during peak hours, intercept requests as long as there are any conditions not met. During off-peak hours, intercept requests that do not meet the three-point standard request headers. During peak traffic, intercept all IPs that do not meet the request header conditions and use low-threshold tags according to user requirements; during off-peak traffic, appropriately relax the policy to reduce the possibility of misblocking normal users; use Elasticsearch Watcher to notify the operation and maintenance team for intervention when abnormal traffic or policy failure is detected. Extract the IP in the host, detect high-frequency access IPs, set gradient-level thresholds according to different levels of crawler behavior, and assign corresponding tags to each threshold level. According to user requirements, once a certain threshold is reached, intercept access even if it meets the specifications (applicable to the case of only querying without booking).

[0063] Step 4: Open the custom rule window. Requests sent to the airline's interfaces will enter different airline modules according to the request type, and the anti-crawler restrictions of each module are different. For example, in the query module, the interface response is fast, consumes less resources, and has a certain positive impact from an operational perspective, so the anti-crawler requirements are relatively low. However, in the booking module, on the one hand, the interface underlying layer is relatively complex and consumes a lot of resources. At the same time, the booking interface is also one of the targets of malicious crawlers. Therefore, the following rules can be customized for this interface separately: If the same IP repeatedly calls the booking interface of the same flight, then attach a first-level anti-crawler tag and store it in the cache.

[0064] In addition, there is another special case. In the operation of the airline project, the ratio of querying to booking is an important data for billing reference. Therefore, the custom rules can also combine the query and booking interfaces for anti-crawling. For example, if within a certain time range, the access frequency of the query interface and the booking interface by the same IP is greater than the acceptable difference ratio range of the airline, then tag this IP and push it to the gateway. For more custom content, it can be customized according to the operating conditions and business strategies of the airline.

[0065] Step 5: Store the problematic IPs identified by the anti-crawler module in the cache and regularly push the crawler behavior IP data to the alarm center to improve the performance of the application program, reduce latency, and save network bandwidth. The specific process is as follows:

[0066] In steps three and four, suspected crawler IPs are tagged with different levels of risk. Using the tags as cache keys, each key corresponds to multiple suspected IPs as values. To save server resources, the key with the highest level of risk is immediately pushed to the alert center. The level of risk decreases in a gradient, and the lower the risk, the slower the push frequency.

[0067] Step six: Push the suspected crawler IPs to the alert center through the cache. Different tags represent different levels of urgency, allowing the alert center to determine the push frequency and method based on the tag levels (for example, the alert center pushes the highest-level tag once a minute, and in addition to pushing to the gateway, it also sends emails, mobile phone numbers, etc. to the configured operation and maintenance teams).

[0068] This application proposes a lightweight automated anti-crawling solution.

[0069] 1. System architecture:

[0070] Airline gateway: The gateway system of the airline, used for load balancing, processing all incoming requests, recording logs, and handling crawler IPs returned by the anti-crawling application.

[0071] Backend service: The backend service of the airline, handling airline business logic and data storage.

[0072] Anti-crawling application: Analyze the data obtained from the ELK Stack to detect crawler behavior and take corresponding measures.

[0073] ELK Stack: Collect, clean, store, and visualize log data.

[0074] Cache: Store the analysis results in the cache and push them to the alert center at regular intervals.

[0075] 2. Location of the anti-crawling application:

[0076] The anti-crawling application should be deployed between the gateway and the backend service of the airline's official website. The anti-crawling application filters the requests received by the gateway for anti-crawling and then forwards them to the backend service. This can intercept and process requests with suspected crawler behavior before they reach the backend service.

[0077] 3. Receiving and processing of data:

[0078] The gateway enables the log module, sends the logs to Logstash for processing and cleaning, extracts relevant fields such as IP address, request path, user agent, the interfaces of the airline's official website called, etc., and then sends the cleaned logs to the anti-crawling application for analysis.

[0079] 4. Implementation of the dynamic adjustment strategy for the anti-crawling strategy

[0080] During peak traffic periods, more stringent crawler detection and blocking measures are adopted; during off-peak traffic periods, the strategy can be appropriately relaxed to reduce the possibility of mistakenly blocking normal users. Therefore, the anti-crawler application sets up a traffic analysis module to accurately analyze the peak and off-peak traffic periods of the current time period based on the number of requests throughout the day, and dynamically adjusts the anti-crawler strategy. The specific implementation is as follows: Use the aggregation function in Elasticsearch to count and analyze the number of requests in different time periods, identify the daily, weekly, and monthly traffic peaks and troughs, output the traffic analysis results, and dynamically adjust the anti-crawler strategy according to the traffic analysis results. At the same time, Elasticsearch Watcher can be used to notify the operation and maintenance team for intervention in a timely manner when abnormal traffic or policy failure is detected.

[0081] 5. Blocking and Feedback Strategies

[0082] Whether it is a regular anti-crawler rule or an anti-crawler rule customized by an airline, gradient-level thresholds (N-level thresholds) will be set according to different degrees of crawler behavior, and corresponding labels will be assigned to each threshold level. For example:

[0083] Suspected crawler behavior: The access frequency is relatively high, but it has not reached an obviously abnormal level.

[0084] Tolerable crawler behavior: The access frequency further increases, but it is still within the tolerable range.

[0085] High crawler behavior: The access frequency is very high, obviously abnormal, and may cause pressure on the system.

[0086] For example, if this IP accesses the same interface 100 times per minute, then assign a first-level label to this IP: Suspected crawler behavior. If this IP accesses the same interface 300 times per minute, then assign a first-level label to this IP: Tolerable crawler behavior.

[0087] In addition, according to the feedback from the maintenance team and the actual operation effect, continuously optimize the anti-crawler strategy and threshold settings to ensure the balance between the effectiveness of the system and the user experience, and regularly analyze the log data to evaluate the execution effect of the anti-crawler strategy, identify new abnormal behavior patterns and adjust the strategy in a timely manner.

[0088] Based on the same inventive concept, the present application also provides a computer-readable storage medium storing one or more programs, which can implement the foregoing method when the one or more programs are executed.

[0089] As Figure 3 shown, the embodiment of the present application also provides a device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory communicate with each other through the communication bus.

[0090] The memory is a computer-readable storage medium for storing one or more programs.

[0091] The processor is configured to execute the programs stored in the computer-readable storage medium.

[0092] The computer-readable storage medium may be included in the device / apparatus described in the foregoing embodiments; or it may exist alone without being assembled into the device / apparatus.

[0093] Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An anti-crawling application method based on ELK log collection, characterized in that, Including: Step 1: Request receiving step. After receiving the traffic request forwarded by the gateway, it enters the log cleaning module of the ELK system and performs log cleaning in Logstash; Step 2: Traffic analysis step. According to the request quantity throughout the entire period, accurately analyze the traffic peak and trough periods of the current period, dynamically adjust the anti-crawler strategy, count and analyze the request quantities in different time periods, and identify the daily, weekly, and monthly traffic patterns; Step 3: Anti-crawler filtering and analysis step. Intercept the User-Agent, Referer, and cookie fields in the request header, compare them with the standard request headers in the database, analyze whether the access path and process of the Referer header conform to the reasonable user portrait information in the database, and determine whether to intercept all or part of them according to the time period; Step 4: Open customization step. The requests sent to the airline's interface will enter different modules of the airline according to the request type; Step 5: Timed push step. Store the identified problem IPs in the cache and push them to the alarm center for crawler behavior IP data at regular intervals.

2. The anti-crawling application method based on ELK log collection according to claim 1, characterized in that, The log cleaning in Logstash further includes: In Logstash, parse and filter the log data, and extract the key information from it. The key information includes the source IP address, the User-Agent information in the request header, and the airline interface accessed by the request.

3. The anti-crawling application method based on ELK log collection according to claim 1, characterized in that, The anti-crawler filtering and analysis step further includes: Extract the IP in the host, detect the high-frequency accessed IPs, set gradient-level thresholds according to different degrees of crawler behavior, assign corresponding tags to each threshold level, and according to the user's requirements, once a certain threshold is reached, even if it meets the specifications, the access will be intercepted.

4. The anti-crawling application method based on ELK log collection according to claim 3, wherein, The anti-crawler filtering and analysis step further includes: Dynamically adjust the interception strategy of the request header according to different time periods. During the traffic peak period, intercept all IPs that do not meet the request header conditions, and use low-threshold tags according to the user's requirements; during the traffic trough period, appropriately relax the strategy to reduce the possibility of misblocking normal users; use Elasticsearch Watcher to notify the operation and maintenance team to intervene when abnormal traffic or policy failure is detected.

5. The anti-crawling application method based on ELK log collection according to claim 4, wherein, The open customization step further includes: Query module. The interface response is fast and the resource consumption is low, which has a certain positive impact from the operation perspective, and the anti-crawler requirements are low; booking module. The underlying layer of the booking interface is relatively complex and the resource consumption is high. The booking interface is also one of the targets of malicious crawlers. Customize the rules for the booking interface as follows: If the same IP repeatedly calls the booking interface of the same flight, then attach a first-level anti-crawler tag and store it in the cache.

6. The anti-crawler application method based on ELK log collection according to claim 5, wherein The timed push step further includes: In steps 3 and 4, different levels of danger tags will be added to the suspected crawler IPs. Use the tag as the cache KEY, and each KEY corresponds to multiple suspected IPs as the value; the KEY with the highest level of danger will be immediately pushed to the alarm center; the level of danger decreases in a gradient, and the lower the danger, the slower the push frequency.

7. The anti-crawling application method based on ELK log collection according to claim 6, characterized in that, Further including: Push suspected crawler IPs to the alarm center through caching. Different tags represent different levels of urgency, enabling the alarm center to determine the push frequency and method based on the tag levels. The alarm center pushes the highest-level tags once a minute. In addition to pushing to the gateway, emails and mobile phone numbers are also sent to the configured operation and maintenance teams.

8. An anti-crawling application system based on ELK log collection, characterized in that, The system includes: A request receiving module. After receiving a traffic request forwarded by the gateway, it enters the log cleaning module of the ELK system for log cleaning in Logstash. A traffic analysis module. It accurately analyzes the traffic peak and trough periods in the current period based on the number of requests over the entire time period, dynamically adjusts the anti-crawling strategy, counts and analyzes the number of requests in different time periods, and identifies the daily, weekly, and monthly traffic patterns. An anti-crawling filtering and analysis module. It intercepts the User-Agent, Referer, and cookie fields in the request header, compares them with the standard request headers in the database, analyzes whether the access path and process of the Referer header conform to the reasonable user profile information in the database, and determines whether to intercept all or part of the requests according to the time period. An open customization module. The requests sent to the airline's interface will enter different modules of the airline according to the request type. A timed push module. It stores the identified problematic IPs in the cache and periodically pushes the crawler behavior IP data to the alarm center.

9. A computer-readable storage medium storing one or more programs, characterized in that when the one or more programs are executed, the anti-crawling application method based on ELK log collection described in any one of claims 1-8 can be implemented.

10. An apparatus, comprising a processor, a communication interface, the computer-readable storage medium according to claim 9, and a communication bus; wherein, The processor, communication interface, and computer-readable storage medium communicate with each other through a communication bus; characterized in that the processor is used to execute the programs stored in the computer-readable storage medium.