Scenario-based self-learning method and system for detecting malicious requests

By adopting a scenario-based self-learning method in malicious request detection, combining a layered detection strategy and an integrated scoring engine, the problem of high false alarm and missed response rates of malicious request detection in the existing technology is solved, and more efficient and accurate malicious request detection is achieved.

WO2025124232A1PCT designated stage expired Publication Date: 2025-06-19CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
PCT/CN2024/136519
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-12-03
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The prior art is difficult to quickly detect and distinguish anthropomorphic and precise bot requests, resulting in high false alarm and missed alarm rates, and poor detection results caused by general configuration.

Method used

The malicious request detection method of scenario-based self-learning is adopted, and the combination of hierarchical detection strategies and integrated scoring engines is used to reduce the rate of missed and false alarms, and the detection strategy is dynamically updated through the traffic self-learning strategy.

Benefits of technology

It improves the accuracy and adaptability of malicious request detection, reduces the false alarm and missed alarm rates, and enhances the detection capabilities of malicious attacks in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136519_19062025_PF_FP_ABST
    Figure CN2024136519_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the field of network technology and security, and in particular, to a scenario-based self-learning method and system for detecting malicious requests. The method mainly comprises: scenario-based one-click configuration, customized configuration of detection granularity and strategies, and an integrated scoring engine that combines anomaly scores obtained from different detection strategies. In addition, the method also comprises generating configuration feedback suggestions on the basis of traffic self-learning and presenting same to a client in a clear and operable manner. The client can adjust configurations on the basis of the feedback suggestions, generate personalized scenario plans for seamless application, and conduct gray box testing before deployment to ensure the usability of the current scenario-based configuration. The present invention achieves the configuration of flexible and adaptive capabilities for service scenarios by finely categorizing requests for different scenarios, applying different traffic restrictions for different scenarios, and combining different detection functions for different scenarios, so that the security detection capability of scenarios is improved, and missed detections and false positives caused by general configurations are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

A scenario-based self-learning malicious request detection method and system

[0001] This application claims priority to Chinese patent application number CN202311703138.3, filed on December 12, 2023, entitled “A scenario-based self-learning method and system for detecting malicious requests,” the entire text of which is hereby incorporated by reference. Technical Field

[0002] The present invention belongs to the field of network technology and security, and in particular relates to a scenario-based self-learning malicious request detection method and system. Background Art

[0003] With the exponential growth of global computer networks and network applications, coupled with the advancement of intelligence, intelligent, automated cyberattacks have surged. Criminals are using automated scripts or tools to mimic human behavior to carry out simple and efficient cyberattacks. With the help of automated tools, the labor-intensive, high-intelligence, and high-cost cyberattacks of the past are no longer the exclusive domain of advanced hackers. Ordinary cybercriminals can now exploit website vulnerabilities in a short period of time, efficiently, and stealthily. Cyberattacks have found a new paradigm shift thanks to AI-based adversarial learning and automated tools, leading to increasingly more humanized and sophisticated attacks. These bots, mimicking human behavior, are smarter, more daring, and more difficult to track and distinguish from real people. Frequent data breaches have resulted in identity theft and trafficking. Cybercriminals can easily exploit this exposed personal data by combining automated scripts or tools to repeatedly perform login verification on hundreds of different websites in a short period of time, attempting to compromise accounts and potentially launch further attacks and profit from the resulting losses. As the conflict between automated attacks and security defenses continues to escalate, a growing number of black and gray market organizations are offering various countermeasures. These services, such as proxy IP services, graphic verification code recognition, SMS verification code collection, group control device pools, and account providers, are readily available. While most automated attack defenses are easily circumvented, this has given rise to new, more anthropomorphic automated attacks. These malicious automated attacks pose a significant threat to system security by using emulators, forged browser environments, user agents (UAs), and distributed IP addresses. The three key characteristics of automated attacks, namely, being free, simple, and highly effective, have exacerbated their prevalence, repeatedly breaching traditional enterprise network security defenses. A Forrester report indicates that current WAF solutions are unable to handle a wider range of application attacks, particularly automated attacks driven by bots, causing significant challenges for enterprise users.

[0004] For example, the existing invention patent with publication number CN115525813A discloses an application scenario-based web crawler detection system, which includes a crawler detection platform, which comprises a user analysis unit, a human-machine identification unit, a secondary verification unit, a space development unit, and a reality integration unit. This detection system divides the enterprise network into different application scenarios based on their usage, and performs triple malicious crawler identification based on these scenarios. This system not only screens malicious crawlers, but also opens personal spaces for registered users and manages these spaces by combining network information related to the application scenarios.

[0005] For example, the existing invention patent with publication number CN106027564A discloses a method and device for detecting the security of an anti-crawler strategy. The method embeds an anti-crawler code for implementing the anti-crawler strategy in the first front-end page of a website; uses the anti-crawler code to detect whether the user accessing the first front-end page is a crawler, and records the user detected as a crawler as the target object; verifies whether the target object is a crawler, and counts the number of times the target object is not a crawler; calculates the false alarm rate of the anti-crawler strategy based on the number of times, and the false alarm rate is used to measure the security of the anti-crawler strategy.

[0006] An analysis of the above-mentioned existing technologies reveals that most existing defense measures restrict requests by combining a general configuration with multiple defense technologies. These methods lack the ability to quickly detect personalized and sophisticated bot requests, and the use of a general configuration for bot request detection in different scenarios leads to excessively high false positives and false negatives. To address these issues, the present invention provides a scenario-based, self-learning method and system for detecting malicious requests. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to address the shortcomings of the existing technology and provide a scenario-based self-learning malicious request detection method and system. This method reduces the missed detection and false alarm rates of detection anomalies by combining a layered detection strategy with an integrated scoring engine. In addition, the use of a traffic self-learning strategy enables the method to have the ability to be dynamically updated, thereby improving the adaptability of the method.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] A scenario-based self-learning method for detecting malicious requests includes the following steps:

[0010] S1: Configure basic scenarios and configure user statistics granularity and detection policies for these scenarios.

[0011] S2: Utilize user statistical granularity and detection strategies to obtain scores for different anomalies in the configuration scenario;

[0012] S3: Configure an integrated scoring engine and use it to combine the different anomaly scores obtained to obtain a comprehensive anomaly score and determine the severity of the anomaly.

[0013] S4: Visualize the severity of the acquired abnormalities;

[0014] S5: Analyze traffic data for potential threats and provide enhanced feedback for detection policy configuration.

[0015] S6: Set up a grayscale test and verify the above scenario configuration through the grayscale test, obtain and fix scenario compatibility or misconfiguration issues, and then deploy the scenario.

[0016] Specifically, the configuration process of user statistics granularity in S1 includes:

[0017] S2.1: The customer configures the user statistics granularity by selecting IP, IP+UA, custom token, and gateway-specific cookie;

[0018] S2.2: Set the following priorities for the user statistics granularity from high to low: custom token, gateway-specific cookie, ip+ua, ip;

[0019] S2.3: When the user statistical granularity at a higher priority level is missing, the user statistical granularity at a lower priority level is used as the location identifier of the user's unique identity.

[0020] Specifically, the detection strategies include front-end confrontation, access speed limiting, intelligence analysis, and intelligent analysis. Front-end confrontation, access speed limiting, intelligence analysis, and intelligent analysis all include preset detection thresholds.

[0021] Specifically, the specific process of obtaining different anomaly scores in S2 includes:

[0022] S4.1: Based on the user statistics granularity configured by the customer, the scenario system automatically collects user statistics granularity data for all access requests in different scenarios;

[0023] S4.2: The collected user statistical granular data of the high-priority scenario is first input into the parallel detection strategy layer of the scenario configuration for analysis and calculation, and the anomaly score is independently set according to the weight of the detection strategy; then the next level scenario is detected, and so on to obtain the anomaly scores of different scenarios.

[0024] Specifically, the comprehensive score in S3 includes a single anomaly comprehensive score and a multiple anomaly comprehensive score. The specific calculation process includes:

[0025] S5.1 When the comprehensive score is a single anomaly comprehensive score, the anomaly score obtained from the single detection strategy is directly used as the comprehensive score;

[0026] S5.2: When the comprehensive score is a composite score of multiple anomalies, the highest score among the anomalies in the detection strategy is used as the highest score for a single anomaly, and the scores of the remaining detection strategies are multiplied by the weights corresponding to the detection strategies and added to the highest score for a single anomaly to obtain a comprehensive score. The formula is as follows:

[0027] M=M1+w2M2+w3M3+w4M4,

[0028] Among them, M is the comprehensive multi-anomaly score, M1 is the highest score of a single anomaly, M2, M3, and M4 are the scores of the second, third, and fourth anomalies respectively, and w2, w3, and w4 are the weights of the corresponding detection strategies of M2, M3, and M4 respectively.

[0029] Specifically, the configuration enhancement suggestion feedback adopts a traffic self-learning strategy. The specific process includes:

[0030] S6.1: Based on the potential threats obtained, use the statistical granularity configured in the scenario to collect and preprocess the threat traffic data.

[0031] S6.2: Extract specific attack-related traffic patterns, abnormal access behaviors, or other suspicious activities from the preprocessed traffic data as attack samples, and use the detection threshold of the detection policy corresponding to the potential threat as the initial threshold;

[0032] S6.3: Input the above attack samples and the initial threshold into a pre-trained support vector machine to obtain the attack pattern and trend of the threat and output the optimal threshold corresponding to the threat;

[0033] S6.4: Use the acquired attack patterns and trends to adjust the configuration and weights of the front-end confrontation, access rate limiting, intelligence analysis, and intelligent analysis, and use the acquired optimal thresholds to update the preset detection thresholds in the front-end confrontation, access rate limiting, intelligence analysis, and intelligent analysis.

[0034] A scenario-based self-learning malicious request detection system, including a scenario-based basic configuration module, an integrated scoring engine module, a configuration suggestion feedback module, and a grayscale testing module;

[0035] Scenario-based basic configuration module: used to provide customers with optional or customized scenario configurations, and configure detection granularity and detection strategies for the selected scenarios;

[0036] Integrated scoring engine module: used to integrate the anomaly scores obtained from different detection strategies to determine the severity of the detected anomaly and provide appropriate measures;

[0037] Configuration suggestion feedback module: used to provide accurate and effective configuration suggestions for each measure;

[0038] Grayscale testing module: used to verify the final configuration and ensure that normal users can successfully pass the verification process.

[0039] Specifically, the scenario-based basic configuration module includes a scenario-based basic setting unit, a user statistics granularity unit, and a detection strategy unit;

[0040] The scenario-based basic setting unit is used to provide customers with scenarios with different configurations and provide one-click configuration functions for different scenarios;

[0041] User statistics granularity unit, used to provide different levels of user identification granularity for customer-configured scenarios;

[0042] The detection strategy unit is used to provide different types of detection strategies for customer-configured scenarios and independently set anomaly scores based on the weights of the detection strategies.

[0043] Specifically, the detection strategy unit includes a front-end confrontation sub-unit, an access rate limiting sub-unit, an intelligence analysis sub-unit, and an intelligent analysis sub-unit;

[0044] The front-end adversarial sub-unit is used to detect abnormal behaviors such as crawler behavior, automated attacks, and page debugging, and obtain anomaly scores;

[0045] The access rate limit sub-unit is used to control access frequency at the user granularity (such as IP address), detect anomalies based on access frequency, and obtain anomaly scores;

[0046] The intelligence analysis subunit is used to manage blacklists and whitelists in the threat intelligence database, perform anomaly detection based on the managed blacklists and whitelists, and obtain anomaly scores;

[0047] The intelligent analysis subunit is used for access behavior analysis and human-machine behavior analysis, and performs anomaly detection through the behavior analysis to obtain anomaly scores.

[0048] Specifically, the configuration suggestion feedback module includes a result display subunit and a traffic self-learning subunit;

[0049] The result display subunit is used to visualize the results with anomaly scores above the threshold, helping customers quickly identify potential threats and take appropriate measures to mitigate these threats;

[0050] The traffic self-learning subunit is used to identify patterns and trends of threat attacks and provide threshold enhancement configuration recommendations for detection policies.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] 1. This invention integrates traditional defense measures to address the problem of missed reports and false positives for malicious requests, and adopts a combination of multiple granularity configurations and detection strategies. It can more flexibly adapt to business scenarios and improve security detection capabilities, while reducing missed reports and false positives caused by general configurations.

[0053] 2. This invention addresses the problems existing in traditional personal configuration models and adopts a combination of traffic self-learning, scoring engine, and configuration feedback strategies to improve the comprehensiveness of risk assessment. It also enhances the performance of detecting malicious attacks in configuration scenarios, fully demonstrating the capabilities of the product.

[0054] 3. The present invention uses grayscale testing to verify the final configuration of the scenario, further ensuring the availability of assets. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0056] FIG1 is a flow chart of a scenario-based self-learning malicious request detection method according to embodiment 1 of the present invention.

[0057] FIG2 is a module diagram of a scenario-based self-learning malicious request detection system according to Example 1 of the present invention. DETAILED DESCRIPTION

[0058] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0059] Example 1:

[0060] Currently, many manufacturers' defense measures restrict requests by combining a general configuration with multiple defense technologies. These methods lack the ability to quickly detect personalized and sophisticated bot requests. Using a general configuration to detect bot requests in different scenarios leads to excessively high false positives and false negatives.

[0061] To solve the above problem, please refer to FIG1 . An embodiment of the present invention is provided: a scenario-based self-learning malicious request detection method, which specifically includes the following steps:

[0062] S1: Configure basic scenarios and configure user statistics granularity and detection policies for these scenarios.

[0063] S2: Utilize user statistical granularity and detection strategies to obtain scores for different anomalies in the configuration scenario;

[0064] S3: Configure an integrated scoring engine and use it to combine the different anomaly scores obtained to obtain a comprehensive anomaly score and determine the severity of the anomaly.

[0065] S4: Visualize the severity of the acquired abnormalities;

[0066] S5: Analyze traffic data for potential threats and provide enhanced feedback for detection policy configuration.

[0067] S6: Set up a grayscale test and verify the above scenario configuration through the grayscale test, obtain and fix scenario compatibility or misconfiguration issues, and then deploy the scenario.

[0068] In the above description, basic scenarios include recommended scenarios such as login, registration, and purchase, as well as custom scenarios. Recommended scenarios are adjusted using common models, which typically include but are not limited to user authentication models, purchase behavior models, and recommendation models. These scenarios can be configured with one click. For custom scenarios, customers can configure custom protection parameters based on the requested characteristics of the desired protection scenario. The priority of each custom scenario is determined based on the configured priority value.

[0069] In this embodiment, the one-key configuration function is implemented by creating a configuration template containing the above configuration information and a button connected to the configuration template.

[0070] Configuring the above scenarios with one click can reduce configuration errors and simplify maintenance. In addition, using the same configuration template can ensure consistent configurations in different scenarios, thereby ensuring system stability and consistency.

[0071] The configuration process of user statistics granularity in S1 includes:

[0072] S2.1: The customer configures the user statistics granularity by selecting IP, IP+UA, custom token, and gateway-specific cookie;

[0073] S2.2: The user statistics granularity is prioritized in descending order as follows: custom token, gateway-specific cookie, ip+ua, ip; wherein, cookie is a small piece of data stored on the user's computer.

[0074] S2.3: When the user statistics granularity at a higher priority level is missing, the user statistics granularity at a lower priority level is used as the location identifier of the user's unique identity. For example, when the custom token expires, a gateway-specific cookie is used as the location identifier of the user's unique identity.

[0075] Detection strategies include front-end confrontation, access rate limiting, intelligence analysis, and intelligent analysis. These strategies all include pre-set detection methods and thresholds. In front-end confrontation, specific detection methods such as crawler traps, automated attack identification, and page debugging prevention are used to detect anomalies and obtain corresponding anomaly scores.

[0076] In access rate limiting, anomaly detection is performed by frequency monitoring and analysis at the user granularity; for example, if a user's access frequency exceeds the preset detection threshold, the system will determine it as abnormal behavior and obtain a corresponding anomaly score.

[0077] Intelligence analysis detects anomalies by matching requests against blacklists and whitelists in the managed threat intelligence database. For example, if an IP address or user is on the blacklist, their behavior will be considered abnormal and receive a corresponding anomaly score. Conversely, if they are on the whitelist, their behavior will generally not be considered abnormal.

[0078] Intelligent analysis uses behavioral analysis and offline engines to detect user click behavior and identify abnormal click behavior. For example, if a user performs a large number of meaningless clicks in a short period of time, or if the click behavior pattern differs significantly from that of a normal user, it will be identified as abnormal and assigned a corresponding anomaly score.

[0079] The specific process of obtaining different anomaly scores in S2 includes:

[0080] S4.1: Based on the user statistics granularity configured by the customer, the scenario system automatically collects user statistics granularity data for all access requests in different scenarios;

[0081] S4.2: The collected user statistical granular data of the high-priority scenario is first input into the parallel detection strategy layer of the scenario configuration for analysis and calculation, and the anomaly score is independently set according to the weight of the detection strategy; then the next level scenario is detected, and so on to obtain the anomaly scores of different scenarios.

[0082] Assume that within a certain period of time, the scenario corresponding to the custom token receives 1000 access requests, the front-end adversarial weight is 0.6, and the detection threshold is 500. Based on this weight, the anomaly score of the scenario corresponding to the custom token under the front-end adversarial detection strategy is calculated to be 600 points, and the detection threshold can be used to determine that it is an anomaly.

[0083] During the same time period, the scenario corresponding to the gateway-specific cookie received 500 access requests, the access rate limit weight was 0.4, and the detection threshold was 400. Based on this weight, the anomaly score for the scenario corresponding to the gateway-specific cookie is calculated to be 200 points, and it is not considered an anomaly.

[0084] The comprehensive score in S3 includes a single anomaly comprehensive score and a multiple anomaly comprehensive score. The specific calculation process includes:

[0085] S5.1 When the comprehensive score is a single anomaly comprehensive score, the anomaly score obtained from the single detection strategy is directly used as the comprehensive score;

[0086] S5.2: When the comprehensive score is a composite score of multiple anomalies, the highest score among the anomalies in the detection strategy is used as the highest score for a single anomaly, and the scores of the remaining detection strategies are multiplied by the weights corresponding to the detection strategies and added to the highest score for a single anomaly to obtain a comprehensive score. The formula is as follows:

[0087] M=M1+w2M2+w3M3+w4M4,

[0088] Among them, M is the comprehensive multi-anomaly score, M1 is the highest score of a single anomaly, M2, M3, and M4 are the scores of the second, third, and fourth anomalies respectively, and w2, w3, and w4 are the weights of the corresponding detection strategies of M2, M3, and M4 respectively.

[0089] The specific visual display method of S4 is:

[0090] The anomaly scores and comprehensive scores corresponding to front-end confrontation, access speed limiting, intelligence analysis, and intelligent analysis are displayed in the form of bar charts, and the detection thresholds corresponding to front-end confrontation, access speed limiting, intelligence analysis, and intelligent analysis are used as standard lines to obtain specific potential threats and the severity of potential threats, and take corresponding measures.

[0091] The following measures can be taken for anomalies obtained by the above different detection strategies:

[0092] To deal with abnormal behaviors acquired by the front-end adversary, we can enhance the defense capability of the scenario by checking whether there are vulnerabilities or malicious attacks in the front-end code, and taking measures such as adding verification codes, limiting access frequency, and preventing malicious crawlers.

[0093] For abnormal behaviors obtained through intelligence analysis, the defense capabilities of the scenario can be enhanced by adding authentication, enabling IP blocking functions, and other measures.

[0094] For abnormal behaviors such as access rate limits, the defense capabilities of the scenario can be enhanced by establishing a more comprehensive logging system and strengthening data encryption and protection.

[0095] For abnormal behaviors obtained by intelligent analysis, the defense capability of the scenario can be enhanced by measures such as enhancing the detection capability of the intelligent analysis model and limiting the number of clicks.

[0096] Configuration enhancement suggestion feedback adopts a traffic self-learning strategy. The specific process includes:

[0097] S6.1: Based on the potential threats obtained, use the statistical granularity configured in the scenario to collect and preprocess the threat traffic data.

[0098] S6.2: Extract specific attack-related traffic patterns, abnormal access behaviors, or other suspicious activities from the preprocessed traffic data as attack samples, and use the detection threshold of the detection policy corresponding to the potential threat as the initial threshold;

[0099] S6.3: Input the above attack samples and the initial threshold into a pre-trained support vector machine to obtain the attack pattern and trend of the threat and output the optimal threshold corresponding to the threat;

[0100] S6.4: Use the acquired attack patterns and trends to adjust the configuration and weights of the front-end confrontation, access rate limiting, intelligence analysis, and intelligent analysis, and use the acquired optimal thresholds to update the preset detection thresholds in the front-end confrontation, access rate limiting, intelligence analysis, and intelligent analysis.

[0101] The specific process of grayscale testing to verify the above scenario configuration is as follows:

[0102] A test IP is configured for the test subset of users. Customers perform functional testing according to the configured scenario-based basic configuration, statistical granularity configuration, detection policy configuration, traffic self-learning and other functional items, and collect user feedback and system logs and other information. The collected feedback information is analyzed to check whether there are any problems or incorrect configurations. If any problems are found, they are fed back to the development team for repair, and then retested until there are no problem feedback from all subset users. Then the verified scenarios are deployed online.

[0103] Example 2:

[0104] To further illustrate the scenario-based self-learning malicious request detection method in Example 1, this example is set up in scenarios with automated crawler attack protection, such as crawler management projects and WAF protection projects. The specific process is as follows:

[0105] S1: Configure scenarios with automated crawler attack protection, including crawler management projects and WAF protection projects, through one-click configuration switches, and assume that the statistical granularity is IP.

[0106] S2: The system in the configuration scenario automatically counts the number of request IPs through the configured user statistical granularity, assuming it is 1000; and detects the collected IPs through the configured front-end confrontation, access rate limit, intelligence analysis and intelligent analysis; in actual applications, the customer can choose the number of detection strategies by combining the above-mentioned detection strategies. In this embodiment, all four are selected for use, and it is assumed that the weights corresponding to the above-mentioned detection strategies and the preset detection thresholds are 0.4,300; 0.2,200; 0.3,200, 0.1,300.

[0107] S3: Calculation shows that the anomaly scores obtained by the four detection strategies are 400, 200, 300, and 100, respectively. Based on the detection threshold, the front-end adversarial and intelligence analysis detect anomalies, and the combined score obtained by configuring the integrated scoring engine is 700.

[0108] S4: Through the bar chart display, it can be obtained that the comprehensive score is far greater than all detection thresholds, and the threat level is serious; and both the front-end confrontation and intelligence analysis detect anomalies. In response to this, the traffic data of the front-end confrontation and intelligence analysis are input into the traffic self-learning model, and the abnormal attack pattern obtained is mainly crawlers, and the risk of being crawled is increasing. Therefore, based on this feedback, this embodiment increases the weights of the front-end confrontation and intelligence analysis to 0.6 and 0.4, and lowers the detection threshold to enhance the defense range of the crawler trap and the black and white list management range, thereby improving the defense capability of the entire scenario.

[0109] Example 3:

[0110] Referring to FIG2 , the present invention provides an embodiment: a scenario-based self-learning malicious request detection system, comprising a scenario-based basic configuration module, an integrated scoring engine module, a configuration suggestion feedback module, and a grayscale testing module;

[0111] Scenario-based basic configuration module: used to provide customers with optional or customized scenario configurations, and configure detection granularity and detection strategies for the selected scenarios;

[0112] Integrated scoring engine module: used to integrate the anomaly scores obtained from different detection strategies to determine the severity of the detected anomaly and provide appropriate measures;

[0113] Configuration suggestion feedback module: used to provide accurate and effective configuration suggestions for each measure;

[0114] Grayscale testing module: used to verify the final configuration and ensure that normal users can successfully pass the verification process.

[0115] Specifically, the scenario-based basic configuration module includes a scenario-based basic setting unit, a user statistics granularity unit, and a detection strategy unit;

[0116] The scenario-based basic setting unit is used to provide customers with scenarios with different configurations and provide one-click configuration functions for different scenarios;

[0117] User statistics granularity unit, used to provide different levels of user identification granularity for customer-configured scenarios;

[0118] The detection strategy unit is used to provide different types of detection strategies for customer-configured scenarios and independently set anomaly scores based on the weights of the detection strategies.

[0119] Specifically, the detection strategy unit includes a front-end confrontation sub-unit, an access rate limiting sub-unit, an intelligence analysis sub-unit, and an intelligent analysis sub-unit;

[0120] The front-end adversarial sub-unit is used to detect abnormal behaviors such as crawler behavior, automated attacks, and page debugging, and obtain anomaly scores;

[0121] The access rate limit sub-unit is used to control access frequency at the user granularity (such as IP address), detect anomalies based on access frequency, and obtain anomaly scores;

[0122] The intelligence analysis subunit is used to manage blacklists and whitelists in the threat intelligence database, perform anomaly detection based on the managed blacklists and whitelists, and obtain anomaly scores;

[0123] The intelligent analysis subunit is used for access behavior analysis and human-machine behavior analysis, and performs anomaly detection through the behavior analysis to obtain anomaly scores.

[0124] Specifically, the configuration suggestion feedback module includes a result display subunit and a traffic self-learning subunit;

[0125] The result display subunit is used to visualize the results with anomaly scores above the threshold, helping customers quickly identify potential threats and take appropriate measures to mitigate these threats;

[0126] The traffic self-learning subunit is used to identify patterns and trends of threat attacks and provide threshold enhancement configuration recommendations for detection policies.

[0127] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A scenario-based self-learning malicious request detection method, characterized in that: The steps include: S1: Configure basic scenarios and configure user statistics granularity and detection strategies for the basic scenarios; S2: Use user statistical granularity and detection strategies to obtain scores of different anomalies in the configuration scenario; S3: Configure an integrated scoring engine, use the integrated scoring engine to integrate the different anomaly scores obtained, obtain a comprehensive anomaly score, and determine the severity of the anomaly; S4: By visualizing the severity of the acquired abnormalities; S5: Provides enhanced feedback for detection policy configuration by analyzing traffic data for potential threats; S6: Set up a grayscale test and verify the above scenario configuration through the grayscale test, obtain and fix the scenario compatibility or misconfiguration issues, and then deploy the scenario.

2. According to the scenario-based self-learning malicious request detection method of claim 1, it is characterized in that: The configuration process of the user statistics granularity in S1 includes: S2.1: The customer configures the user statistics granularity by selecting IP, IP+UA, custom token and gateway specific cookie; S2.2: setting the following priorities for the user statistics granularity from high to low: custom token, gateway specific cookie, ip+ua, ip; S2.3: When the user statistical granularity at a higher priority level is missing, the user statistical granularity at a lower priority level is used as the location identifier of the user's unique identity.

3. According to the scenario-based self-learning malicious request detection method of claim 2, it is characterized in that: The detection strategy includes front-end confrontation, access speed limit, intelligence analysis and intelligent analysis, and the front-end confrontation, access speed limit, intelligence analysis and intelligent analysis all include preset detection thresholds.

4. According to the scenario-based self-learning malicious request detection method of claim 3, it is characterized in that: The specific process of obtaining different anomaly scores in S2 includes: S4.1: Based on the user statistics granularity configured by the customer, the system where the scenario is located automatically counts the user statistics granularity data of all access requests in different scenarios; S4.2: The collected user statistical granular data of the high-priority scene is first input into the parallel detection strategy layer of the scene configuration for analysis and calculation, and the anomaly score is independently set according to the weight of the detection strategy; then the next-level scene is detected, and so on to obtain the anomaly scores of different scenes.

5. According to the scenario-based self-learning malicious request detection method of claim 4, it is characterized in that: The comprehensive score in S3 includes a single anomaly comprehensive score and a multiple anomaly comprehensive score, and the specific calculation process includes: S5.1 When the comprehensive score is a single anomaly comprehensive score, the anomaly score obtained from a single detection strategy is directly used as the comprehensive score; S5.2: When the comprehensive score is a comprehensive score of multiple anomalies, the highest score of the anomaly in the detection strategy is used as the highest score of the single anomaly, and the scores of the remaining detection strategies are multiplied by the weights corresponding to the detection strategies and added to the highest score of the single anomaly to obtain a comprehensive score, the formula of which is as follows: M=M1+w2M2+w3M3+w4M4, Among them, M is the comprehensive multi-anomaly score, M1 is the highest score of a single anomaly, M2, M3, M4 are the scores of the second, third and fourth anomalies respectively, and w2, w3, w4 are the weights of the corresponding detection strategies of M2, M3, M4 respectively.

6. A scenario-based self-learning malicious request detection method according to claim 5, characterized in that: The configuration enhancement suggestion feedback adopts a traffic self-learning strategy. The specific process includes: S6.1: According to the acquired potential threat, using the statistical granularity configured in the scenario, collect the traffic data of the threat and perform preprocessing; S6.2: Extract specific traffic patterns, abnormal access behaviors or other suspicious activities related to the attack from the pre-processed traffic data as attack samples, and use the detection threshold of the detection strategy corresponding to the potential threat as the initial threshold; S6.3: Input the above attack samples and the initial threshold into a pre-trained support vector machine to obtain the attack mode and trend of the threat, and output the optimal threshold corresponding to the threat; S6.4: Use the acquired attack patterns and trends to adjust the configuration and weight of the front-end confrontation, access speed limit, intelligence analysis and intelligent analysis, and use the acquired optimal threshold to update the preset detection threshold in the front-end confrontation, access speed limit, intelligence analysis and intelligent analysis.

7. A scenario-based self-learning malicious request detection system, which is implemented based on a scenario-based self-learning malicious request detection method as described in any one of claims 1-6, characterized in that: The system includes a scenario-based basic configuration module, an integrated scoring engine module, a configuration suggestion feedback module, and a grayscale testing module; The scenario-based basic configuration module is used to provide customers with selectable scenario configurations or customized scenario configurations, and configure detection granularity and detection strategies for the selected scenarios; The integrated scoring engine module is used to integrate the anomaly scores obtained in different detection strategies to determine the severity of the detected anomaly and provide appropriate measures; The configuration suggestion feedback module is used to provide accurate and effective configuration suggestions for each measure; The grayscale test module is used to verify the final configuration and ensure that normal users can successfully pass the verification process.

8. A scenario-based self-learning malicious request detection system according to claim 7, characterized in that: The scenario-based basic configuration module includes a scenario-based basic setting unit, a user statistical granularity unit and a detection strategy unit; The scenario-based basic setting unit is used to provide customers with scenarios with different configurations and provide a one-key configuration function for different scenarios; The user statistics granularity unit is used to provide different levels of user identification granularity for customer-configured scenarios; The detection strategy unit is used to provide different types of detection strategies for scenarios configured by customers, and independently set anomaly scores according to the weights of the detection strategies.

9. A scenario-based self-learning malicious request detection system according to claim 8, characterized in that: The detection strategy unit includes a front-end confrontation subunit, an access speed limit subunit, an intelligence analysis subunit and an intelligent analysis subunit; The front-end confrontation sub-unit is used to detect abnormal behaviors such as crawler behaviors, automated attacks, and page debugging, and obtain abnormal scores; The access rate limiting subunit is used to control the access frequency at the user granularity (such as IP address), and to perform anomaly detection through the access frequency and obtain anomaly scores; The intelligence analysis subunit is used to manage the blacklist and whitelist of the threat intelligence database, and to perform anomaly detection through the managed blacklist and whitelist, and obtain anomaly scores; The intelligent analysis subunit is used for access behavior analysis and human-machine behavior analysis, and performs anomaly detection through the behavior analysis to obtain anomaly scores.

10. A scenario-based self-learning malicious request detection system according to claim 9, characterized in that: The configuration suggestion feedback module includes a result display subunit and a traffic self-learning subunit; The result display subunit is used to visualize the results with abnormal scores higher than the threshold, so as to assist customers to quickly identify potential threats and take appropriate measures to mitigate these threats; The traffic self-learning subunit is used to determine the pattern and trend of threat attacks, thereby providing threshold enhancement configuration suggestions for the detection strategy.

Citation Information

Patent Citations

  • Anomaly detection method and device for a network security scene, equipment and medium

    CN112565275A

  • Web crawler detection system based on application scene

    CN115525813A

  • Scenarized threat modeling method based on multi-library fusion

    CN116663022A

  • Scenarized self-learning malicious request detection method and system

    CN117857121A

  • Network anomaly detection apparatus, network anomaly detection system, and network anomaly detection method

    US20200186557A1

Cited By

  • Scenarized self-learning malicious request detection method and system

    CN117857121A

  • A scene-based self-learning malicious request detection method and system

    CN117857121B

  • Multi-scene identity authentication and data sharing system based on real-name DID

    CN120979767A