Abnormity detection method based on multi-dimensional features

Through the anomaly detection method of multi-dimensional features, combined with the preset dimension data of the request log and cluster analysis, the problems of missed detection and false alarm caused by single parameter judgment are solved, and more accurate anomaly detection is achieved.

CN120811698APending Publication Date: 2025-10-17WEBANK (CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511056686.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In the prior art, the method of judging whether a request is malicious based on a single parameter is prone to missed detections and false positives, resulting in low accuracy in identifying abnormal requests.

Method used

An anomaly detection method based on multidimensional features is adopted. By calculating the preset dimension data and evaluation dimension scores of the request log, combined with cluster analysis, disposal instructions are generated, and anomaly judgments are made in multiple dimensions to prevent missed detections and false alarms.

Benefits of technology

It improves the accuracy of anomaly detection, effectively prevents missed detections and false alarms during normal business fluctuations, and improves the accuracy of anomaly detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811698A_ABST
    Figure CN120811698A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer network security, and discloses a multidimensional feature-based anomaly detection method, which comprises the following steps of: judging whether preset dimension data is abnormal or not based on the preset dimension data of a request log and a corresponding threshold value, and judging whether an evaluation dimension score is abnormal or not based on an evaluation dimension score of the request log and a score threshold value; the evaluation dimension score is obtained based on different request identifier numbers, different device numbers and geographical variation coefficients of the request log, and a parameter pair is extracted from the request log for clustering to obtain a cluster; and judging whether the number of the same evaluation dimension parameter values is abnormal or not based on the number corresponding to each same evaluation dimension parameter value in the clustering cluster and a number threshold value, judging whether the clustering cluster is abnormal or not according to a number judgment result, and finally generating a disposal instruction based on the preset dimension data, the evaluation dimension score and whether the clustering cluster is abnormal or not. Abnormality judgment is carried out from multiple dimensions, missed detection and false alarm during normal service fluctuation can be prevented, and the accuracy of an abnormality detection result is high.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer network security, and in particular to an abnormality detection method based on multi-dimensional features. BACKGROUND

[0002] With the development of the Internet, various e-commerce platforms, Internet finance, social new user activities or other preferential activities may be exploited by abnormal users, resulting in platform losses, and ordinary users cannot grab them, failing to achieve the expected marketing effect.

[0003] In order to solve the above problems, the Internet platform usually adopts the following way to identify abnormal attacks: after the request submitted by the user reaches the rule engine, the rule engine judges whether the request is malicious through the preset rule, if yes, the rule engine blocks the malicious request, otherwise, the normal request is released. The preset rule is usually as follows: setting a threshold for the number of requests of a single request IP, if the number of requests of a single request IP exceeds the set threshold, the request is judged to be a malicious request, otherwise, the request is judged to be a normal request; or, judging by matching known malicious value patterns through regular expressions; or, identifying and blocking suspicious devices based on device identifiers.

[0004] The prior art solution is based on a single parameter to judge whether the request is malicious, which is prone to missed detection, and may cause false positives during normal business fluctuations, and the accuracy of identifying abnormal requests is low. SUMMARY

[0005] In order to solve the above problems, the present application provides an abnormality detection method based on multi-dimensional features.

[0006] According to an aspect of an embodiment of the present application, an abnormality detection method based on multi-dimensional features is disclosed, the method comprising:

[0007] calculating preset dimension data of a request log, and judging whether the preset dimension data is abnormal based on the preset dimension data and a corresponding threshold value, wherein the preset dimension includes at least one of request identifier concentration, geographic concentration, device concentration, and traffic mutation degree, the request identifier concentration represents the concentration degree of the request log on the source request identifier, the geographic concentration represents the concentration degree of the request log on the source geographic location, the device concentration represents the concentration degree of the request log on the source device, and the traffic mutation degree represents the traffic deviation degree of the request log;

[0008] calculating an evaluation dimension score of the request log, and judging whether the evaluation dimension score is abnormal based on the evaluation dimension score and a score threshold value, wherein the evaluation dimension score is obtained based on different request identifier numbers, different device numbers, and geographic concentration of the request log;

[0009] extracting parameter pairs from the request logs, the parameter pairs comprising values of a plurality of evaluation dimension parameters, clustering based on the parameter pairs to obtain a cluster, calculating a number of each same evaluation dimension parameter value in the cluster, determining whether the number of each same evaluation dimension parameter value is abnormal based on the number of each same evaluation dimension parameter value and a number threshold, and determining whether the cluster is abnormal according to a number determination result;

[0010] generating a treatment instruction based on the preset dimension data, the evaluation dimension score, and whether the cluster is abnormal.

[0011] In some embodiments, the calculating the evaluation dimension score of the request logs comprises: calculating the evaluation dimension score of the request logs based on a relationship obtaining the evaluation dimension score of the request logs; wherein S is the evaluation dimension score, w1, w2, and w3 are weight coefficients, IP is a number of different request identifiers, τ is a number threshold, Device is a number of different devices, and σ is a geographic variation coefficient. coint is the number of different request identifiers, τ ip is the number threshold, Device count is the number of different devices, is the geographic variation coefficient, σ geo is a spherical distance standard deviation, μ geo is an average spherical distance.

[0012] In some embodiments, the clustering based on the parameter pairs to obtain a cluster comprises: performing a hash operation on each parameter value in the parameter pairs to obtain a hash value corresponding to each parameter value in the parameter pairs, and constructing a log vector of each request log based on the hash value corresponding to each parameter value in the parameter pairs; calculating a Jaccard distance between each pair of the log vectors; and grouping request logs corresponding to the log vectors with a Jaccard distance less than a distance threshold into the same cluster.

[0013] In some embodiments, the plurality of evaluation dimension parameters comprises a request identifier, a device identifier and a monitored parameter carried in the request log. The calculating the number of each same evaluation dimension parameter value in the cluster comprises: calculating the number of different request identifiers in the cluster to obtain a request identifier number, calculating the number of different device identifiers in the cluster to obtain a device number, and calculating the number of requests of the monitored parameter to obtain a request frequency; judging whether the request identifier number is abnormal based on the request identifier number and an identifier number threshold, judging whether the device number is abnormal based on the device number and a device number threshold, and judging whether the request frequency is abnormal based on the request frequency and a frequency threshold; and judging whether the cluster is abnormal based on the judging results of the request identifier number, the device number and the request frequency.

[0014] In some embodiments, the judging whether the cluster is abnormal based on the judging results of the request identifier number, the device number and the request frequency comprises: determining that the cluster is abnormal if the judging results of the request identifier number, the device number and the request frequency are all abnormal; and determining that the cluster is normal if any one of the judging results of the request identifier number, the device number and the request frequency is normal.

[0015] In some embodiments, the number threshold is a dynamic threshold, and the method further comprises: establishing the number threshold by using an exponential weighted moving average algorithm.

[0016] In some embodiments, the calculating the preset dimension data of the request log comprises: calculating the number of different device identifiers in the request log and the number of different request identifiers in the request log; and obtaining the device concentration based on the quotient of the number of different device identifiers and the number of different request identifiers; and the judging whether the preset dimension data is abnormal based on the preset dimension data and a corresponding threshold comprises: judging whether the device concentration is abnormal based on the device concentration and a device concentration threshold.

[0017] In some embodiments, the calculating the preset dimension data of the request log comprises: calculating the quotient of the spherical distance standard deviation and the average spherical distance in the request log to obtain a geographic variation coefficient; and taking the geographic variation coefficient as a geographic concentration; and the judging whether the preset dimension data is abnormal based on the preset dimension data and a corresponding threshold comprises: judging whether the geographic concentration is abnormal based on the geographic concentration and a geographic concentration threshold.

[0018] In some embodiments, the handling instruction comprises one of a release, a global blocking instruction, a start processing instruction; wherein the start processing instruction comprises: limiting a request frequency and triggering a human-computer verification.

[0019] In some embodiments, the preset dimension comprises a request identity concentration, a geographical concentration, a device concentration, and a traffic mutation degree. The generating of the handling instruction based on the preset dimension data, the evaluation dimension score, and whether the clustering cluster is abnormal comprises: generating a global blocking instruction if at least two of the request identity concentration, the geographical concentration, the device concentration, the traffic mutation degree, and the clustering cluster are abnormal; and generating a start processing instruction if any one of the request identity concentration, the geographical concentration, the device concentration, the traffic mutation degree, and the clustering cluster is abnormal and the evaluation dimension score is abnormal.

[0020] The technical scheme provided by the embodiments of the present application at least has the following beneficial effects:

[0021] The scheme disclosed in the present application judges whether the preset dimension data is abnormal based on the preset dimension data of the request log and the corresponding threshold value, judges whether the evaluation dimension score is abnormal based on the evaluation dimension score of the request log and the score threshold value, obtains the evaluation dimension score based on the different request identifier numbers, the different device numbers, and the geographical variation coefficient of the request log, further extracts parameter pairs from the request log to obtain a clustering cluster, judges whether the number of each same evaluation dimension parameter value in the clustering cluster is abnormal based on the number of each same evaluation dimension parameter value in the clustering cluster and the number threshold value, judges whether the clustering cluster is abnormal according to the number judgment result, and finally generates a handling instruction based on whether the preset dimension data, the evaluation dimension score, and the clustering cluster are abnormal. The abnormality is judged from multiple dimensions, which can prevent missed detection and false positives during normal business fluctuations, and the accuracy of the abnormal detection result is high. BRIEF DESCRIPTION OF DRAWINGS

[0022] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate preferred embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.

[0023] Figure 1 A system architecture diagram of an existing abnormal defense system is shown;

[0024] Figure 2 A flowchart of a multi-dimensional feature-based abnormal detection method according to an embodiment of the present application is shown;

[0025] Figure 3 A detailed flowchart of step S201 in an embodiment of the present application is shown; Figure 2

[0026] Figure 4 ​An embodiment of the present application is shown Figure 2 A flow chart showing details of the clustering step of step S203 in the embodiment of the present application

[0027] Figure 5 An embodiment of the present application is shown Figure 2 A flow chart showing details of the clustering step of step S203 in the embodiment of the present application

[0028] Figure 6 A flow chart showing details of the clustering step of step S203 in the embodiment of the present application

[0029] Figure 7 A flow chart showing details of the clustering step of step S203 in the embodiment of the present application

[0030] Figure 8 A flow chart showing details of the clustering step of step S203 in the embodiment of the present application

[0031] The reference signs are explained as follows:

[0032] 700, computer device; 701, processor; 702, memory; 800, computer system; 801,

[0033] CPU; 802, ROM; 803, RAM; 804, bus; 805, I / O interface; 806, input part; 807, output part; 808, storage part; 809, communication part; 810, drive; 811, detachable medium. DETAILED DESCRIPTION

[0034] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these example implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art. Like reference numerals may refer to like elements throughout.

[0035] The terms "first", "second", etc. are used only for the purpose of description, and should not be understood as indicating or implying relative importance or an indicated number of technical features. Thus, features defined with "first", "second", etc. can explicitly or implicitly include one or more features.

[0036] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the application can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the application.

[0037] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further decomposed, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to actual conditions.

[0038] Prior art solutions judge whether a request is malicious based on isolated single parameters, such as Figure 1 As shown, after a user submits an HTTP request for registration or coupon use, the request reaches the rule engine, which analyzes the user device, IP, request parameters, and request mode through various preset rules. Important indicators include IP counters and device identification libraries (fingerprint libraries), which are used to determine whether the request is malicious. Finally, the rule engine blocks malicious requests and releases normal requests.

[0039] Since an attacker can submit a business request using a large number of different proxy IPs, such as in the coupon cancellation scenario, each IP independently submits a request with a coupon code parameter (couponCode = SUMMER2023), which can easily bypass the IP counter detection method, resulting in missed detection of the above solution. The above solution may also produce false positives during normal business fluctuations, such as when a popular coupon is used during a promotion and is mistakenly identified as abnormal (e.g., 100 real users collectively using the same coupon code). The accuracy of identifying abnormal requests is low.

[0040] To this end, the application provides an abnormality detection method based on multi-dimensional features to improve the accuracy of abnormality detection results. The method determines whether the preset dimension data of the request log is abnormal based on the preset dimension data and the corresponding threshold value, determines whether the evaluation dimension score of the request log is abnormal based on the evaluation dimension score and the score threshold value, obtains the evaluation dimension score based on the different request identifier numbers, different device numbers, and geographic variation coefficients of the request log, further extracts parameter pairs from the request log to obtain clustering clusters, determines whether the number of each same evaluation dimension parameter value in the clustering cluster is abnormal based on the number of the same evaluation dimension parameter value and the number threshold value, determines whether the clustering cluster is abnormal according to the number determination result, and finally generates a handling instruction based on whether the preset dimension data, the evaluation dimension score, and the clustering cluster are abnormal. The abnormality is determined from multiple dimensions, which can prevent missed detection and false positives during normal business fluctuations, and the accuracy of the abnormality detection result is high.

[0041] First, some terms involved in the application are explained:

[0042] Web Application Firewall (WAF): a network security system for monitoring and filtering HTTP / HTTPS traffic.

[0043] Parameter Name: the key part of the key-value pair structure in the HTTP request.

[0044] Parameter Value: the value part of the key-value pair structure in the HTTP request.

[0045] Jaccard distance: a method for measuring the similarity between two sets.

[0046] HLL (HyperLogLog): an efficient approximate counting algorithm for estimating the number of different elements in a dataset, particularly suitable for large data scenarios. It can estimate 10^9 independent elements with 1KB of memory with an error of 1.5%.

[0047] HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise): an extension of DBSCAN by converting it into a hierarchical clustering algorithm, then based on clustering stability, using techniques to extract planar clusters. HDBSCAN is a density-based hierarchical clustering algorithm that can handle clusters with different densities.

[0048] Exponentially Weighted Moving-Average (EWMA): A time series forecasting model with the formula: τ t =αx t +(1-α)τ t-1 .

[0049] Geographic dispersion / aggregation: This is an indicator that measures the degree of geographical dispersion / aggregation of request sources, calculated as the standard deviation of coordinate points.

[0050] Cold start protection: A mechanism that uses a conservative strategy to avoid misjudgment during system initialization.

[0051] The following is a detailed description of the implementation details of the technical solution of the embodiment of the present application:

[0052] Figure 2 A flowchart of an anomaly detection method based on multidimensional features according to an embodiment of the present application is shown. Figure 2 As shown, the anomaly detection method includes at least a preset dimension data anomaly detection step, an evaluation dimension score anomaly detection step, a cluster anomaly detection step, and a disposal instruction generation step, which correspond to the following steps S201 to S204, respectively, and are described in detail as follows:

[0053] In step S201 , preset dimension data of the request log is calculated, and based on the preset dimension data and the corresponding threshold, it is determined whether the preset dimension data is abnormal.

[0054] The preset dimensions include at least one of request identifier concentration, geographic concentration, device concentration, and traffic mutation. Request identifier concentration indicates the concentration of request logs around the source request identifier, geographic concentration indicates the concentration of request logs around the source geographic location, device concentration indicates the concentration of request logs around the source device, and traffic mutation indicates the degree of traffic deviation in the request logs.

[0055] In some embodiments, the preset dimension includes request identifier concentration. In step S201, calculating the preset dimension data of the request log includes: calculating the number of different request identifiers in the request log to obtain the request identifier concentration. Determining whether the preset dimension data is abnormal based on the preset dimension data and a corresponding threshold includes: determining whether the request identifier concentration is abnormal based on the request identifier concentration and the identifier concentration threshold.

[0056] The identification concentration threshold value can be an identification number value, which can include an upper limit value and a lower limit value. Accordingly, the determination of whether the request identification concentration is abnormal based on the request identification concentration and the identification concentration threshold value can be a determination of whether the request identification concentration is abnormal according to whether the request identification concentration falls within the interval range formed by the upper limit value and the lower limit value. Specifically, when the request identification concentration falls within the interval range formed by the upper limit value and the lower limit value, it is determined that the request identification concentration is normal; otherwise, it is determined that the request identification concentration is abnormal.

[0057] Of course, in other embodiments, the identification concentration threshold value can also be set to only one upper limit value or only one lower limit value.

[0058] In some embodiments, the identification concentration threshold value is a dynamic threshold value, and an exponential weighted moving average algorithm is used to establish the identification concentration threshold value. Using the exponential weighted moving average algorithm to set the identification concentration threshold value can dynamically weight the historical data, more sensitively reflect the latest changes, effectively smooth the noise, automatically relax the threshold value during the promotion period, and reduce false positives.

[0059] Specifically, the identification concentration threshold value is established based on the following formula:

[0060] τ t = αx t + (1-α)τ t-1

[0061] wherein τ t is the identification concentration threshold value of the current time window, x t is the observation value of the current time window, i.e., the number of different request identification numbers of the current time window, α is a smoothing coefficient (usually 0.2-0.3), and τ t-1 is the identification concentration threshold value of the previous time window.

[0062] In some embodiments, the preset dimension includes geographical concentration. In step S201, calculating the preset dimension data of the request log includes: calculating the quotient of the spherical distance standard deviation and the average spherical distance in the request log to obtain a geographical coefficient of variation; and taking the geographical coefficient of variation as the geographical concentration. Based on the preset dimension data and the corresponding threshold value, determining whether the preset dimension data is abnormal includes: based on the geographical concentration and the geographical concentration threshold value, determining whether the geographical concentration is abnormal.

[0063] In some embodiments, the spherical distance standard deviation is calculated using the following formula:

[0064]

[0065] wherein σ geo is the spherical distance standard deviation, and d ia spherical distance of a source geographical position of a request log to a center point of the source of the request log, an average spherical distance, i.e. n is a number of spherical distance data, i.e. a number of requests of the request log.

[0066] The geographical coefficient of variation is calculated by using the following formula:

[0067]

[0068] wherein, GCV is the geographical coefficient of variation, σ geo is a spherical distance standard deviation, μ geo is an average spherical distance.

[0069] The geographical coefficient of variation is taken as the geographical concentration, and correspondingly, the geographical concentration threshold can be a geographical coefficient of variation threshold. The geographical coefficient of variation represents the degree of dispersion of two groups of data, i.e. the degree of dispersion of the request log in the source geographical position. Based on the geographical concentration and the geographical concentration threshold, whether the geographical concentration is abnormal can be determined. Specifically, when the geographical coefficient of variation is less than the geographical coefficient of variation threshold, it is determined that the request identifier concentration is abnormal; otherwise, it is determined that the request identifier concentration is normal.

[0070] In some embodiments, the preset dimension includes device concentration. In step S201, calculating the preset dimension data of the request log includes: calculating the number of different device identifiers in the request log and the number of different request identifiers in the request log; and obtaining the device concentration based on the quotient of the number of different device identifiers and the number of different request identifiers. Based on the preset dimension data and the corresponding threshold, whether the preset dimension data is abnormal is determined, including: based on the device concentration and the device concentration threshold, whether the device concentration is abnormal is determined.

[0071] In some embodiments, based on the device concentration and the device concentration threshold, whether the device concentration is abnormal can specifically include: subtracting the value 1 from the device concentration to obtain a difference value; comparing the difference value with the device concentration threshold; when the difference value is greater than the device concentration threshold, it is determined that the device concentration is abnormal; otherwise, it is determined that the device concentration is normal.

[0072] In some embodiments, the device concentration threshold is a dynamic threshold, and the exponential weighted moving average algorithm is used to establish the device concentration threshold. Using the exponential weighted moving average algorithm to set the device concentration threshold can dynamically weight the historical data, more sensitively reflect the latest changes, effectively smooth the noise, automatically relax the threshold in the promotion period, and reduce false positives.

[0073] Specifically, the device concentration threshold is established based on the following formula:

[0074] τt = ax t + (1 - a) t t-1

[0075] wherein t t is the device concentration threshold of the current time window, x t is the observation value of the current time window, i.e., the number of different device identifiers in the current time window, a is a smoothing coefficient (usually 0.2-0.3), t t-1 is the device concentration threshold of the previous time window.

[0076] In some embodiments, the preset dimension includes traffic mutation. In step S201, calculating the preset dimension data of the request log includes: calculating the traffic mutation of the request log using the CUSUM (Cumulative Sum) algorithm. Based on the preset dimension data and the corresponding threshold, determining whether the preset dimension data is abnormal includes: based on the traffic mutation and the traffic deviation threshold, determining whether the traffic mutation is abnormal.

[0077] The traffic mutation can include traffic surge and traffic crash, and can be the traffic mutation of the request log calculated using the following formula:

[0078]

[0079] wherein: is the upper deviation cumulative value, used to detect traffic surge, is the lower deviation cumulative value, used to detect traffic crash, x t is the tth observation value, μ is the historical mean value, K is the allowed fluctuation amplitude, usually σ / 2, and σ is the historical standard deviation.

[0080] In the attack detection scenarios such as coupon and invitation registration system, detecting traffic surge is enough. The traffic mutation of the request log is calculated using the CUSUM algorithm, specifically, the traffic upper deviation cumulative value of the request log is calculated using the CUSUM algorithm; based on the traffic mutation and the traffic deviation threshold, determining whether the traffic mutation is abnormal, specifically, when the traffic upper deviation cumulative value is greater than the traffic deviation threshold, it is determined that the traffic mutation is abnormal; otherwise, it is determined that the traffic mutation is normal.

[0081] In other attack detection scenarios, traffic crash can also be detected. Correspondingly, the traffic mutation of the request log is calculated using the CUSUM algorithm, specifically, the traffic lower deviation cumulative value of the request log is calculated using the CUSUM algorithm; based on the traffic mutation and the traffic deviation threshold, determining whether the traffic mutation is abnormal, specifically, when the traffic lower deviation cumulative value is less than the traffic deviation threshold, it is determined that the traffic mutation is abnormal; otherwise, it is determined that the traffic mutation is normal.

[0082] In the above embodiment, the CUSUM algorithm is used to calculate the traffic mutation degree of the request log, which can amplify small changes through cumulative deviation, and can achieve sensitive detection of abnormalities.

[0083] In some embodiments, the preset dimensions can include two or more of the request identifier concentration, the geographic concentration, the device concentration, and the traffic mutation degree, to make abnormality judgment from more dimensions and improve the accuracy of the abnormality detection result. In an exemplary embodiment, the preset dimensions include the request identifier concentration, the geographic concentration, the device concentration, and the traffic mutation degree, as shown in FIG. 1, the preset dimension data abnormality detection step includes steps S301-S304, which are described in detail as follows: Figure 3

[0084] In step S301, the number of different request identifiers in the request log is calculated to obtain the request identifier concentration; and based on the request identifier concentration and the identifier concentration threshold, it is determined whether the request identifier concentration is abnormal.

[0085] In step S302, the quotient of the spherical distance standard deviation and the average spherical distance in the request log is calculated to obtain the geographic variation coefficient; the geographic variation coefficient is taken as the geographic concentration; and based on the geographic concentration and the geographic concentration threshold, it is determined whether the geographic concentration is abnormal.

[0086] In step S303, the number of different device identifiers in the request log and the number of different request identifiers in the request log are calculated; based on the quotient of the number of different device identifiers and the number of different request identifiers, the device concentration is obtained; and based on the device concentration and the device concentration threshold, it is determined whether the device concentration is abnormal.

[0087] In step S304, the CUSUM algorithm is used to calculate the traffic mutation degree of the request log; and based on the traffic mutation degree and the traffic deviation threshold, it is determined whether the traffic mutation degree is abnormal.

[0088] Under normal circumstances, the number of different request identifiers generally tends to be equal to the number of request logs, and under abnormal attack circumstances, the number of different request identifiers can be much smaller than the number of request logs, so the request identifier concentration can be used for abnormality judgment. Similarly, under normal circumstances, the number of different device identifiers generally tends to be equal to the number of request logs, and under abnormal attack circumstances, the number of different device identifiers can be much smaller than the number of request logs, so the device concentration can be used for abnormality judgment. Under normal circumstances, the geographic location of the request source usually does not exhibit obvious geographic concentration, so the geographic concentration can also be used for abnormality judgment. Under normal circumstances, there is no obvious traffic mutation, so the traffic mutation degree of the request log can also be used for abnormality judgment. In the embodiment shown in FIG. 1, dimension data abnormality judgment is made from four dimensions, which helps to improve the accuracy of the abnormality detection result. Figure 3 In the embodiment shown in FIG. 1, dimension data abnormality judgment is made from four dimensions, which helps to improve the accuracy of the abnormality detection result.​

[0089] In step S202 , the evaluation dimension score of the request log is calculated, and based on the evaluation dimension score and the score threshold, it is determined whether the evaluation dimension score is abnormal.

[0090] Among them, the evaluation dimension scores are obtained based on the number of different request identifiers, the number of different devices, and the geographical concentration in the request log.

[0091] In some embodiments, a weighted calculation is performed on the quotient of the number of different request identifiers in the request log and the number threshold, the quotient of the number of different devices and the number of different request identifiers, and the geographic variation coefficient to obtain the evaluation dimension score for the request log. By assigning appropriate weights to each evaluation dimension, the evaluation dimension score can be made more realistic, helping to improve the accuracy of anomaly detection results.

[0092] Specifically, based on the relationship Get the evaluation dimension score of the request log; where S is the evaluation dimension score, w1, w2, w3 are weight coefficients, IP count is the number of different request identifiers, τ ip is the number threshold, Device count For different number of devices, is the coefficient of geographical variation, σ geo is the standard deviation of spherical distance, μ geo is the average spherical distance. For example, w1, w2, and w3 are 0.6, 0.3, and 0.1, respectively.

[0093] In some embodiments, based on the evaluation dimension score and the score threshold, whether the evaluation dimension score is abnormal can be determined by comparing the evaluation dimension score with the score threshold. If the evaluation dimension score is greater than the score threshold, the evaluation dimension score is determined to be abnormal; otherwise, the evaluation dimension score is determined to be normal.

[0094] In step S203, parameter pairs are extracted from the request log, and clustering is performed based on the parameter pairs to obtain clusters, where the parameter pairs include values ​​of multiple evaluation dimension parameters. The number corresponding to each identical evaluation dimension parameter value in the cluster is calculated, and based on the number of identical evaluation dimension parameter values ​​and the number threshold, it is determined whether the number of identical evaluation dimension parameter values ​​is abnormal, and whether the cluster is abnormal is determined based on the number judgment result.

[0095] In some embodiments, as Figure 4 As shown, clustering is performed based on the parameter pairs to obtain cluster clusters, including the following steps S401 to S403, which are detailed as follows:

[0096] In step S401, hash operations are performed on each parameter value in the parameter pair to obtain a hash value corresponding to each parameter value in the parameter pair, and a log vector of each request log is constructed based on the hash value corresponding to each parameter value in the parameter pair.

[0097] In step S402, the Jaccard distance is calculated for all log vectors two by two.

[0098] In step S403, the request logs corresponding to the log vectors with a Jaccard distance less than the distance threshold are classified into the same cluster.

[0099] During the attack, the parameter values of the same parameter in the attack request logs are the same, and there are a large number of same hashes between the attack request logs, such as the same source geographic location hash and the same source device hash. Therefore, the constructed log vectors are highly similar, and the Jaccard distance is very small. Therefore, the request logs corresponding to the log vectors with a Jaccard distance less than the distance threshold can be classified into the same cluster.

[0100] This method can automatically discover abnormal aggregation based on density without prior knowledge of the attack mode, and accurately mark abnormal requests.

[0101] In some embodiments, HDBSCAN clustering is used. In the density hierarchical tree, the request logs with highly similar log vectors are naturally clustered into the same cluster due to a low "mutual distance"; HDBSCAN identifies that the cluster "stably exists" under multiple density thresholds (high cluster stability), and disperses the logs of other normal users as noise or other clusters.

[0102] In some embodiments, the plurality of evaluation dimension parameters include a request identifier, a device identifier, and a monitored parameter carried in the request log. As shown in Figure 4 The number of each same evaluation dimension parameter value in the cluster is calculated, and based on the number of each same evaluation dimension parameter value and the number threshold, it is determined whether the number of each same evaluation dimension parameter value is abnormal, and based on the number determination result, it is determined whether the cluster is abnormal, including the following steps S501-S503, which are described in detail as follows:

[0103] In step S501, the number of different request identifiers in the cluster is calculated to obtain the request identifier number, the number of different device identifiers in the cluster is calculated to obtain the device number, and the number of requests for the monitored parameter is calculated to obtain the request frequency.

[0104] The monitored parameter is the parameter to be abnormally detected, for example, a coupon code.

[0105] The request identification number can be obtained by an accurate calculation or an estimation. For example, in some detection scenarios with small data volume or for high-value parameters, the request identification number can be calculated by Redis to improve the accuracy of the calculation result. In some detection scenarios with large data volume, the request identification number can be obtained by estimation to reduce memory occupation and improve calculation efficiency.

[0106] In an example embodiment, the HyperLogLog (HLL) algorithm is used to estimate the number of different request identifications in the cluster, and the request identification number is obtained by applying a hash function to each source request identification (IP address) to generate a 128-bit hash value, performing bucket processing, setting the number of registers to 2 b (b=4, i.e., 16 buckets), taking the first b bits of the hash value as the bucket number (e.g., taking the first 12 bits to determine the bucket number), and calculating the number of leading zeros ρ (the number of consecutive zeros from the b+1th bit +1) from the remaining bits; each bucket records the current maximum number of leading zeros ρ, and if the newly calculated ρ is greater than the original value in the bucket, the bucket value is updated to ρ; the number of different request identifications in the cluster is estimated by using the harmonic mean correction formula:

[0107]

[0108] wherein E is the number of different request identifications in the cluster, m is the number of buckets (16), α m is a correction factor, and R j is the maximum number of leading zeros +1 recorded in the jth bucket.

[0109] In step S502, based on the request identification number and the identification number threshold, it is determined whether the request identification number is abnormal, based on the device number and the device number threshold, it is determined whether the device number is abnormal, and based on the request frequency and the frequency number threshold, it is determined whether the request frequency is abnormal.

[0110] In some embodiments, the number threshold in step S502 is a dynamic threshold, and each number threshold is established by using an exponential weighted moving average algorithm. By using the exponential weighted moving average algorithm to set the number threshold, the historical data can be dynamically weighted, the latest changes can be more sensitively reflected, noise can be effectively smoothed, the threshold can be automatically relaxed during the promotion period, and false positives can be reduced.

[0111] Specifically, the identification number threshold is established based on the following formula:

[0112] τ t = αx t + (1-α) τ t-1

[0113] wherein τ tx is the threshold of the number of identities for the current time window t is the observed value of the current time window, i.e., the number of different identities for the current time window, a is a smoothing coefficient (usually 0.2-0.3), and τ t-1 is the threshold of the number of identities for the previous time window.

[0114] The threshold of the number of devices is established based on the following formula:

[0115] τ t = a x t + (1-a) τ t-1

[0116] wherein τ t is the threshold of the number of devices for the current time window, x t is the observed value of the current time window, i.e., the number of devices for the current time window, a is a smoothing coefficient (usually 0.2-0.3), and τ t-1 is the threshold of the number of devices for the previous time window.

[0117] The threshold of the number of frequencies is established based on the following formula:

[0118] τ t = a x t + (1-a) τ t-1

[0119] wherein τ t is the threshold of the number of frequencies for the current time window, x t is the observed value of the current time window, i.e., the frequency of requests for the current time window, a is a smoothing coefficient (usually 0.2-0.3), and τ t-1 is the threshold of the number of frequencies for the previous time window.

[0120] In some embodiments, determining whether the number of identities is abnormal based on the number of identities and the threshold of the number of identities can be: calculating the quotient of the number of identities and the threshold of the number of identities to obtain a first multiple, and determining that the number of identities is abnormal if the first multiple is greater than a multiple threshold; otherwise, determining that the number of identities is normal.

[0121] In some embodiments, determining whether the number of devices is abnormal based on the number of devices and the threshold of the number of devices can be: calculating the quotient of the number of devices and the threshold of the number of devices to obtain a second multiple, and determining that the number of devices is abnormal if the second multiple is greater than a multiple threshold; otherwise, determining that the number of devices is normal.

[0122] In some embodiments, based on the request frequency and the frequency number threshold, determining whether the request frequency is abnormal can be: calculating the quotient of the request frequency and the frequency number threshold to obtain a third multiple, and if the third multiple is greater than a multiple threshold, determining that the request frequency is abnormal; otherwise, determining that the request frequency is normal.

[0123] In the foregoing embodiments, the number of evaluation dimension parameters is determined to be abnormal based on the multiple relationship, which can avoid misjudgment caused by small fluctuations in the number of evaluation dimension parameters. Of course, in other embodiments, the number of evaluation dimension parameter values can be directly compared with the number threshold, and whether the number of evaluation dimension parameter values is abnormal can be determined according to the comparison result.

[0124] In step S503, based on the determination results of the request identifier number, the device number and the request frequency, it is determined whether the clustering cluster is abnormal.

[0125] In some embodiments, step S503 includes: if the determination results of the request identifier number, the device number and the request frequency are all abnormal, determining that the clustering cluster is abnormal; and if any one of the determination results of the request identifier number, the device number and the request frequency is normal, determining that the clustering cluster is normal.

[0126] In step S204, based on the preset dimension data, the evaluation dimension score and whether the clustering cluster is abnormal, a treatment instruction is generated.

[0127] In some embodiments, the treatment instruction includes one of a release instruction, a global blocking instruction and a start processing instruction.

[0128] In some embodiments, the start processing instruction includes: limiting the request frequency and triggering the CAPTCHA.

[0129] In some embodiments, the start processing instruction includes: limiting the request frequency, triggering the CAPTCHA and synchronizing the device identifier to all edge nodes.

[0130] By limiting the request frequency and triggering the CAPTCHA, batch attacks by automated tools can be prevented, and abuse behavior can be reduced. By synchronizing the device identifier to all edge nodes, the device identifier can be directly marked as an attack device by subsequent related models, which helps to prevent abnormal attacks.

[0131] In some embodiments, the preset dimensions include request identity concentration, geographical concentration, device concentration, and traffic mutation. In step S204, if at least two of the request identity concentration, the geographical concentration, the device concentration, the traffic mutation, and the cluster are abnormal, a global blocking instruction is generated; if any one of the request identity concentration, the geographical concentration, the device concentration, the traffic mutation, and the cluster is abnormal, and the evaluation dimension score is abnormal, a start processing instruction is generated; and for other cases, no processing is performed, and the request is released.

[0132] It can be understood that the rules for generating the handling instruction can be flexibly set, and specifically, the rules for generating the handling instruction can be flexibly adjusted according to the number of preset dimensions.

[0133] Exemplarily, the preset dimensions include request identity concentration, geographical concentration, and device concentration. In step S204, if any one of the request identity concentration, the geographical concentration, the device concentration, and the cluster is abnormal, and the evaluation dimension score is abnormal, a global blocking instruction is generated; if any one of the request identity concentration, the geographical concentration, the device concentration, and the cluster is abnormal, a start processing instruction is generated; and for other cases, no processing is performed, and the request is released.

[0134] Exemplarily, the preset dimensions include request identity concentration, geographical concentration, and traffic mutation. In step S204, if at least two of the request identity concentration, the geographical concentration, the traffic mutation, and the cluster are abnormal, a global blocking instruction is generated; if any one of the request identity concentration, the geographical concentration, the traffic mutation, and the cluster is abnormal, and the evaluation dimension score is abnormal, a start processing instruction is generated; and for other cases, no processing is performed, and the request is released.

[0135] Next, taking the monitored parameter as a coupon and the attack feature as a coupon verification request submitted by 185 IPs (request identities) and 170 devices in 3 geographical regions as an example, the implementation process of the abnormality detection method of the present application is described in combination with the decision logic diagram shown in FIG. 2. Figure 6

[0136] First, the monitored parameter is configured as {"param_name":"couponCode", "value_type":"string", "baseline_group":"promo"}.

[0137] The exponential weighted moving average algorithm is used to establish a dynamic threshold, and the calculation formula is: τ t = αxt + (1-α)τ t-1 , where τ t is the dynamic threshold of the current time window, x t ​is the observation value of the current time window, α is the smoothing coefficient (usually 0.2-0.3), τ t-1 Dynamic threshold for the previous time window.

[0138] Assume that the number of request identifiers in the current period is 55, the dynamic threshold for the historical period (such as the past 7 days) is 50, set α = 0.25, and calculate the new dynamic threshold (identifier number threshold / identifier concentration threshold): τ t =0.25×55+0.75×50=51.25≈51.

[0139] Assume that the request frequency in the current period is 38, the dynamic threshold in the historical period (such as the past 7 days) is 42, set α = 0.25, and calculate the new dynamic threshold (frequency number threshold): τ t =0.25×38+0.75×42=41.

[0140] Assume that the number of devices in the current period is 110, the dynamic threshold of the historical period (such as the past 7 days) is 80, set α = 0.25, and calculate the new dynamic threshold (device number threshold): τ t =0.25×110+0.75×80=87.5≈88.

[0141] Calculate the evaluation dimension score, the calculation formula is based on the relationship Among them, S is the evaluation dimension score, w1, w2, w3 are weight coefficients, IP count is the number of different request identifiers, τ ip is the number threshold, Device count For different number of devices, is the coefficient of geographical variation, σ geo is the standard deviation of spherical distance, μ geo is the mean spherical distance.

[0142] Assuming the request identification concentration is 185 / 51 = 3.63, the device concentration is 1-170 / 185 = 0.081, and the geographic variation coefficient is 0.35, substitute the formula into: S = 0.6 × 3.63 + 0.3 × 0.081 + 0.1 × 0.35 = 2.18 + 0.024 + 0.035 = 2.239. When S > the score threshold of 1.5, it is determined to be abnormal (triggering an alarm).

[0143] The CUSUM algorithm is used to calculate the traffic mutation degree of the request log. The calculation formula is: in, is the cumulative value of the upper deviation, used to detect traffic surges, x tFor the tth observation, μ is the historical mean, K is the allowed fluctuation range, usually σ / 2, σ is the historical standard deviation.

[0144] Assume the request volume of the couponCode parameter is 50 in 15 minutes, the historical mean μ is 50, and the actual observation x t is 240, and the historical standard deviation σ is 20, calculate the flow mutation degree: When , it is determined to be abnormal (trigger an alarm).

[0145] Each request log extracts parameter pairs (such as couponCode, request identifier, device identifier, geographic location, etc.), and constructs a log vector after MD5 hashing.

[0146] During the attack, the 185 different IPs, 170 devices, and 3 regions have the same parameter values, which are highly similar in the log vector.

[0147] Calculate the Jaccard distance between all log vectors pairwise. Since the attack logs share a large number of identical hashes (the same couponCode, the same geographic location / device identifier hash), the Jaccard distance is very small, representing a "high similarity" cluster.

[0148] In the density hierarchical tree, these 185 highly similar request logs will naturally cluster in the same cluster due to low "mutual distance"; HDBSCAN will identify that this cluster "stably exists" under multiple density thresholds, while other normal user scattered logs are considered noise or other clusters.

[0149] Within the cluster, count the number of request identifiers for the same couponCode, the frequency of couponCode requests, and the number of devices:

[0150] The number of request identifiers in the cluster is 185, the threshold for the number of identifiers is 51, and 185 / 51≈3.63, which is determined to be abnormal; the count of the same couponCode in the cluster is 400, the threshold for the number of frequencies is 41, and 400 / 41≈9.7, which is determined to be abnormal; the number of devices in the cluster is 170, the threshold for the number of devices is 41, and 170 / 88≈9.7, which is determined to be abnormal.

[0151] When the cluster stability is high and the number of request identifiers, the number of devices, and the frequency of requests are all abnormal, it is determined to be an "abuse attack cluster", and all request logs in the cluster are labeled with the same "attack cluster" label, and the cluster stability score and high-frequency features {couponCode=SUMMER2023:185,...} are output for subsequent interception or alarm.

[0152] The abnormal detection data is shown in the following table:

[0153]

[0154] When any two or more models alarm at the same time, a global blocking instruction is generated; when any one model alarms and the evaluation dimension score S is greater than the score threshold, a start processing instruction is generated, wherein the start processing instruction includes: limiting the use frequency of the coupon code to 1 time / minute, triggering the CAPTCHA, and synchronizing the attack device identifier to all edge nodes.

[0155] In summary, the application performs abnormality judgment from multiple dimensions, fully utilizes the request IP, device identifier and geographic information naturally carried in the request log, improves the detection accuracy, and can reduce the false negative rate to 2%; uses a dynamic threshold, and the threshold can be automatically relaxed during the promotion period, which can reduce the false positive rate of the promotion period by 89%, and ensures that the normal business is not affected during the big promotion period.

[0156] Next, referring to Figure 7 The embodiment provides a computer device 700, which comprises one or more processors 701 and a memory 702, and the memory 702 is used for storing one or more programs, when the one or more programs are executed by the one or more processors 701, the computer device 700 realizes the abnormality detection method of the application.

[0157] Figure 8 A structural block diagram of a computer system for implementing some embodiments of the application is shown, and it should be noted that, Figure 8 The computer system shown is only an example, and should not bring any limitation to the function and use range of the embodiments of the application.

[0158] As Figure 8 shown, the computer system 800 comprises a CPU (Central Processing Unit, central processing unit) 801, which can perform various appropriate actions and processes according to the program stored in the ROM (Read-Only Memory, read-only memory) 802 or the program loaded from the storage part 808 to the RAM (Random Access Memory, random access memory) 803, such as the abnormality detection method in the above embodiment. In the RAM 803, various programs and data required for system operation are also stored. The CPU 801, the ROM 802 and the RAM 803 are connected to each other through a bus 804. The I / O (Input / Output, input / output) interface 805 is also connected to the bus 804.

[0159] The following components are connected to the I / O interface 805: an input part 806 including a keyboard, a mouse, etc.; an output part 807 including a display such as a CRT (Cathode Ray Tube), an LCD (Liquid Crystal Display), etc., and a speaker, etc.; a storage part 808 including a hard disk, etc.; and a communication part 809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication part 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as necessary. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 810 as necessary, so that a computer program read therefrom is installed in the storage part 808 as necessary.

[0160] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present application. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing computer programs for executing all or part of the steps shown in the flowcharts in the anomaly detection method. In such embodiments, the computer program can be downloaded and installed from a network by the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, various functions defined in the system of the present application are executed.

[0161] It should be noted that the computer-readable medium in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In this application, the computer-readable signal medium can include a data signal carrying computer-readable computer programs in a baseband or as a part of a carrier wave. Such a propagated data signal can take on various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium that can transmit, propagate or transport programs for use by or in connection with an instruction execution system, device or apparatus. The computer programs contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, or the like, or any suitable combination thereof.

[0162] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In the flowcharts or block diagrams, each block can represent a module, a program segment or a part of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that shown in the drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the involved functions. It should also be noted that each block in the block diagrams or flowcharts, and the combination of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0163] The units described in the embodiments of the present application can be implemented by software, or by hardware, or by a combination of software and hardware. The units described can also be located in a single processor. In some cases, the names of the units do not limit the units themselves.

[0164] As another aspect, the present application provides a computer readable medium, which can be included in the computer device described in the above embodiments, or can exist separately without being assembled into the computer device. The computer readable medium carries one or more programs, which, when executed by the computer device, cause the computer device to implement the method described in the above embodiments.

[0165] It should be noted that although several modules or units for performing actions are mentioned in the above detailed description, the division into the modules or units is not mandatory. In fact, according to the embodiments of the present application, features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, features and functions of one module or unit described above can be further divided into a plurality of modules or units.

[0166] From the above description of the embodiments, those skilled in the art will readily appreciate that the example embodiments described herein can be implemented by software and / or by hardware. Accordingly, the technical solutions of the embodiments of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, or the like) or on a network, and includes a number of instructions for causing a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to perform the methods according to the embodiments of the present application.

[0167] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the application following, in general, the principles of the application and including such departures from the present disclosure as come within known or customary practice in the art to which the application pertains. The specification and examples are to be regarded as exemplary only, and the true scope and spirit of the application are indicated by the appended claims.

Claims

1. A method for anomaly detection based on multidimensional features, characterized in that: include: Calculate preset dimension data of the request log, and determine whether the preset dimension data is abnormal based on the preset dimension data and the corresponding threshold value, wherein the preset dimension includes at least one of request identifier concentration, geographic concentration, device concentration, and traffic mutation degree, the request identifier concentration indicates the concentration degree of the request log on the source request identifier, the geographic concentration indicates the concentration degree of the request log on the source geographic location, the device concentration indicates the concentration degree of the request log on the source device, and the traffic mutation degree indicates the traffic deviation degree of the request log; Calculating an evaluation dimension score for the request log, and determining whether the evaluation dimension score is abnormal based on the evaluation dimension score and a score threshold, wherein the evaluation dimension score is obtained based on the number of different request identifiers, the number of different devices, and the geographic concentration in the request log; Extracting parameter pairs from the request log, clustering the parameter pairs to obtain clusters, wherein the parameter pairs include values ​​of multiple evaluation dimension parameters, calculating the number corresponding to each identical evaluation dimension parameter value in the clusters, judging whether the number of each identical evaluation dimension parameter value is abnormal based on the number of each identical evaluation dimension parameter value and a number threshold, and judging whether the clusters are abnormal based on the number judgment result; Based on the preset dimension data, the evaluation dimension scores, and whether the clusters are abnormal, a handling instruction is generated.

2. The method according to claim 1, characterized in that Calculating the evaluation dimension score of the request log includes: Based on the relation Obtaining an evaluation dimension score of the request log; Among them, S is the evaluation dimension score, w1, w2, w3 are weight coefficients, IP count is the number of different request identifiers, τ ip is the number threshold, Device count is the number of different devices, is the coefficient of geographical variation, σ geo is the standard deviation of spherical distance, μ geo is the mean spherical distance.

3. The method according to claim 1, characterized in that The clustering based on the parameter pairs to obtain cluster clusters includes: Performing a hash operation on each parameter value in the parameter pair to obtain a hash value corresponding to each parameter value in the parameter pair, and constructing a log vector for each request log based on the hash value corresponding to each parameter value in the parameter pair; Calculate the Jaccard distance between all the log vectors; The request logs corresponding to the log vectors whose Jaccard distance is less than a distance threshold are classified into the same cluster.

4. The method according to claim 1, wherein The multiple evaluation dimension parameters include the request identifier, device identifier and monitored parameters carried in the request log; The calculating the number corresponding to each identical evaluation dimension parameter value in the cluster, judging whether the number of each identical evaluation dimension parameter value is abnormal based on the number of each identical evaluation dimension parameter value and a number threshold, and judging whether the cluster is abnormal according to the number judgment result, includes: Calculating the number of different request identifiers in the cluster to obtain the number of request identifiers, calculating the number of different device identifiers in the cluster to obtain the number of devices, and calculating the number of requests for the monitored parameters to obtain the request frequency; Based on the number of request identifiers and the identifier number threshold, determine whether the number of request identifiers is abnormal; based on the number of devices and the device number threshold, determine whether the number of devices is abnormal; based on the request frequency and the frequency number threshold, determine whether the request frequency is abnormal; Whether the cluster is abnormal is determined based on the determination result of the number of request identifiers, the determination result of the number of devices, and the determination result of the request frequency.

5. The method according to claim 4, characterized in that The determining whether the cluster is abnormal based on the determination result of the number of request identifiers, the determination result of the number of devices, and the determination result of the request frequency includes: If the determination result of the number of request identifiers, the determination result of the number of devices, and the determination result of the request frequency are all abnormal, then the cluster is determined to be abnormal; If any one of the determination result of the number of request identifiers, the determination result of the number of devices, and the determination result of the request frequency is normal, then the cluster is determined to be normal.

6. The method according to claim 1, characterized in that The number threshold is a dynamic threshold, and the method further includes: The exponentially weighted moving average algorithm is used to establish the number threshold.

7. The method according to claim 1, characterized in that The preset dimension data of the calculation request log includes: Calculating the number of different device identifiers in the request log and the number of different request identifiers in the request log; Obtaining the device concentration based on a quotient of the number of the different device identifiers and the number of the different request identifiers; The determining whether the preset dimension data is abnormal based on the preset dimension data and the corresponding threshold value includes: determining whether the device concentration is abnormal based on the device concentration and the device concentration threshold value.

8. The method according to claim 1, characterized in that The preset dimension data of the calculation request log includes: Calculating the quotient of the standard deviation of the spherical distance and the average spherical distance in the request log to obtain a geographic variation coefficient; The geographic variation coefficient is taken as the geographic aggregation; The determining whether the preset dimensional data is abnormal based on the preset dimensional data and the corresponding threshold value includes: determining whether the geographic concentration is abnormal based on the geographic concentration degree and the geographic concentration degree threshold value.

9. The method according to any one of claims 1 to 8, characterized in that The handling instruction includes one of a release instruction, a global blocking instruction, and a start processing instruction; The startup processing instruction includes: limiting the request frequency and triggering human-machine verification.

10. The method according to claim 9, characterized in that The preset dimensions include request identification concentration, geographical concentration, device concentration, and traffic mutation; The generating of a handling instruction based on the preset dimension data, the evaluation dimension score, and whether the cluster is abnormal includes: If at least two of the request identification concentration, geographical concentration, device concentration, traffic mutation, and clustering are abnormal, a global blocking instruction is generated; If any one of the request identification concentration, geographic concentration, device concentration, traffic mutation, and clustering is abnormal, and the evaluation dimension score is abnormal, a start processing instruction is generated.