Abnormal behavior detection method and device and electronic equipment

By combining the isolated forest algorithm, the local anomaly factor algorithm, and the graph behavior analysis algorithm, we have achieved accurate detection of both individual and group abnormal behaviors, which solves the problem of insufficient detection accuracy in existing technologies and improves the reliability and flexibility of abnormal behavior detection.

CN121786687APending Publication Date: 2026-04-03CHINA MOBILE INTERNET CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies lack effective methods for detecting abnormal behavior, especially when faced with complex and ever-changing user behavior patterns, resulting in issues of missed and false alarms, and failing to accurately detect abnormal behavior in individuals and groups.

Method used

Unsupervised learning is performed using the Isolation Forest algorithm and the Local Anomaly Factor algorithm, combined with the Graph Behavior Analysis algorithm, to obtain the detection results of individual and group abnormal behaviors through multi-level anomaly detection.

Benefits of technology

It improves the accuracy, reliability, and flexibility of abnormal behavior detection, effectively identifying abnormal behavior in individuals and groups, enhancing communication security, and optimizing service management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786687A_ABST
    Figure CN121786687A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an abnormal behavior detection method and device and electronic equipment. The scheme comprises the following steps: acquiring to-be-detected behavior data of a target user, and acquiring a target behavior feature data set according to the to-be-detected behavior data; carrying out anomaly detection by adopting an isolation forest algorithm to obtain a first detection result, and carrying out anomaly detection by adopting a local anomaly factor algorithm to obtain a second detection result; obtaining an individual abnormal behavior detection result according to the first detection result and the second detection result; and performing anomaly detection by adopting a graph behavior analysis algorithm in response to that the individual abnormal behavior detection result is abnormal, and obtaining a group abnormal behavior detection result. According to the embodiment of the invention, the isolation forest algorithm, the local abnormal factor algorithm and the graph behavior analysis algorithm are effectively combined, double accurate detection of individual abnormal behaviors and group abnormal behaviors is realized through multi-level abnormal detection, and the accuracy, reliability and flexibility of an abnormal behavior detection process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to data processing techniques, and more particularly to a method, apparatus, and electronic device for detecting abnormal behavior. Background Technology

[0002] Driven by technology, market forces, and wireless mobile application services, the wireless mobile communication industry has developed rapidly, and mobile communication products have become necessities for people's daily work and life. As mobile communication has an increasingly significant impact on people's lives, research into abnormal user behavior related to mobile communication products has received increasing attention. The main purpose of researching abnormal user behavior is to improve communication security, optimize service management, and reduce potential risks by identifying and analyzing abnormal mobile phone number usage behaviors.

[0003] However, there is currently no perfect method for detecting abnormal behavior. Therefore, how to accurately and reliably detect abnormal behavior has become a pressing issue that needs to be addressed. Summary of the Invention

[0004] This disclosure addresses some of the shortcomings mentioned in the background art by providing an abnormal behavior detection method, apparatus, and electronic device.

[0005] In a first aspect, embodiments of this disclosure provide an abnormal behavior detection method, comprising: acquiring target user behavior data to be detected, and acquiring a target behavior feature dataset for the target user based on the target behavior data; performing anomaly detection using an isolated forest algorithm based on the target behavior feature dataset to obtain a first detection result, and performing anomaly detection using a local anomaly factor algorithm based on the target behavior feature dataset to obtain a second detection result; acquiring individual abnormal behavior detection results for the target user based on the first detection result and the second detection result; and, in response to the individual abnormal behavior detection result indicating the presence of anomalies, performing anomaly detection using a graph behavior analysis algorithm to obtain group abnormal behavior detection results for the target user, wherein the group abnormal behavior detection results include information about at least one associated user corresponding to the target user.

[0006] In a second aspect, embodiments of this disclosure provide an abnormal behavior detection device, comprising: a behavior feature dataset acquisition unit, configured to acquire target user behavior data to be detected, and acquire a target behavior feature dataset for the target user based on the target behavior data; an intermediate detection result acquisition unit, configured to perform anomaly detection using an isolated forest algorithm based on the target behavior feature dataset to acquire a first detection result, and perform anomaly detection using a local anomaly factor algorithm based on the target behavior feature dataset to acquire a second detection result; an individual abnormal behavior detection result acquisition unit, configured to acquire individual abnormal behavior detection results for the target user based on the first and second detection results; and a group abnormal behavior detection result acquisition unit, configured to perform anomaly detection using a graph behavior analysis algorithm in response to the individual abnormal behavior detection results indicating anomaly, and acquire group abnormal behavior detection results for the target user, wherein the group abnormal behavior detection results include information of at least one associated user corresponding to the target user.

[0007] In a third aspect, embodiments of this disclosure provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the first aspect.

[0008] In a fourth aspect, embodiments of this disclosure provide a processor-readable storage medium storing a computer program for causing a processor to perform the method described in the first aspect.

[0009] In a fifth aspect, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0010] The embodiments provided in this disclosure have at least the following beneficial technical effects: According to an embodiment of this disclosure, an abnormal behavior detection method can obtain target user behavior data to be detected, and obtain a target behavior feature dataset for the target user based on the target behavior data. Then, based on the target behavior feature dataset, an isolation forest algorithm is used for anomaly detection to obtain a first detection result. Based on the target behavior feature dataset, a local anomaly factor algorithm is used for anomaly detection to obtain a second detection result. Further, based on the first and second detection results, individual abnormal behavior detection results of the target user are obtained. Finally, in response to the individual abnormal behavior detection results indicating the existence of anomalies, a graph behavior analysis algorithm is used for anomaly detection to obtain group abnormal behavior detection results of the target user. The group abnormal behavior detection results include information on at least one associated user corresponding to the target user. Therefore, the embodiments of this disclosure no longer rely on fixed static detection rules and supervised learning algorithms with inefficiencies that need improvement. Instead, they effectively combine two unsupervised learning algorithms, namely the Isolation Forest algorithm with more prominent global detection capabilities and the Local Anomaly Factor algorithm with more prominent local detection capabilities, as well as graph algorithms. Through multi-level anomaly detection, they achieve accurate detection of both individual and group abnormal behaviors, improving the accuracy, reliability, flexibility, and robustness of the anomaly detection process. This lays the foundation for improving communication security, optimizing service management, and reducing potential risks.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a flowchart illustrating an abnormal behavior detection method; Figure 2 This is a schematic diagram of part of the sample classification process using the isolated forest algorithm; Figure 3 This is a flowchart illustrating another method for detecting abnormal behavior; Figure 4 This is a flowchart illustrating another method for detecting abnormal behavior; Figure 5 This is a schematic diagram of part of the processing steps for classifying samples using the Local Anomaly Factor algorithm; Figure 6 This is a flowchart illustrating another method for detecting abnormal behavior; Figure 7 This is a flowchart illustrating another method for detecting abnormal behavior; Figure 8This is a flowchart illustrating another method for detecting abnormal behavior; Figure 9 This is a flowchart illustrating another method for detecting abnormal behavior; Figure 10 This is a flowchart illustrating another method for detecting abnormal behavior; Figure 11 This is a schematic diagram of an abnormal behavior detection device. Figure 12 This is a schematic diagram of another abnormal behavior detection device; Figure 13 It is a block diagram of an electronic device. Detailed Implementation

[0013] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present disclosure and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the drawings, not the entire structure.

[0014] It should be noted that in existing technologies, abnormal behavior detection is mainly carried out in the following two ways: detection based on fixed rule settings and detection based on supervised machine learning (ML).

[0015] Methods relying on fixed rules for detection typically use static thresholds to determine if a phone number / its corresponding user behavior is abnormal. For example, specific thresholds are set for call duration, number of SMS messages sent, etc. However, these fixed-rule detection methods are inadequate when dealing with complex, diverse, and ever-changing behavioral patterns, such as criminal groups using a large number of phone numbers. In such cases, detection relying on fixed rules often suffers from numerous false negatives and missed false positives.

[0016] Supervised machine learning-based detection methods suffer from drawbacks such as the need for a large amount of labeled sample data during training and the difficulty in adapting to new abnormal behavior patterns, which also prevent the accuracy of abnormal behavior detection results from being guaranteed.

[0017] Furthermore, existing abnormal behavior detection methods not only fail to meet the accuracy requirements, but also, because they can only focus on processing the abnormal behavior of a single mobile phone number or the user corresponding to that mobile phone number, they are even more helpless in detecting abnormal behavior in groups.

[0018] Therefore, this disclosure proposes an abnormal behavior detection method that effectively combines unsupervised learning algorithms and graph algorithms. Through multi-level anomaly detection, it achieves accurate detection of both individual and group abnormal behaviors, thereby improving the accuracy, reliability, flexibility, and robustness of the abnormal behavior detection process.

[0019] Figure 1 This is a schematic diagram based on the first embodiment of this disclosure. (See diagram below.) Figure 1 As shown, an abnormal behavior detection method proposed in this disclosure embodiment will be explained and described, specifically including the following steps: S101. Obtain the target user's behavior data to be detected, and based on the target user's behavior data, obtain the target behavior feature dataset for the target user.

[0020] Among them, the target user can be any user to be detected; the behavior data to be detected can be the target user's call, SMS, registration, login, and other behavior data; the target behavior feature dataset can be the behavior feature dataset obtained after feature extraction from the behavior data to be detected.

[0021] It should be noted that this disclosure does not impose any restrictions on the specific settings of the target behavioral feature dataset, which can be selected according to the actual situation. For example, the target behavioral feature dataset may include at least call duration, SMS sending volume, active periods, registration frequency, and behavioral periodicity.

[0022] It should also be noted that this disclosure does not limit the specific method for obtaining the target user's behavior data to be detected, and for obtaining the target behavior feature dataset for the target user based on the behavior data to be detected, and can be set according to the actual situation.

[0023] One possible approach is to acquire historical data of the target user within a pre-defined target period, and select behavioral data such as calls, sending and receiving SMS messages, screen activation, registration, and login as the behavioral data to be detected. Furthermore, feature extraction can be performed on the behavioral data to be detected, and at least a portion of the extracted behavioral features (such as call duration, SMS sending volume, active periods, registration frequency, and behavioral periodicity) can be used as the target behavioral feature dataset for the target user, according to pre-defined target feature categories.

[0024] Optionally, call duration feature data can be obtained by extracting features from the target user's call behavior data, such as the cumulative call time of the target user's mobile phone number over a period of time (e.g., the total duration of a day or a week).

[0025] Optionally, the characteristic data of SMS sending volume can be obtained by extracting features from the SMS sending and receiving behavior data of the target user. For example, it can be obtained based on the total number of SMS messages sent by the target user's mobile phone number within a certain period of time.

[0026] Optionally, the active time period feature data can be obtained by extracting features from the screen-lighting behavior data of the target users. The active time period feature data can be the main active time period (e.g., the main active time period is concentrated in the late night, or the main active time period is mainly distributed during commuting time).

[0027] Optionally, registration frequency feature data can be obtained by extracting features from the registration behavior data of target users. The registration frequency feature data can be the frequency of registering multiple accounts in a short period of time.

[0028] Optionally, for behavioral periodicity feature data, at least one of the aforementioned behavioral data can be integrated and analyzed first, and features can be extracted from the behavioral data obtained from the integrated analysis. The behavioral periodicity feature data can be whether there is a specific periodic behavior (such as making a call or sending a text message at a set time every day).

[0029] S102. Based on the target behavior feature dataset, the isolated forest algorithm is used for anomaly detection to obtain the first detection result. Based on the target behavior feature dataset, the local anomaly factor algorithm is used for anomaly detection to obtain the second detection result.

[0030] In this embodiment of the disclosure, the detection of individual abnormal behavior of target users no longer relies on fixed detection parameters, but effectively combines two unsupervised learning algorithms, the Isolation Forest algorithm and the Local Anomaly Factor algorithm, for detection.

[0031] Isolation Forest is a tree-based unsupervised learning algorithm used for anomaly detection. Its core idea is to recursively and randomly select features and split points to separate data; anomalous data is typically easier to segment and thus "isolate" earlier. The algorithm's advantage lies in its efficient processing of high-dimensional data, making it suitable for large-scale detection of anomalies in user phone number behavior.

[0032] The Local Outlier Factor (LOF) algorithm is an unsupervised local anomaly detection algorithm that determines whether a data point is an anomaly by calculating the local density difference between a data point and its neighbors. The core of the LOF algorithm lies in its consideration of not only the anomaly of a data point within the overall context but also its anomaly within a local region. This makes it particularly effective for detecting anomalous phone number behavior within a local population.

[0033] S103. Based on the first detection result and the second detection result, obtain the individual abnormal behavior detection result of the target user.

[0034] In this embodiment, after detection using two unsupervised learning algorithms—the Isolation Forest algorithm and the Local Anomaly Factor algorithm—the detection results obtained from the two methods can be comprehensively evaluated to obtain the individual abnormal behavior detection result of the target user. The Isolation Forest algorithm has a more prominent global detection capability, while the Local Anomaly Factor algorithm has a more prominent local detection capability.

[0035] It should be noted that this disclosure does not limit the specific method for obtaining the individual abnormal behavior detection results of the target user based on the first detection result and the second detection result, and can be set according to the actual situation.

[0036] As one possible implementation, if at least one of the first detection result and the second detection result is abnormal, then the individual abnormal behavior detection result of the target user is determined to be abnormal; correspondingly, if neither the first detection result nor the second detection result is abnormal, then the individual abnormal behavior detection result of the target user is determined to be abnormal.

[0037] As another possible implementation, a pre-defined first weight set can be obtained, and the first and second detection results can be weighted according to this first weight set. Based on the processing result and a pre-defined third threshold, the individual abnormal behavior detection result of the target user can be determined. For example, the pre-defined first weight set might include: a first detection result weight of 0.4 and a second detection result weight of 0.6; a first detection result indicating no abnormality (value 0) and a second detection result indicating an abnormality (value 1). In this case, the processing result would be 0.4. 0+0.6 If 1 = 0.6, reaching the third threshold of 0.6, then the individual abnormal behavior detection result of the target user is determined to be abnormal.

[0038] S104. In response to the individual abnormal behavior detection result indicating the existence of an anomaly, a graph behavior analysis algorithm is used to detect the anomaly and obtain the group abnormal behavior detection result of the target user. The group abnormal behavior detection result includes information of at least one associated user corresponding to the target user.

[0039] It should be noted that users with abnormal individual behavior may often exhibit abnormal group behavior. After obtaining the individual abnormal behavior detection results of the target user by combining two unsupervised learning algorithms, this disclosure can combine graph algorithms to further detect whether there is abnormality in the group behavior of the target users with abnormal individual behavior, and obtain the group abnormal behavior detection results of the target users.

[0040] The abnormal behavior detection results include information on at least one associated user corresponding to the target user. The associated user can be a user with similar abnormal behavior to the target user, or a potential user who may have the possibility of exhibiting similar abnormal behavior to the target user.

[0041] For example, if user A's individual abnormal behavior detection result is abnormal, a graph behavior analysis algorithm can be used for anomaly detection to obtain the group abnormal behavior detection result of the target user. If the group abnormal behavior detection result is abnormal, users with similar abnormal behavior to the target user can be identified as associated users. In this case, the group abnormal behavior detection result includes information such as the associated users' mobile phone numbers and behavioral data. If the group abnormal behavior detection result is not abnormal, potential users who are likely to exhibit similar abnormal behavior to the target user can be identified as associated users. In this case, the group abnormal behavior detection result includes information such as the associated users' mobile phone numbers, behavioral data, and potential anomaly index.

[0042] According to an embodiment of this disclosure, an abnormal behavior detection method can obtain target user behavior data to be detected, and obtain a target behavior feature dataset for the target user based on the target behavior data. Then, based on the target behavior feature dataset, an isolation forest algorithm is used for anomaly detection to obtain a first detection result. Based on the target behavior feature dataset, a local anomaly factor algorithm is used for anomaly detection to obtain a second detection result. Further, based on the first and second detection results, individual abnormal behavior detection results of the target user are obtained. Finally, in response to the individual abnormal behavior detection results indicating the existence of anomalies, a graph behavior analysis algorithm is used for anomaly detection to obtain group abnormal behavior detection results of the target user. The group abnormal behavior detection results include information on at least one associated user corresponding to the target user. Therefore, the embodiments of this disclosure no longer rely on fixed static detection rules and supervised learning algorithms with inefficiencies that need improvement. Instead, they effectively combine two unsupervised learning algorithms, namely the Isolation Forest algorithm with more prominent global detection capabilities and the Local Anomaly Factor algorithm with more prominent local detection capabilities, as well as graph algorithms. Through multi-level anomaly detection, they achieve accurate detection of both individual and group abnormal behaviors, improving the accuracy, reliability, flexibility, and robustness of the anomaly detection process. This lays the foundation for improving communication security, optimizing service management, and reducing potential risks.

[0043] The following examples illustrate in detail the detection process using two unsupervised learning algorithms, the Isolation Forest algorithm and the Local Anomaly Factor algorithm, as well as a graph algorithm.

[0044] To facilitate understanding, the main processing steps for classifying samples using the Isolation Forest algorithm will be explained first.

[0045] like Figure 2 As shown, the Isolation Forest algorithm recursively partitions each data point by constructing multiple binary trees. A data point that is easily separated (i.e., isolated in fewer partitioning steps) is considered an anomaly because its behavior differs from the majority of data points. Specifically, the algorithm randomly selects features (such as call duration or SMS volume for a mobile phone number) and corresponding split points, recursively dividing the data. For the behaviors corresponding to normal users' mobile phone numbers, these behaviors are typically distributed in the middle of most feature values, requiring more partitioning steps to isolate. However, for the behaviors corresponding to anomaly users' mobile phone numbers, because they deviate from the normal distribution, they are often isolated in fewer partitioning steps. The anomaly score range for the Isolation Forest algorithm is typically set to 0-1, with scores closer to 1 indicating a higher likelihood of an anomaly.

[0046] As one possible implementation, such as Figure 3 As shown, based on the above embodiments, the specific process of "using the isolated forest algorithm to perform anomaly detection based on the target behavior feature dataset and obtaining the first detection result" in step 102 includes the following steps: S301. Obtain the first reference behavior feature dataset, wherein the first reference behavior feature dataset consists of reference behavior feature data corresponding to multiple first reference users, and the first reference behavior feature dataset has the same feature category as the target behavior feature dataset.

[0047] It should be noted that the target behavior feature dataset and the first reference behavior feature dataset are the foundation for detection using the Isolation Forest algorithm. Therefore, before using the Isolation Forest algorithm for detection, a feature data set including the target behavior feature dataset and the first reference behavior feature dataset can be constructed. The first reference user can be a user who can serve as a reference object for the target user.

[0048] It should also be noted that, in order to ensure the accuracy of the detection process, the first reference behavioral feature dataset and the target behavioral feature dataset have the same feature categories.

[0049] For example, if the target behavioral feature dataset includes the following five behavioral feature data: call duration, SMS volume, active time period, registration frequency, and behavioral periodicity, then the first reference behavioral feature dataset can also include the following five behavioral feature data: call duration, SMS volume, active time period, registration frequency, and behavioral periodicity.

[0050] S302. Based on the target behavior feature dataset and the first reference behavior feature dataset, the isolation forest algorithm is used to perform anomaly detection and obtain the first detection result.

[0051] As one possible implementation, such as Figure 4 As shown, based on the above embodiments, the specific process of using the isolated forest algorithm to perform anomaly detection and obtain the first detection result according to the target behavior feature dataset and the first reference behavior feature dataset includes the following steps: S401. Based on the target behavior feature dataset and the first reference behavior feature dataset, construct the corresponding binary tree for any feature category.

[0052] It should be noted that, in this disclosure, after obtaining the target behavioral feature dataset and the first reference behavioral feature dataset, a data matrix containing the mobile phone numbers of all users (target users and all first reference users) and their corresponding behavioral features can be constructed first, providing input for subsequent anomaly detection. Specifically, the rows of the matrix can be set to represent each mobile phone number, and the columns can be set to represent different behavioral features.

[0053] Optionally, multiple random binary trees (usually dozens to hundreds) are constructed based on the target behavioral feature dataset and the first reference behavioral feature dataset, with each tree randomly selecting one feature (such as call duration).

[0054] S402. Obtain the split point corresponding to each binary tree, and recursively split the corresponding binary tree according to the split point to obtain the average split depth of the target user in all binary trees.

[0055] In this disclosure, after constructing the corresponding binary tree, random split points can be obtained for each binary tree. Based on these split points, the feature data corresponding to the binary tree is recursively divided into two parts. The tree construction process is repeated until all data points are isolated. Furthermore, after the tree construction is complete, the average split depth of the target user across all binary trees can be obtained.

[0056] S403. Obtain the first detection result based on the average segmentation depth.

[0057] As one possible implementation, a first detection score for the target user can be obtained based on the average segmentation depth. Further, a first threshold can be obtained, and a first detection result can be derived based on the first detection score and the first threshold.

[0058] It should be noted that for each user's mobile phone number, an anomaly score (first detection score) is calculated based on how early it is isolated across multiple trees. The higher the score, the more abnormal the behavior of that mobile phone number.

[0059] As one possible implementation, the average segmentation depth can be standardized, and the first detection score of the target user can be obtained based on the processed data. The first detection score is the score calculated using the isolated forest algorithm.

[0060] For example, abnormal scores S(x,n) The calculation is based on the average segmentation depth of the data point in the tree. h(x) And it is standardized using the following formula:

[0061] Among them, E( h(x) The ) represents the average segmentation depth of the data point. c(n) It is a normalization constant, which is related to the size of the dataset. n Related.

[0062] In this case, abnormal scores S(x,n) The closer a value is to 1, the more likely the user (phone number) is to exhibit abnormal behavior; the closer a value is to 0, the more likely the user (phone number) is to exhibit normal behavior.

[0063] It should also be noted that, in order to distinguish between normal and abnormal user behavior, a threshold corresponding to an abnormality score (first detection score) can be preset. This disclosure does not limit the specific method for setting the first threshold; it can be selected according to the actual situation.

[0064] As one possible approach, an appropriate adaptive threshold can be set based on business needs and the distribution of actual data, through statistical distribution or experimental analysis.

[0065] For example, if the target user belongs to the top 5% of users with the highest first detection score, the target user's first detection result is marked as having an anomaly; if the target user does not belong to the top 5% of users with the highest first detection score, the target user's first detection result is marked as not having an anomaly.

[0066] It should also be noted that the isolated forest algorithm proposed in this disclosure can be widely applied to various anomaly detection scenarios, especially for scenarios such as abnormally active mobile phone number detection, high-frequency operation anomaly detection, and registration attack protection, where it has more significant effects. Specifically, for abnormally active mobile phone number detection, it can detect target users who are frequently active during abnormal time periods (such as late at night); for high-frequency operation anomaly detection, it can detect target users who frequently make calls, send a large number of text messages, or frequently register on multiple platforms in a short period of time; and for registration attack protection, it can identify target users who frequently register on the same or multiple different platforms, especially target users who complete a large number of registrations in a short period of time.

[0067] As can be seen from the above, compared with traditional anomaly detection methods, the isolated forest algorithm used in this embodiment to obtain the first detection result has advantages such as high processing efficiency, no need for massive annotation, and the ability to effectively handle high-dimensional data. Specifically, the isolated forest algorithm has linear time complexity, can run efficiently on large-scale datasets, is suitable for processing large amounts of mobile phone number behavior data, and does not require pre-labeled data, making it suitable for handling scenarios without abnormal labels. It can automatically discover abnormal behavior from the data, making it particularly suitable for detecting constantly changing gray and black market behaviors. Furthermore, the isolated forest algorithm can handle combinations of multi-dimensional behavioral features (such as call duration, SMS sending volume, registration frequency, etc.), greatly improving the flexibility of anomaly detection and further enhancing the accuracy, reliability, flexibility, and robustness of the abnormal behavior detection process.

[0068] To facilitate understanding, the main processing steps for classifying samples using the Local Anomaly Factor (LOF) algorithm will first be explained.

[0069] like Figure 5 As shown, the Local Anomaly Factor (LOFF) algorithm determines whether a data point (a user's phone number) is an anomaly by calculating the local density difference between it and its neighbors. The core of the LEFF algorithm lies in its consideration of not only the anomaly of a data point within the overall context but also its anomaly within a local region. This is particularly effective for detecting anomalous phone number behavior within a local group.

[0070] It should be noted that the local outlier factor algorithm for anomaly detection is mainly based on the concepts of neighborhood, local reachability density (LRD), and local outlier factor (LOF) to identify local anomalies.

[0071] In this context, "neighborhood" refers to defining the surrounding data points of a given data point by setting a radius or a number of nearest neighbors; these are other data points with similar characteristics to the given data point. "Locally accessible density" represents the density of a data point relative to its neighbors. If a data point's density is significantly lower than that of its neighbors, it may be a local anomaly. "Local anomaly factor" calculates the degree of anomaly by comparing a data point's locally accessible density with that of its neighbors. A higher local anomaly factor value indicates a greater likelihood that the data point is an anomaly.

[0072] As one possible implementation, such as Figure 6 As shown, based on the above embodiments, the specific process of "using the local anomaly factor algorithm to perform anomaly detection based on the target behavior feature dataset and obtaining the second detection result" in step 102 includes the following steps: S601. Obtain the second reference behavior feature dataset, wherein the second reference behavior feature dataset consists of reference behavior feature data corresponding to multiple second reference users, and the second reference behavior feature dataset has the same feature categories as the target behavior feature dataset.

[0073] It should be noted that the target behavior feature dataset and the second reference behavior feature dataset form the basis for detection using the Local Anomaly Factor (LOF) algorithm. Therefore, before using the LEF algorithm for detection, a feature dataset including the target behavior feature dataset and the second reference behavior feature dataset can be constructed. The second reference user can be a user that can serve as a reference object for the target user.

[0074] It should also be noted that, in order to ensure the accuracy of the detection process, the second reference behavioral feature dataset and the target behavioral feature dataset have the same feature categories, and multiple second reference users can be completely consistent, partially consistent, or completely inconsistent with multiple first reference users.

[0075] For example, if the target behavioral feature dataset includes the following five behavioral features: call duration, SMS volume, active time period, registration frequency, and behavioral periodicity, then the second reference behavioral feature dataset can also include the following five behavioral features: call duration, SMS volume, active time period, registration frequency, and behavioral periodicity.

[0076] S602. Based on the target behavior feature dataset and the second reference behavior feature dataset, use the local anomaly factor algorithm to perform anomaly detection and obtain the second detection result.

[0077] As one possible implementation, such as Figure 7As shown, based on the above embodiments, the specific process of using the local anomaly factor algorithm to perform anomaly detection and obtain the second detection result according to the target behavior feature dataset and the second reference behavior feature dataset includes the following steps: S701. Based on the target behavioral feature dataset and the second reference behavioral feature dataset, obtain the similarity between the target user and each second reference user.

[0078] It should be noted that this disclosure does not limit the specific method for obtaining the similarity between the target user and each second reference user based on the target behavioral feature dataset and the second reference behavioral feature dataset, and can be set according to the actual situation.

[0079] As one possible approach, the similarity between the target user and each second reference user can be obtained through metrics such as Euclidean distance and cosine similarity.

[0080] S702. Based on all similarities, select at least one neighbor of the target user from all second reference users, and obtain the local reachability density between the target user and all its neighbors.

[0081] Here, "neighbors" refers to finding the k nearest neighbors of the target user (k is an integer greater than or equal to 1) based on similarity. These neighbors are a set of data points that are most similar to the target user in terms of behavioral characteristics.

[0082] It should be noted that this disclosure does not limit the specific method of selecting at least one neighbor of the target user from all second reference users based on all similarities, and can be set according to the actual situation.

[0083] For example, if the behavioral characteristics of target user A are (call duration, SMS sending volume, registration frequency) = (300 seconds, 5 messages, 2 times), and the behavioral characteristics of second reference users B, C, and D are (290 seconds, 6 messages, 1 time), (310 seconds, 4 messages, 2 times), and (20 seconds, 1 message, 1 time), respectively, then second reference users B and C might be defined as neighbors of target user A, while second reference user C is not a neighbor of target user A.

[0084] Furthermore, after identifying the target user's neighbors, the local reachability density between the target user and all its neighbors can be obtained.

[0085] Local reachability density refers to the reciprocal of the average distance between a data point and its neighbors, representing the density of a point relative to its neighborhood group. LRD(A) It can be obtained using the following formula:

[0086] in, A This is the current data point. i yes A Any of its neighbors.

[0087] For example, consider two neighbors of target user A: second reference user B and second reference user C. The distance from target user A to second reference user B is 10, and the distance from target user A to second reference user C is 12. In this case, what is the local reachability density of target user A? LRD(A) for:

[0088] S703. Obtain the second detection result based on the locally accessible density.

[0089] As one possible implementation, a second detection score for the target user can be obtained based on the local reachability density. Furthermore, a second threshold can be obtained, and a second detection result can be derived based on the second detection score and the second threshold.

[0090] It should be noted that if a user's mobile phone number has a significantly lower local reachability density than its neighbors, it is likely an outlier. Optionally, a local anomaly factor (second detection score) can be calculated based on the local reachability density.

[0091] One possible implementation is to obtain the average local reachability density of all neighbors and then compare it with the local reachability density of the target user. LRD(A) The quotient is used as the second test score LOF(A) Specifically, it can be obtained using the following formula:

[0092] Among them, the second test score LOF(A) A value greater than 1 indicates that the data point is an outlier, and the larger the value, the stronger the outlier the point. LOF(A) A value close to 1 indicates that the user's behavior is similar to that of their neighbors and is not an outlier.

[0093] For example, if the local reachability density of target user A... LRD(A) The value is 0.05, and the average local reachability density of the second reference user B and the second reference user C is 0.1. In this case, LOF(A) =0.1 / 0.05=2, which means LOF(A) The value of 2 indicates that target user A is a local outlier.

[0094] The second threshold refers to the threshold corresponding to the pre-set second detection score. For example, the second threshold can be set to 1.5 or 2.

[0095] It should be noted that the second threshold proposed in this disclosure is not a fixed data, but a dynamic adaptive threshold obtained by comprehensively considering factors such as application scenarios and business needs.

[0096] For example, for application scenarios with high risk levels, a second threshold can be set to 1.5. In this case, if the second detection score is greater than or equal to 1.5, the second detection result of the target user is marked as abnormal. For application scenarios with low risk levels, a second threshold can be set to 2. In this case, if the second detection score is greater than or equal to 2, the second detection result of the target user is marked as abnormal.

[0097] It should also be noted that the local anomaly factor algorithm proposed in this disclosure can be widely applied to various anomaly detection scenarios, particularly for anomaly detection of small-scale gray and black market groups and the detection of local high-frequency operations, where it has more significant effects. Specifically, for anomaly detection of small-scale gray and black market groups, these groups may exhibit anomalies in certain local behaviors, but these behaviors are not obvious globally; the local anomaly factor algorithm can effectively capture these local anomalies. For the detection of local high-frequency operations, it can detect phone numbers that stand out compared to neighboring phone numbers during specific time periods or specific behaviors. For example, some users' phone numbers frequently register within a specific time period, while other users in the same group do not exhibit the same behavior.

[0098] As can be seen from the above, compared with traditional anomaly detection methods, the embodiments of this disclosure employ a local anomaly factor algorithm to obtain the second detection result, which has advantages such as strong local detection capability, high flexibility, and strong adaptability. Specifically, the local anomaly factor algorithm pays more attention to changes in local data distribution, and can effectively capture mobile phone numbers that behave abnormally in specific groups, which is particularly useful for detecting small-scale gangs in the gray and black market. At the same time, it can not only detect overall anomalies, but also identify mobile phone numbers whose behavior seems normal globally, but is abnormal in local environments. For example, in a certain type of behavioral group, the behavior of some mobile phone numbers deviates significantly from the normal range, which can be effectively detected by the local anomaly factor algorithm. Furthermore, the local anomaly factor algorithm can flexibly adjust the second threshold according to different application scenarios, thereby improving the accuracy and recall rate of detection, and further enhancing the accuracy, reliability, flexibility, and robustness of the abnormal behavior detection process.

[0099] For detection using graph behavior analysis algorithms, to facilitate understanding, we will first explain the main processing steps of classifying samples using graph behavior analysis algorithms.

[0100] As one possible implementation, such as Figure 8 As shown, based on the above embodiments, the specific process of "using graph behavior analysis algorithms to perform anomaly detection and obtain the group abnormal behavior detection results of target users" in step 104 includes the following steps: S801. Obtain the third reference behavior feature dataset, wherein the third reference behavior feature dataset consists of reference behavior feature data corresponding to multiple third reference users, and the third reference behavior feature dataset has the same feature categories as the target behavior feature dataset.

[0101] It should be noted that the target behavior feature dataset and the third reference behavior feature dataset are the foundation for detection using graph behavior analysis algorithms. Therefore, before using graph behavior analysis algorithms for detection, a feature data set including the target behavior feature dataset and the third reference behavior feature dataset can be constructed. The third reference user can be a user that can serve as a reference object for the target user.

[0102] It should also be noted that, in order to ensure the accuracy of the detection process, the third reference behavioral feature dataset and the target behavioral feature dataset have the same feature categories, and multiple third reference users can be completely consistent, partially consistent, or completely inconsistent with multiple first reference users and multiple second reference users.

[0103] For example, if the target behavioral feature dataset includes the following five behavioral features: call duration, SMS volume, active time period, registration frequency, and behavioral periodicity, then the third reference behavioral feature dataset can also include the following five behavioral features: call duration, SMS volume, active time period, registration frequency, and behavioral periodicity.

[0104] S802. Based on the target behavior feature dataset and the third reference behavior feature dataset, use the graph behavior analysis algorithm to perform anomaly detection and obtain the group abnormal behavior detection results.

[0105] As one possible implementation, such as Figure 9 As shown, based on the above embodiments, the specific process of using a graph behavior analysis algorithm to perform anomaly detection based on the target behavior feature dataset and the third reference behavior feature dataset, and obtaining the group abnormal behavior detection results, includes the following steps: S901. Construct a target undirected graph with the target user and all third reference users as nodes, and the line connecting any two nodes with related behaviors as edges.

[0106] It should be noted that this disclosure does not impose any restrictions on the specific settings of related behaviors, and choices can be made according to the actual situation.

[0107] As one possible implementation, associated behaviors can be defined to include at least the following: simultaneous registration on the same platform, frequent dialing of the same number, and similar activity times. Specifically, simultaneous registration on the same platform refers to multiple users' corresponding mobile phone numbers registering on the same platform or application within a similar timeframe, potentially indicating that these numbers are controlled by the same group; frequent dialing of the same number refers to multiple users' corresponding mobile phone numbers frequently dialing the same phone number or jointly contacting a group of numbers, suggesting they may be performing the same task; and similar activity times refer to users' mobile phone numbers exhibiting consistent behavior within the same time period, such as frequent activity late at night, which may indicate collective action by a group.

[0108] It should also be noted that the edges in the target undirected graph in this disclosure have corresponding weights, and the weights of the edges represent the strength of the association. The target undirected graph refers to a weighted undirected graph.

[0109] For example, considering target user A, third reference user B, and third reference user C, the following behavioral associations exist between them: Target user A and third reference user B frequently register on multiple identical platforms within similar time periods, with a correlation degree of 0.8; Third reference user B and third reference user C frequently contact the same batch of phone numbers, with a correlation degree of 0.6; Target user A and third reference user C have no obvious behavioral association, therefore there are no edges. In this case, the following weighted undirected graph (target undirected graph) can be constructed: A --(0.8)-- B, B --(0.6)-- C.

[0110] S902. Obtain the abnormal behavior detection results of the group based on the target undirected graph.

[0111] As one possible implementation, such as Figure 10 As shown, based on the above embodiments, the specific process of obtaining the detection results of abnormal group behavior according to the target undirected graph includes the following steps: S1001. Perform vector transformation on each node in the target undirected graph to obtain the embedding vector of each node.

[0112] In this embodiment of the disclosure, after obtaining the target undirected graph, each node in the target undirected graph can be processed by vector transformation, that is, each node in the target undirected graph is converted into a vector representation to obtain the embedding vector of each node.

[0113] It should be noted that this disclosure does not limit the specific settings for performing vector transformation processing on each node in the target undirected graph to obtain the embedding vector of each node, and can be selected according to the actual situation.

[0114] One possible implementation is the DeepWalk algorithm. DeepWalk is a classic algorithm for learning network representations, used to learn the vector representations of vertices in a network. Specifically, the DeepWalk algorithm captures the relationships between nodes by performing random walks on the graph. Starting from each node, the algorithm performs random walks along edges, recording the paths between nodes, and thus learning the feature representations of the nodes. These feature representations reflect the relationships between the nodes and their neighbors, as well as their positions in the graph.

[0115] Optionally, multiple random walks can be performed starting from each node to capture the node's neighborhood structure. Each walk generates a path, and the sequence of nodes along the path can help identify the node's local structure. Furthermore, the embedding vector for each node can be derived by processing the paths generated by these random walks. Nodes with shorter paths to each other will have more similar embedding vectors.

[0116] Therefore, the DeepWalk algorithm can be used to convert users' phone numbers into a set of high-dimensional vectors, representing their positions in the graph and their relationships with other nodes.

[0117] As another possible implementation, the GraphSAGE (graph sampling aggregation) algorithm can be used. GraphSAGE refers to an inductive graph neural network algorithm framework, primarily used to generate embedded representations of nodes in a graph. It solves the problem that traditional graph embedding methods struggle to generate embeddings for unfamiliar nodes when dealing with dynamic or large-scale graphs. Specifically, the GraphSAGE algorithm generates node representations by aggregating information about node neighbors. It learns the vector representation of each node based on the features of its neighbors, enabling each node to express not only its own features but also the behavioral information of its neighbors.

[0118] Optionally, a fixed number of neighboring nodes can be sampled from each node, and the features of the neighboring nodes can be aggregated (e.g., by taking the average or maximum value) to generate a feature representation of the node. Further, the final embedding vector of the node can be generated by aggregating the features of the neighboring nodes, representing the node's relative position and importance in the entire graph.

[0119] For example, using the GraphSAGE algorithm, for target user A, third reference user B, and third reference user C, three corresponding embedding vectors can be obtained: A: 0.2, 0.8, 0.5, B: 0.1, 0.9, 0.6, C: 0.3, 0.7, 0.4.

[0120] Therefore, by using the GraphSAGE approach, the association behavior and relationships of each user's mobile phone number in the graph can be represented by the embedding vector of each node.

[0121] S1002. Perform clustering on all embedded vectors to obtain the clustering results.

[0122] In this embodiment of the disclosure, after obtaining the embedding vector of each node, a clustering algorithm can be used to analyze these embedding vectors to obtain clustering results, so as to identify groups of mobile phone numbers with closely related behaviors, which may reveal gray and black market gangs.

[0123] It should be noted that this disclosure does not limit the specific method for clustering all embedded vectors to obtain the clustering results, and the method can be selected according to the actual situation.

[0124] One possible implementation is the K-Means clustering algorithm. K-Means is an iterative clustering analysis algorithm that divides the embedding vectors of nodes into k clusters to find the most closely related groups of users' phone numbers, and identifies anomalous groups by minimizing the distance from the node embedding vector to the cluster center.

[0125] Optionally, k initial cluster centers can be randomly selected, and the embedding vector of each user's mobile phone number can be assigned to the nearest cluster center. Further, the average value of the node embedding vectors in each cluster can be calculated as the new cluster center, and the process of assigning nodes and updating cluster centers can be repeated until the clustering results stabilize.

[0126] Therefore, by using the K-Means algorithm, we can assign mobile phone numbers of users whose embedding vectors are similar to each other to the same cluster, indicating that these users' mobile phone numbers have unusually similar behaviors and may belong to the same gray and black market gang.

[0127] As another possible implementation, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm can be used. DBSCAN is a density-based clustering algorithm suitable for detecting clusters of arbitrary shapes, particularly well-suited for irregular or noisy data. This algorithm defines core and boundary nodes based on the number and distance of a node's neighbors, forming clusters of high-density regions.

[0128] Optionally, a radius ε and a minimum number of points MinPts can be defined for each node, and other nodes can be found within its neighborhood. Further, if a node has more than MinPts neighbors, it is defined as a core node. Further, starting from the core node, the cluster is expanded to include its neighboring nodes in the same cluster. In this case, nodes not included in any cluster are considered noise nodes, which may be isolated abnormal user phone numbers.

[0129] Therefore, using the DBSCAN algorithm, if target user A, third reference user B, and third reference user C are assigned to the same high-density cluster, it indicates that they have highly correlated behaviors and may form a gray and black market gang.

[0130] S1003. Based on the clustering results, determine at least one associated user corresponding to the target user, and obtain the abnormal behavior detection results of the group based on the information of all associated users.

[0131] In this embodiment of the disclosure, based on the clustering results, groups of mobile phone numbers with closely related behaviors can be identified. These groups may be controlled by gray and black market gangs. By analyzing the mobile phone numbers in the clusters, the common behavioral patterns of the gangs can be identified, such as jointly contacting the same number or registering multiple times in similar time periods.

[0132] It should be noted that, in order to prevent and control detected abnormal behavior in a timely and effective manner, corresponding prevention and control measures can be implemented when abnormal individual behavior and / or abnormal group behavior of target users are detected.

[0133] As one possible implementation, in response to the detection result of individual abnormal behavior being found to be abnormal, and / or in response to the detection result of group abnormal behavior being found to be abnormal, prevention and control measures are performed, wherein the prevention and control measures include at least one of error reporting, blocking, and requiring identity authentication.

[0134] It should be noted that this disclosure does not limit the specific methods for implementing prevention and control measures, and the appropriate methods can be selected based on the actual situation.

[0135] For example, for application scenarios with high risk levels, the same prevention and control measures can be set for both individual abnormal behavior detection results and group abnormal behavior detection results at different stages, in order to eliminate any form of abnormal behavior as much as possible; for application scenarios with low risk levels, corresponding, different and targeted prevention and control measures can be set for both individual abnormal behavior detection results and group abnormal behavior detection results at different stages.

[0136] Furthermore, in order to improve the effectiveness and flexibility of the anomaly detection process, the detection process using two unsupervised learning algorithms, the Isolation Forest algorithm and the Local Anomaly Factor algorithm, as well as the graph algorithm, can be optimized and iteratively processed.

[0137] As one possible implementation, feedback information can be obtained and optimization processing can be performed based on the feedback information. The optimization processing includes optimizing at least one of the first reference behavioral feature dataset, the second reference behavioral feature dataset, and the third reference behavioral feature dataset.

[0138] For example, individual and group abnormal behavior detection results can be combined with feedback from real-world business scenarios to adjust model parameters. For instance, adjusting the number of binary trees in the Isolation Forest algorithm or the number of neighbors in the Local Anomaly Factor algorithm can make the algorithm more adaptable to the data distribution of the actual scenario. Furthermore, continuous optimization can be achieved through iteration. After each anomaly detection, feature extraction and hyperparameters can be fine-tuned based on feedback to improve detection accuracy and recall.

[0139] Furthermore, to improve the reliability of the anomaly detection process, the original behavioral data of the target user can be preprocessed before acquiring the target user's behavioral data to be detected.

[0140] One possible approach is to acquire the target user's raw behavioral data and target preprocessing strategy, preprocess the raw behavioral data according to the target preprocessing strategy, and use the processed data as the target user's behavioral data to be detected.

[0141] For example, raw behavioral data of target users and a target preprocessing strategy can be obtained. The raw behavioral data can then be cleaned according to the target preprocessing strategy, including handling missing values, removing duplicate data, and eliminating outliers. In this case, if behavioral data for some users' phone numbers is missing, or if there are abnormally high call durations or unusually high registration frequency, appropriate data can be filled in or removed.

[0142] Corresponding to the abnormal behavior detection method provided in the above embodiments, an embodiment of this disclosure also provides an abnormal behavior detection device. Since the abnormal behavior detection device provided in this disclosure corresponds to the abnormal behavior detection method provided in the above embodiments, the implementation of an abnormal behavior detection method is also applicable to the abnormal behavior detection device provided in this embodiment, and will not be described in detail in this embodiment.

[0143] Figure 11-12 This is a schematic diagram of an abnormal behavior detection device according to an embodiment of the present disclosure.

[0144] like Figure 11 As shown, the abnormal behavior detection device 2000 includes: a behavior feature dataset acquisition unit 1110, an intermediate detection result acquisition unit 1120, an individual abnormal behavior detection result acquisition unit 1130, and a group abnormal behavior detection result acquisition unit 1140.

[0145] The system includes a behavior feature dataset acquisition unit 1110, which acquires the target user's behavior data to be detected and, based on the behavior data, acquires a target behavior feature dataset for the target user; an intermediate detection result acquisition unit 1120, which performs anomaly detection using the isolated forest algorithm based on the target behavior feature dataset to acquire a first detection result and performs anomaly detection using the local anomaly factor algorithm based on the target behavior feature dataset to acquire a second detection result; an individual abnormal behavior detection result acquisition unit 1130, which acquires the target user's individual abnormal behavior detection result based on the first and second detection results; and a group abnormal behavior detection result acquisition unit 1140, which, in response to the individual abnormal behavior detection result indicating the presence of anomalies, performs anomaly detection using a graph behavior analysis algorithm to acquire the target user's group abnormal behavior detection result, wherein the group abnormal behavior detection result includes information about at least one associated user corresponding to the target user.

[0146] According to an embodiment of this disclosure, the intermediate detection result acquisition unit 1120 is further configured to: acquire a first reference behavior feature dataset, wherein the first reference behavior feature dataset consists of reference behavior feature data corresponding to multiple first reference users, and the first reference behavior feature dataset has the same feature category as the target behavior feature dataset; and perform anomaly detection using the isolated forest algorithm based on the target behavior feature dataset and the first reference behavior feature dataset to acquire a first detection result.

[0147] According to an embodiment of this disclosure, the intermediate detection result acquisition unit 1120 is further configured to: construct a corresponding binary tree for any feature category based on the target behavior feature dataset and the first reference behavior feature dataset; obtain the segmentation point corresponding to each binary tree, and perform recursive segmentation processing on the corresponding binary tree based on the segmentation point to obtain the average segmentation depth of the target user in all binary trees; and obtain the first detection result based on the average segmentation depth.

[0148] According to one embodiment of this disclosure, the intermediate detection result acquisition unit 1120 is further configured to: acquire a first detection score of the target user based on the average segmentation depth; acquire a first threshold, and acquire a first detection result based on the first detection score and the first threshold.

[0149] According to an embodiment of this disclosure, the intermediate detection result acquisition unit 1120 is further configured to: acquire a second reference behavior feature dataset, wherein the second reference behavior feature dataset consists of reference behavior feature data corresponding to multiple second reference users, and the second reference behavior feature dataset has the same feature category as the target behavior feature dataset; and perform anomaly detection using a local anomaly factor algorithm based on the target behavior feature dataset and the second reference behavior feature dataset to acquire a second detection result.

[0150] According to one embodiment of this disclosure, the intermediate detection result acquisition unit 1120 is further configured to: acquire the similarity between the target user and each second reference user based on the target behavior feature dataset and the second reference behavior feature dataset; select at least one neighbor of the target user from all the second reference users based on all the similarities, and acquire the local reachability density between the target user and all the neighbors; and acquire the second detection result based on the local reachability density.

[0151] According to one embodiment of this disclosure, the intermediate detection result acquisition unit 1120 is further configured to: acquire a second detection score of the target user based on the local reachability density; acquire a second threshold, and acquire a second detection result based on the second detection score and the second threshold.

[0152] According to one embodiment of this disclosure, the group abnormal behavior detection result acquisition unit 1140 is further configured to: acquire a third reference behavior feature dataset, wherein the third reference behavior feature dataset consists of reference behavior feature data corresponding to multiple third reference users, and the third reference behavior feature dataset has the same feature category as the target behavior feature dataset; and perform anomaly detection using a graph behavior analysis algorithm based on the target behavior feature dataset and the third reference behavior feature dataset to acquire the group abnormal behavior detection result.

[0153] According to one embodiment of this disclosure, the group abnormal behavior detection result acquisition unit 1140 is further configured to: construct a target undirected graph with the target user and all third reference users as nodes and the line connecting any two nodes with related behaviors as edges; and acquire the group abnormal behavior detection result based on the target undirected graph.

[0154] According to one embodiment of this disclosure, the group abnormal behavior detection result acquisition unit 1140 is further configured to: perform vector transformation processing on each node in the target undirected graph to obtain the embedding vector of each node; perform clustering processing on all the embedding vectors to obtain the clustering result; determine at least one associated user corresponding to the target user based on the clustering result, and acquire the group abnormal behavior detection result based on the information of all associated users.

[0155] According to one embodiment of this disclosure, such as Figure 12 As shown, Figure 11 The abnormal behavior detection device 2000 further includes a prevention and control processing unit 1150. The prevention and control processing unit 1150 is used to: in response to the detection result of an individual abnormal behavior being abnormal, and / or, in response to the detection result of a group abnormal behavior being abnormal, perform prevention and control processing, wherein the prevention and control processing includes at least one of error reporting, blocking, and requiring identity authentication.

[0156] According to one embodiment of this disclosure, such as Figure 12 As shown, Figure 11 The abnormal behavior detection device 2000 further includes an optimization processing unit 1160. The optimization processing unit 1160 is used to: acquire feedback information and, based on the feedback information, perform optimization processing, wherein the optimization processing includes optimizing at least one of a first reference behavior feature dataset, a second reference behavior feature dataset, and a third reference behavior feature dataset.

[0157] According to one embodiment of this disclosure, the behavior feature dataset acquisition unit 1110 is further configured to: acquire the original behavior data of the target user and the target preprocessing strategy; preprocess the original behavior data according to the target preprocessing strategy, and use the processed data as the target user's behavior data to be detected.

[0158] An abnormal behavior detection device according to an embodiment of this disclosure can acquire target user behavior data to be detected, and acquire a target behavior feature dataset for the target user based on the target behavior data. Then, based on the target behavior feature dataset, an isolation forest algorithm is used for anomaly detection to obtain a first detection result. Based on the target behavior feature dataset, a local anomaly factor algorithm is used for anomaly detection to obtain a second detection result. Further, based on the first and second detection results, an individual abnormal behavior detection result for the target user is obtained. Finally, in response to the individual abnormal behavior detection result indicating the existence of anomalies, a graph behavior analysis algorithm is used for anomaly detection to obtain a group abnormal behavior detection result for the target user. The group abnormal behavior detection result includes information about at least one associated user corresponding to the target user. Therefore, the embodiments of this disclosure no longer rely on fixed static detection rules and supervised learning algorithms with inefficiencies that need improvement. Instead, they effectively combine two unsupervised learning algorithms, namely the Isolation Forest algorithm with more prominent global detection capabilities and the Local Anomaly Factor algorithm with more prominent local detection capabilities, as well as graph algorithms. Through multi-level anomaly detection, they achieve accurate detection of both individual and group abnormal behaviors, improving the accuracy, reliability, flexibility, and robustness of the anomaly detection process. This lays the foundation for improving communication security, optimizing service management, and reducing potential risks.

[0159] The method and apparatus are based on the same concept of the application. Since the methods and apparatus solve problems in similar ways, the implementation of the apparatus and methods can refer to each other, and the repeated parts will not be described again.

[0160] It should be noted that the division of units in the embodiments of this disclosure is illustrative and only represents one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.

[0161] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0162] According to embodiments of this disclosure, this disclosure also provides an electronic device 3000, such as... Figure 13 As shown, it includes a memory 400, a processor 500, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the aforementioned abnormal behavior detection method.

[0163] According to embodiments of this disclosure, a processor-readable storage medium is also provided. This processor-readable storage medium stores a computer program that causes the processor to perform the aforementioned abnormal behavior detection method.

[0164] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0165] According to embodiments of this disclosure, this disclosure also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the aforementioned abnormal behavior detection method.

[0166] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0167] The specific embodiments described herein do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for detecting abnormal behavior, characterized in that, include: Obtain the target user's behavior data to be detected, and based on the target user's behavior data, obtain the target behavior feature dataset for the target user; Based on the target behavior feature dataset, the isolated forest algorithm is used for anomaly detection to obtain a first detection result, and based on the target behavior feature dataset, the local anomaly factor algorithm is used for anomaly detection to obtain a second detection result. Based on the first detection result and the second detection result, obtain the individual abnormal behavior detection result of the target user; In response to the individual abnormal behavior detection result indicating the presence of anomalies, a graph behavior analysis algorithm is used to perform anomaly detection to obtain the group abnormal behavior detection result of the target user, wherein the group abnormal behavior detection result includes information of at least one associated user corresponding to the target user.

2. The method according to claim 1, characterized in that, The step of performing anomaly detection using the isolated forest algorithm based on the target behavior feature dataset to obtain a first detection result includes: Obtain a first reference behavior feature dataset, wherein the first reference behavior feature dataset consists of reference behavior feature data corresponding to multiple first reference users, and the first reference behavior feature dataset has the same feature category as the target behavior feature dataset; Based on the target behavior feature dataset and the first reference behavior feature dataset, anomaly detection is performed using the isolated forest algorithm to obtain the first detection result.

3. The method according to claim 2, characterized in that, The step of performing anomaly detection using the isolated forest algorithm based on the target behavior feature dataset and the first reference behavior feature dataset, and obtaining the first detection result, includes: Based on the target behavior feature dataset and the first reference behavior feature dataset, construct a corresponding binary tree for any of the feature categories; Obtain the segmentation point corresponding to each binary tree, and recursively segment the corresponding binary tree according to the segmentation point to obtain the average segmentation depth of the target user in all the binary trees; The first detection result is obtained based on the average segmentation depth.

4. The method according to claim 3, characterized in that, The step of obtaining the first detection result based on the average segmentation depth includes: Based on the average segmentation depth, the first detection score of the target user is obtained; Obtain a first threshold, and obtain the first detection result based on the first detection score and the first threshold.

5. The method according to claim 1, characterized in that, The step of performing anomaly detection using the local anomaly factor algorithm based on the target behavior feature dataset to obtain a second detection result includes: Obtain a second reference behavior feature dataset, wherein the second reference behavior feature dataset consists of reference behavior feature data corresponding to multiple second reference users, and the second reference behavior feature dataset has the same feature categories as the target behavior feature dataset; Based on the target behavior feature dataset and the second reference behavior feature dataset, anomaly detection is performed using the local anomaly factor algorithm to obtain the second detection result.

6. The method according to claim 5, characterized in that, The step of performing anomaly detection using the isolated forest algorithm based on the target behavior feature dataset and the second reference behavior feature dataset, and obtaining the second detection result, includes: Based on the target behavior feature dataset and the second reference behavior feature dataset, the similarity between the target user and each of the second reference users is obtained; Based on all the similarities, at least one neighbor of the target user is selected from all the second reference users, and the local reachability density between the target user and all the neighbors is obtained; The second detection result is obtained based on the locally reachable density.

7. The method according to claim 6, characterized in that, The step of obtaining the second detection result based on the locally reachable density includes: Based on the local reachability density, obtain the second detection score of the target user; Obtain a second threshold, and based on the second detection score and the second threshold, obtain the second detection result.

8. The method according to claim 1, characterized in that, The step of using graph behavior analysis algorithms for anomaly detection to obtain the group abnormal behavior detection results of the target users includes: Obtain a third reference behavior feature dataset, wherein the third reference behavior feature dataset consists of reference behavior feature data corresponding to multiple third reference users, and the third reference behavior feature dataset has the same feature categories as the target behavior feature dataset; Based on the target behavior feature dataset and the third reference behavior feature dataset, anomaly detection is performed using a graph behavior analysis algorithm to obtain the abnormal behavior detection results of the group.

9. The method according to claim 8, characterized in that, The step of performing anomaly detection using a graph behavior analysis algorithm based on the target behavior feature dataset and the third reference behavior feature dataset to obtain the abnormal behavior detection results of the group includes: Using the target user and all the third reference users as nodes, and the lines connecting any two nodes with related behaviors as edges, construct a target undirected graph; Based on the target undirected graph, the abnormal behavior detection results of the group are obtained.

10. The method according to claim 9, characterized in that, The step of obtaining the abnormal behavior detection result of the group based on the target undirected graph includes: Perform vector transformation processing on each node in the target undirected graph to obtain the embedding vector of each node; Clustering is performed on all the embedding vectors to obtain the clustering results; Based on the clustering results, at least one associated user corresponding to the target user is determined, and based on the information of all the associated users, the abnormal behavior detection results of the group are obtained.

11. The method according to claim 1, characterized in that, The method further includes: In response to the detection result of the individual abnormal behavior being found to be abnormal, and / or in response to the detection result of the group abnormal behavior being found to be abnormal, a prevention and control process is executed, wherein the prevention and control process includes at least one of error reporting, blocking, and requiring identity authentication.

12. The method according to claim 1, characterized in that, The method further includes: Obtain feedback information and perform optimization processing based on the feedback information, wherein the optimization processing includes optimizing at least one of the first reference behavioral feature dataset, the second reference behavioral feature dataset, and the third reference behavioral feature dataset.

13. The method according to claim 1, characterized in that, Before acquiring the target user's behavior data to be detected, the process also includes: Obtain the target user's raw behavioral data and target preprocessing strategy; The original behavioral data is preprocessed according to the target preprocessing strategy, and the processed data is used as the target user's behavioral data to be detected.

14. An abnormal behavior detection device, characterized in that, include: The behavior feature dataset acquisition unit is used to acquire the target user's behavior data to be detected, and to acquire the target behavior feature dataset for the target user based on the target behavior data. The intermediate detection result acquisition unit is used to perform anomaly detection using the isolated forest algorithm based on the target behavior feature dataset to obtain a first detection result, and to perform anomaly detection using the local anomaly factor algorithm based on the target behavior feature dataset to obtain a second detection result. An individual abnormal behavior detection result acquisition unit is used to acquire the individual abnormal behavior detection result of the target user based on the first detection result and the second detection result; The group abnormal behavior detection result acquisition unit is used to detect abnormality by using a graph behavior analysis algorithm in response to the individual abnormal behavior detection result indicating the presence of anomalies, and to acquire the group abnormal behavior detection result of the target user, wherein the group abnormal behavior detection result includes information of at least one associated user corresponding to the target user.

15. The apparatus according to claim 14, characterized in that, The intermediate detection result acquisition unit is further configured to: Obtain a first reference behavior feature dataset, wherein the first reference behavior feature dataset consists of reference behavior feature data corresponding to multiple first reference users, and the first reference behavior feature dataset has the same feature category as the target behavior feature dataset; Based on the target behavior feature dataset and the first reference behavior feature dataset, anomaly detection is performed using the isolated forest algorithm to obtain the first detection result.

16. The apparatus according to claim 15, characterized in that, The intermediate detection result acquisition unit is further configured to: Based on the target behavior feature dataset and the first reference behavior feature dataset, construct a corresponding binary tree for any of the feature categories; Obtain the segmentation point corresponding to each binary tree, and recursively segment the corresponding binary tree according to the segmentation point to obtain the average segmentation depth of the target user in all the binary trees; The first detection result is obtained based on the average segmentation depth.

17. The apparatus according to claim 16, characterized in that, The intermediate detection result acquisition unit is further configured to: Based on the average segmentation depth, the first detection score of the target user is obtained; Obtain a first threshold, and obtain the first detection result based on the first detection score and the first threshold.

18. The apparatus according to claim 14, characterized in that, The intermediate detection result acquisition unit is further configured to: Obtain a second reference behavior feature dataset, wherein the second reference behavior feature dataset consists of reference behavior feature data corresponding to multiple second reference users, and the second reference behavior feature dataset has the same feature categories as the target behavior feature dataset; Based on the target behavior feature dataset and the second reference behavior feature dataset, anomaly detection is performed using the local anomaly factor algorithm to obtain the second detection result.

19. The apparatus according to claim 18, characterized in that, The intermediate detection result acquisition unit is further configured to: Based on the target behavior feature dataset and the second reference behavior feature dataset, the similarity between the target user and each of the second reference users is obtained; Based on all the similarities, at least one neighbor of the target user is selected from all the second reference users, and the local reachability density between the target user and all the neighbors is obtained; The second detection result is obtained based on the locally reachable density.

20. The apparatus according to claim 19, characterized in that, The intermediate detection result acquisition unit is further configured to: Based on the local reachability density, obtain the second detection score of the target user; Obtain a second threshold, and based on the second detection score and the second threshold, obtain the second detection result.

21. The apparatus according to claim 14, characterized in that, The group abnormal behavior detection result acquisition unit is also used for: Obtain a third reference behavior feature dataset, wherein the third reference behavior feature dataset consists of reference behavior feature data corresponding to multiple third reference users, and the third reference behavior feature dataset has the same feature categories as the target behavior feature dataset; Based on the target behavior feature dataset and the third reference behavior feature dataset, anomaly detection is performed using a graph behavior analysis algorithm to obtain the abnormal behavior detection results of the group.

22. The apparatus according to claim 21, characterized in that, The group abnormal behavior detection result acquisition unit is also used for: Using the target user and all the third reference users as nodes, and the lines connecting any two nodes with related behaviors as edges, construct a target undirected graph; Based on the target undirected graph, the abnormal behavior detection results of the group are obtained.

23. The apparatus according to claim 22, characterized in that, The group abnormal behavior detection result acquisition unit is also used for: Perform vector transformation processing on each node in the target undirected graph to obtain the embedding vector of each node; Clustering is performed on all the embedding vectors to obtain the clustering results; Based on the clustering results, at least one associated user corresponding to the target user is determined, and based on the information of all the associated users, the abnormal behavior detection results of the group are obtained.

24. The apparatus according to claim 14, characterized in that, The device further includes a prevention and control processing unit for: In response to the detection result of the individual abnormal behavior being found to be abnormal, and / or in response to the detection result of the group abnormal behavior being found to be abnormal, a prevention and control process is executed, wherein the prevention and control process includes at least one of error reporting, blocking, and requiring identity authentication.

25. The apparatus according to claim 14, characterized in that, The device further includes an optimization processing unit for: Obtain feedback information and perform optimization processing based on the feedback information, wherein the optimization processing includes optimizing at least one of the first reference behavioral feature dataset, the second reference behavioral feature dataset, and the third reference behavioral feature dataset.

26. The apparatus according to claim 14, characterized in that, The behavioral feature dataset acquisition unit is also used for: Obtain the target user's raw behavioral data and target preprocessing strategy; The original behavioral data is preprocessed according to the target preprocessing strategy, and the processed data is used as the target user's behavioral data to be detected.

27. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-13.