Abnormal user detection method and device
By combining the probability and similarity scoring method of the abnormal user detection model, the problem of low detection accuracy of abnormal user detection in the prior art is solved, and a more accurate and intuitive user rating is achieved.
Patent Information
- Application Number
- CN202210568270.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-05-24
AI Technical Summary
The detection results of existing abnormal user detection models are low in accuracy, making it difficult to quickly and accurately adapt to changes in fraudulent behavior.
The preset detection model predicts the probability that the user is an abnormal user, and combines the similarity between the user and the abnormal user group to determine the maximum similarity, obtain user scores, and use multi-dimensional information to score to improve judgment accuracy.
It improves the accuracy and quantitative interpretability of abnormal user judgments, and overcomes the problem of low accuracy of model detection results.
Smart Images

Figure CN114912534B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of communication technology, and specifically relates to a method and device for detecting abnormal users. Background Art
[0002] Abnormal user detection is a feature that identifies users engaging in unusual activities, including transaction fraud, online fraud, phone fraud, card and account theft, and telemarketing. Abnormal user detection is a key focus for carriers, helping users identify unusual activities like fraudulent calls and text messages. This effectively prevents financial loss and improves user security, well-being, and user experience.
[0003] For abnormal user detection models, the selection of abnormal behavior features is very important. However, in reality, abnormal behaviors such as fraud are constantly changing rapidly, while the abnormal behavior features learned by the abnormal user detection model are difficult to change quickly and accurately. There is a lag, resulting in the abnormal user detection model being unable to accurately determine whether the behavior is abnormal, regardless of whether the behavior features are directly obtained from the business or derived from the business. In other words, the detection results of the abnormal user detection model have the problem of low accuracy. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a method and device for detecting abnormal users, so as to solve the problem of low detection accuracy of abnormal user detection models in the prior art.
[0005] In a first aspect, an embodiment of the present application provides a method for detecting abnormal users, the method comprising:
[0006] Predicting the probability that the first target user is an abnormal user through a preset detection model;
[0007] Determine the similarity between the first target user and N abnormal user groups respectively; wherein each abnormal user group corresponds to an abnormal user type, and N is an integer greater than or equal to 1;
[0008] Determining a maximum similarity among similarities between the first target user and the N abnormal user groups;
[0009] Obtaining a user score for the first target user based on the probability that the first target user is an abnormal user and the maximum similarity predicted by the preset detection model;
[0010] Determine whether the first target user is an abnormal user based on the user score.
[0011] In a second aspect, an embodiment of the present application provides an abnormal user detection device, the device comprising:
[0012] A prediction module, configured to predict the probability that the first target user is an abnormal user through a preset detection model;
[0013] A first determination module is configured to determine similarities between the first target user and N abnormal user groups, respectively; wherein each abnormal user group corresponds to an abnormal user type, and N is an integer greater than or equal to 1;
[0014] A second determining module is configured to determine a maximum similarity among similarities between the first target user and the N abnormal user groups;
[0015] A scoring module, configured to obtain a user score of the first target user based on the probability that the first target user is an abnormal user and the maximum similarity predicted by the preset detection model;
[0016] The third determination module is configured to determine whether the first target user is an abnormal user based on the user score.
[0017] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps in the abnormal user detection method described in the first aspect are implemented.
[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps in the abnormal user detection method described in the first aspect are implemented.
[0019] In an embodiment of the present application, for abnormal user detection, the first target user is scored based on the model output result (i.e., the probability that the first target user predicted by the model is an abnormal user) and the similarity between the first target user and the abnormal user group. First, by scoring the first target user through multi-dimensional information (i.e., probability and similarity), the user score can be made more accurate, thereby making the judgment of abnormal users more accurate. Moreover, the abnormal user group contains abnormal users who have been detected. Using the behavioral patterns (i.e., behavioral characteristics) of known abnormal users as prior knowledge for abnormal behavior detection can also improve the accuracy of user scores, which is conducive to overcoming the problem of low accuracy of model detection results. Secondly, compared with probability, user scores are more intuitive and have higher quantitative interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A flowchart of the abnormal user detection method provided in an embodiment of the present application;
[0021] Figure 2A schematic diagram of an example process provided in an embodiment of the present application;
[0022] Figure 3 A schematic block diagram of an abnormal user detection device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described below are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0024] It should be understood that references to "one embodiment" or "an embodiment" in this specification mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, the appearance of "in one embodiment" or "in an embodiment" throughout this specification does not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0025] In the various embodiments of the present application, it should be understood that the serial numbers of the steps do not mean an absolute order of execution. The execution order of each step should be determined by its function and internal logic. Therefore, the serial numbers of each step should not constitute an absolute limitation on the implementation process of the embodiments of the present application.
[0026] The abnormal user detection method provided in the embodiment of the present application is described in detail below with reference to specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0027] An embodiment of the present application provides an abnormal user detection method, which is applied to an electronic device, which may be a server or a terminal device.
[0028] like Figure 1 As shown, the abnormal user detection method may include:
[0029] Step 101: Predict the probability that the first target user is an abnormal user through a preset detection model.
[0030] The preset detection model described here is used to predict the probability that a user is an abnormal user, and an abnormal user generally refers to a user with abnormal behavior. Therefore, predictions can be made based on user behavior, that is, the input data of the model can be the user's behavioral feature data. Therefore, in an embodiment of the present application, the behavioral feature data of the first target user can be obtained first, and then the behavioral feature data of the first target user can be input into the preset detection model to obtain the output result of the preset detection model, that is, the probability that the first target user is an abnormal user.
[0031] The behavioral characteristic data described herein may include but is not limited to at least one of: the first target user's call frequency, call time period, SMS sending frequency, SMS sending time period, and the number of suspected fraudsters among the first target user's contacts.
[0032] The model is pre-trained, and its training process can be described as follows:
[0033] Step 1: Use user behavior data as features.
[0034] Step 2: Define abnormal user behavior (or abnormal users) and abnormal user behavior (or normal users) as target variables.
[0035] Step 3: Combine features and target variables into samples and use classification algorithms to train the model.
[0036] Step 102: Determine the similarity between the first target user and N abnormal user groups respectively.
[0037] For example, the number of abnormal user groups is 3, namely group A, group B and group C. Determining the similarity between the first target user and the N abnormal user groups respectively means: determining the similarity between the first target user and group A, the similarity between the first target user and group B, and the similarity between the first target user and group C respectively.
[0038] The N abnormal user groups described here are obtained by pre-dividing known abnormal users according to their types. Each abnormal user group corresponds to a specific abnormal user type, i.e., the abnormal users in each abnormal user group belong to the abnormal user type corresponding to that abnormal user group. An abnormal user group includes at least one abnormal user. N is an integer greater than or equal to 1.
[0039] The higher the similarity between a user's behavioral characteristics and those of abnormal users, the greater the probability that the user is an abnormal user. However, different types of abnormal users will have different behavioral levels, that is, different behavioral characteristics. Therefore, the embodiment of the present application can group abnormal users with different behavioral characteristics, and each group after grouping corresponds to an abnormal user type. Then, the similarity between the first target user and each abnormal user group in the N abnormal user groups is determined separately, that is, the similarity between the first target user and different types of abnormal users is determined separately, so that abnormal users can be judged more comprehensively and accurately.
[0040] Among them, the abnormal user types may include but are not limited to: at least one of telephone fraud, telephone sales, etc.
[0041] Step 103: Determine the maximum similarity among the similarities between the first target user and the N abnormal user groups.
[0042] In this embodiment of the present application, after obtaining the similarities between the first target user and each abnormal user group, the maximum value, i.e., the maximum similarity, is determined. The higher the similarity between the first target user and the abnormal user group, the greater the probability that the first target user belongs to the abnormal user type corresponding to the abnormal user group. Determining abnormal users based on the maximum similarity can improve the accuracy of abnormal user determination.
[0043] Step 104: Obtain a user rating of the first target user based on the probability and maximum similarity of the first target user being an abnormal user predicted by the preset detection model.
[0044] In an embodiment of the present application, based on the output result of the preset detection model (i.e., the probability that the first target user predicted by the model is an abnormal user), the first target user is scored based on the maximum similarity between the first target user and the abnormal user group to obtain a user score for the first target user.
[0045] Since the preset detection model itself has a certain error, and the behavioral characteristics that the model has learned have a lag, the accuracy of the model output result is reduced, and the probability is poor in terms of quantitative interpretability. In the embodiment of the present application, the first target user is scored by multi-dimensional information (i.e., probability and similarity), which can make the user score more accurate, thereby making the judgment of abnormal users more accurate. And the abnormal user group is composed of abnormal users who have been detected. Using the behavioral patterns (i.e., behavioral characteristics) of known abnormal users as prior knowledge for abnormal behavior detection can also improve the accuracy of user scores, which is conducive to overcoming the problem of low accuracy of model detection results. Secondly, compared with probability, user scores are more intuitive and have higher quantitative interpretability.
[0046] Step 105: Determine whether the first target user is an abnormal user based on the user score of the first target user.
[0047] In the embodiment of the present application, the user rating can be a positive rating of the user, that is, the higher the rating, the smaller the probability that the user is an abnormal user, and conversely, the lower the rating, the greater the probability that the user is an abnormal user.
[0048] In the embodiment of the present application, a preset score threshold can be used to determine whether a user is an abnormal user. For example, if the user score of the first target user is greater than or equal to the score threshold, the first target user is determined to be a normal user; if the user score of the first target user is less than the score threshold, the first target user is determined to be an abnormal user.
[0049] Among them, different scoring criteria (i.e., scoring thresholds) can be set according to different business scenarios to adapt to business needs.
[0050] As an optional embodiment, step 102: determining the similarity between the first target user and the N abnormal user groups respectively may include:
[0051] Step A1: Determine the similarity between the first target user and each abnormal user in each abnormal user group.
[0052] For example, the number of abnormal user groups is 3, namely group A, group B and group C. Then, respectively determining the similarity between the first target user and each abnormal user in each abnormal user group means: respectively determining the similarity between the first target user and each abnormal user in group A, the similarity between the first target user and each abnormal user in group B, and the similarity between the first target user and each abnormal user in group C.
[0053] Step A2: determining an average similarity between the first target user and the abnormal users in each abnormal user group according to the similarity between the first target user and each abnormal user in each abnormal user group.
[0054] After obtaining the similarity between the first target user and each abnormal user in each abnormal user group in step A1, the average similarity between the first target user and the abnormal users in each abnormal user group is calculated.
[0055] Continuing with the example in step A1 above, the example of determining the average similarity between the first target user and the abnormal users in each abnormal user group is as follows: averaging the similarities between the first target user and each abnormal user in group A to obtain the average similarity between the first target user and the abnormal users in group A; averaging the similarities between the first target user and each abnormal user in group B to obtain the average similarity between the first target user and the abnormal users in group B; and averaging the similarities between the first target user and each abnormal user in group C to obtain the average similarity between the first target user and the abnormal users in group C.
[0056] Step A3: Determine the average similarity between the first target user and the abnormal users in each abnormal user group as the similarity between the first target user and the corresponding abnormal user group.
[0057] The mean represents the central tendency of a group of data and is representative to a certain extent. Therefore, in the embodiment of the present application, the average similarity between the first target user and the abnormal users in each abnormal user group is determined as the similarity between the first target user and the corresponding abnormal user group.
[0058] As an optional embodiment, step 102: determining the similarity between the first target user and the N abnormal user groups respectively may include:
[0059] Step B1: Determine the similarity between the first target user and the preset abnormal users in each abnormal user group respectively.
[0060] Step B2: Determine the similarity between the first target user and the preset abnormal users in each abnormal user group as the similarity between the first target user and the corresponding abnormal user group.
[0061] In an embodiment of the present application, a representative abnormal user (i.e., a preset abnormal user) can be specified in each abnormal user group. The preset abnormal user has a high or highest matching degree with the abnormal user type corresponding to the group, and its behavioral characteristics are relatively consistent with or most consistent with the abnormal user type corresponding to the group.
[0062] Since the preset abnormal users are representative, in the embodiment of the present application, the similarity between the first target user and the preset abnormal users in each abnormal user group can be determined as the similarity between the first target user and the corresponding abnormal user group.
[0063] The preset abnormal user may be manually specified based on experience, or a preset detection model may be used to detect abnormal users in each group, and the abnormal user with the largest output result may be determined as the preset abnormal user in the group.
[0064] As an optional embodiment, step 104: obtaining a user score of the first target user based on the probability and maximum similarity of the first target user being an abnormal user predicted by the preset detection model may include:
[0065] The probability that the first target user is an abnormal user predicted by the preset detection model and the maximum similarity among the similarities between the first target user and N abnormal user groups are substituted into the first preset formula to obtain a user score of the first target user.
[0066] Among them, the first preset formula is a functional relationship between the scoring weight, the probability that the user is an abnormal user and the user score.
[0067] In the embodiment of the present application, the probability that the first target user is an abnormal user and the maximum similarity among the similarities between the first target user and N abnormal user groups can be converted into a user score of the first target user through a first preset formula.
[0068] This embodiment of the present application provides a specific embodiment of a first preset formula, as described below:
[0069] The first preset formula is:
[0070]
[0071] Among them, score user represents the user score of the first target user; m represents the maximum value of the user score value range. For example, if the score value range is [0,100], then m is 100; odds represents the boundary probability that the user is an abnormal user, which can be 0.5 or other values according to actual needs; score odds Indicates the user's basic score when the boundary probability of the user being an abnormal user is odds. For example, when odds is 0.5 and the score range is [0,100], score odds The value can be 50; P 异常 represents the probability that the first target user is an abnormal user, P 正常 represents the probability that the first target user is a normal user, P 正常 =1-P 异常 ;simility max Represents the maximum similarity among the similarities between the first target user and N abnormal user groups.
[0072] In the embodiment of the present application, the abnormal probability predicted by the model can be corrected and mapped into a user score through a first preset formula, and the corrected score is used instead of the abnormal probability as the criterion for judging abnormal users, thereby reducing the error in judging abnormal users by the abnormal probability.
[0073] It is understandable that the first preset formula can be designed according to actual needs. For example, the basic score in the first preset formula can be removed. odds , only the part of the formula for converting probability and maximum similarity into user ratings is retained.
[0074] As an optional embodiment, in the embodiment of the present application, abnormal users in the abnormal user set can be divided into N abnormal user groups by using a preset clustering algorithm (such as the Kmeans algorithm). Therefore, before step 102: determining the similarity between the first target user and the N abnormal user groups, the method may further include:
[0075] Step C1: N preset abnormal users in the abnormal user set are set as N initial cluster centers respectively.
[0076] Each pre-set abnormal user corresponds to an abnormal user type, and within the set of abnormal users, the pre-set abnormal user has the highest degree of match with its corresponding abnormal user type compared to other abnormal users. In other words, each pre-set abnormal user is a representative user of its corresponding abnormal user type. The pre-set abnormal user can be an abnormal user manually determined based on experience to have the highest degree of match with its corresponding abnormal user type, or a pre-set detection model can detect abnormal users in each group and determine the abnormal user with the highest output result as having the highest degree of match with its corresponding abnormal user type.
[0077] Wherein, N is determined according to the number of abnormal user types. For example, if the number of abnormal user types is 3, then N=3.
[0078] Each initial cluster center represents a cluster, so the initial cluster center can also be called the cluster initial center. A cluster corresponds to a group.
[0079] Step C2: According to the initial cluster centers and the preset clustering algorithm, the abnormal users in the abnormal user set are divided into N abnormal user groups.
[0080] The selection of initial cluster centers has a significant impact on the clustering results. Conventional clustering algorithms (such as the Kmeans algorithm) randomly select k objects from the objects to be clustered as initial cluster centers, resulting in an unclear clustering target. However, in the embodiments of the present application, initial cluster centers are specified, and the specified initial cluster centers are representative, with a clear clustering target. This facilitates accurate clustering, reduces the number of clustering attempts, improves clustering efficiency, and enhances the accuracy of clustering results.
[0081] As an optional embodiment, step C2: dividing abnormal users in the abnormal user set into N abnormal user groups according to the initial cluster center and the preset clustering algorithm may include:
[0082] Step C21: grouping other abnormal users according to the first similarity between other abnormal users and each initial cluster center to obtain N first abnormal user groups.
[0083] Among them, other abnormal users are abnormal users in the abnormal user set except the abnormal user serving as the initial cluster center.
[0084] In the embodiment of the present application, clustering and grouping can be performed based on the similarity between users.
[0085] For the first clustering grouping, the specific implementation method can be described as follows: calculate the similarity between each abnormal user (specifically, the abnormal users in the abnormal user set except the abnormal users serving as the initial clustering center) and the N initial clustering centers respectively, and divide the abnormal user into a group with the initial clustering center corresponding to the maximum similarity, that is, the abnormal user is divided into a group with the initial clustering center with which the abnormal user has the greatest similarity, thus completing the first clustering grouping and obtaining N first abnormal user groups.
[0086] Step C22: Determine a new cluster center for each first abnormal user group.
[0087] After completing the first clustering grouping, it is necessary to determine the new cluster center of each first abnormal user group respectively. The specific implementation method can be: calculate the behavioral feature mean of all abnormal users in each first abnormal user group respectively, and determine the abnormal user represented by the behavioral feature mean as the new cluster center. For example, the behavioral features include: call frequency and SMS sending frequency, then calculate the call frequency mean and SMS sending frequency mean of all abnormal users in each first abnormal user group, and then use the call frequency mean and SMS sending frequency mean as the new cluster center of the first abnormal user group. Among them, the new cluster center is a virtual cluster center (also called a cluster virtual center), that is, the abnormal user corresponding to the cluster center is not a real user.
[0088] Step C23: Group the other abnormal users according to the weighted sum of their first similarity to each initial cluster center and their second similarity to each new cluster center to obtain N second abnormal user groups, and determine a new cluster center for each of the second abnormal user groups.
[0089] Step C24: Repeat the steps of "determining a new cluster center for each second abnormal user group" and step C23 in sequence until the number of groupings reaches a preset number or the difference between the cluster center determined after each grouping and the cluster center determined after the previous grouping is less than or equal to a preset value.
[0090] For the cluster grouping after the first cluster grouping, the embodiment of the present application performs cluster grouping based on the similarity between the abnormal user and each initial cluster center (i.e., the first similarity) and the similarity with each new cluster center (i.e., the second similarity). The specific implementation method can be described as follows: respectively calculate the second similarity between each abnormal user (specifically, the abnormal users in the abnormal user set except the abnormal users as the initial cluster center) and each new cluster center, and then perform weighted summation of the first similarity and the second similarity based on the first preset weight of the initial cluster center and the second preset weight of the new cluster center. Finally, the abnormal user is divided into a group with the initial cluster center and the new cluster center corresponding to the maximum weighted sum result, that is, the abnormal user is divided into a group with the initial cluster center and the new cluster center in which the weighted sum of the initial cluster center and the new cluster center in the first abnormal user group has the largest similarity. In this way, another cluster grouping is completed to obtain N second abnormal user groups.
[0091] For cluster grouping after the second cluster grouping, refer to the second cluster grouping method and repeat it until the number of groupings reaches the preset number or the difference between the cluster center determined after each grouping and the cluster center determined after the previous grouping is less than or equal to the preset value.
[0092] It should be noted that the group obtained after the second clustering can also be called the second abnormal user group. The first abnormal user group and the second abnormal user group mentioned here are only used to distinguish the group obtained after the first clustering and the group obtained after the first clustering.
[0093] The first similarity and the second similarity can be weighted and summed using a second preset formula. The second preset formula is:
[0094] Among them, d 总 Represents the weighted sum result, which can also be understood as the total similarity; d(x,centre initial ) represents other abnormal users and the initial cluster center centre initial The first similarity of d(x,centre invented ) represents other abnormal users and the new cluster center centre invented The second similarity; Indicates the weight of the initial cluster center (i.e., the first preset weight), which is generally set to 0.3; Represents the weight of the new cluster center (i.e., the second preset weight), When the value is 0.3, The value is 0.7.
[0095] In the embodiment of the present application, when clustering, a weighted calculation of the similarity with the initial cluster center and the new cluster center is used to comprehensively measure the distance between the abnormal user and the cluster. This can prevent the abnormal users in a single cluster from being too discrete from the initial cluster center after multiple iterations of clustering, reduce the degree of discreteness between the abnormal users in a single cluster and the initial cluster center, and thus improve the accuracy of clustering. Furthermore, because the abnormal user grouping is more accurate, when calculating the abnormal similarity between the first target user and the abnormal user group (i.e., the abnormal user group), the similarity result obtained can also be more accurate.
[0096] As an optional embodiment, in step 105: after determining whether the first target user is an abnormal user based on the user score of the first target user, the method may further include: sending a reminder message to the second target user if it is determined that the first target user is an abnormal user.
[0097] The second target user is the communication object when the first target user is the communication initiator, that is, the contact of the first target user (such as a telephone contact, a text message contact, etc.).
[0098] In an embodiment of the present application, when the first target user is determined to be an abnormal user based on the user score, a reminder message can be generated and sent to the second target user to inform the second target user that the first target user may be an abnormal user, so that the second target user will pay attention to the first target user and reduce the occurrence of being deceived.
[0099] The reminder message may also include information about the abnormal user type of the first target user. For example, the reminder message may be: "[Abnormal Number Reminder] Dear customer, hello! The previous contact number: 1873160xxxx may be a sales call or a scam call. Please be careful."
[0100] As an optional embodiment, after the aforementioned step of sending the reminder information to the second target user, the method may further include:
[0101] Step D1: receiving feedback information from the second target user marking the first target user as abnormal.
[0102] Step D2: According to the feedback information, the number of users who mark the first target user as abnormal is counted.
[0103] Step D3: when the number of users who mark the first target user as abnormal is greater than a first threshold, the first target user is added to the abnormal user set.
[0104] In this embodiment of the present application, a user-friendly "abnormal user marking" function can be provided. When a user receives a call or text message from an abnormal user, the user can proactively choose to mark the abnormal call or text message. When a phone number is marked a certain number of times, it can be added to the abnormal user collection to update the abnormal user data.
[0105] In this embodiment of the present application, user feedback on abnormal user behavior is collected through interaction with users, thereby enhancing the user's real experience and perception during abnormal user detection. When the number of tags for the same abnormal user exceeds a first threshold, the abnormal user is added to the abnormal user set. This ensures that the abnormal user set can promptly include the latest abnormal behavior patterns, reducing the lag in the acquisition of abnormal behavior feature data. In this way, when abnormal user data in the abnormal user set is used for abnormal user detection, the accuracy of detection can be improved.
[0106] Optionally, the aforementioned step of adding the first target user to the abnormal user set when the number of users who mark the first target user as abnormal is greater than a first threshold may include:
[0107] When the number of users who mark the first target user as abnormal is greater than a first threshold, the first target user is added to the abnormal mark set; when the number of users in the abnormal mark set is greater than a second threshold, the abnormal users including the first target user in the abnormal mark set are added to the abnormal user set, and the abnormal mark set is cleared.
[0108] In an embodiment of the present application, if the number of users who have marked a first target user as abnormal exceeds a first threshold, the data of the first target user can be temporarily stored in an abnormal marking set. If the number of users stored in the abnormal marking set exceeds a second threshold, abnormal users in the abnormal marking set, including the first target user, are added to the abnormal user set, thereby updating the abnormal user data in the abnormal user set. This can reduce the number of abnormal user set data updates and the amount of data processing.
[0109] The first threshold and the second threshold are integers greater than or equal to 1, and their specific values can be determined according to actual needs.
[0110] Finally, in order to better understand the technical solutions provided by the embodiments of the present application, the technical solutions provided by the embodiments of the present application are comprehensively described below using an exemplary embodiment.
[0111] like Figure 2 As shown, this example includes the following process:
[0112] Step 201: Construct an abnormal user set T using the abnormal users that have been detected.
[0113] Step 202: Perform cluster analysis on the set T using a clustering algorithm, and group abnormal users in the set T.
[0114] Step 203: Calculate the similarity between the user to be detected x and each user in a single cluster (i.e., group), and then calculate the average similarity between user x and all abnormal users in a single cluster to form an average similarity group {Similarity1, Similarity2,…, Similarity N}.
[0115] Step 204: Find the maximum abnormal similarity in the average similarity group max .
[0116] Step 205: Use the preset detection model to predict the probability P that user x is an abnormal user 异常 .
[0117] Step 206: Use the score mapping formula (i.e., the aforementioned formula 1) to convert the probability P 异常 Map it to a score in the interval [0,100] and use the score as the user rating.
[0118] Step 207 : Determine whether the user score of user x is less than the abnormal score threshold threshoid. If yes, proceed to step 208 .
[0119] Step 208: Determine user x as an abnormal user.
[0120] If the user score of user x is greater than or equal to the abnormal score threshold threshoid, user x is determined to be a normal user.
[0121] Step 209: When it is determined that user x is an abnormal user, an early warning reminder message is sent to the contacts of user x (by phone, text message, etc.).
[0122] The contactor can choose whether to mark user x as abnormal.
[0123] Step 210: Determine whether the contact has marked user x. If yes, proceed to step 211.
[0124] If the contact confirms that user x is an abnormal user, the contact may choose to mark user x as abnormal.
[0125] Step 211 : Determine whether the number of abnormal marks of user x is greater than a predetermined threshold. If yes, proceed to step 212 .
[0126] Step 212: Add user x to the abnormal tag set U.
[0127] Step 214: Determine whether the number of users in the abnormal mark set U is greater than a predetermined threshold. If yes, proceed to step 214.
[0128] Step 214: Add the users in set U to the abnormal user set T, update the abnormal user set T, and clear the set U. Then, repeat steps 202 to 214.
[0129] Finally, it should be noted that in the embodiments of this application, the Java language and the Spark distributed computing framework can be used to distribute and encapsulate the relevant technical details. Using a distributed database for data storage can improve the real-time performance of the overall system calculation. In addition, the similarity between users described in the embodiments of this application refers to the similarity between user behavior characteristics, and the similarity calculation method can specifically adopt the cosine similarity calculation method.
[0130] The above is a description of the abnormal user detection method provided in the embodiment of the present application.
[0131] In summary, the technical solution provided by the embodiment of the present application, for abnormal user detection, scores the first target user based on the model output result (i.e., the probability that the first target user predicted by the model is an abnormal user) and the similarity between the first target user and the abnormal user group. First, by scoring the first target user through multi-dimensional information (i.e., probability and similarity), the user score can be made more accurate, thereby making the judgment of abnormal users more accurate. Moreover, the abnormal user group contains abnormal users who have been detected. Using the behavioral patterns (i.e., behavioral characteristics) of known abnormal users as prior knowledge for abnormal behavior detection can also improve the accuracy of user scores, which is conducive to overcoming the problem of low accuracy of model detection results. Secondly, compared with probability, user scores are more intuitive and have higher quantitative interpretability.
[0132] The above describes the abnormal user detection method provided by the embodiment of the present application. The following will describe the abnormal user detection device provided by the embodiment of the present application with reference to the accompanying drawings.
[0133] like Figure 3 As shown, an embodiment of the present application also provides an abnormal user detection device, which is applied to electronic equipment.
[0134] The abnormal user detection device may include:
[0135] The prediction module 301 is configured to predict the probability that the first target user is an abnormal user by using a preset detection model.
[0136] The first determining module 302 is configured to respectively determine similarities between the first target user and N abnormal user groups.
[0137] The N abnormal user groups are pre-divided according to the types of abnormal users, each abnormal user group corresponds to one abnormal user type, each abnormal user group includes at least one abnormal user, and N is an integer greater than or equal to 1.
[0138] The second determining module 303 is configured to determine the maximum similarity among the similarities between the first target user and the N abnormal user groups.
[0139] The scoring module 304 is configured to obtain a user score of the first target user based on the probability of the first target user being an abnormal user predicted by the preset detection model and the maximum similarity.
[0140] The third determination module 305 is configured to determine whether the first target user is an abnormal user based on the user score.
[0141] Optionally, the first determining module 302 may include:
[0142] The first determining unit is configured to respectively determine a similarity between the first target user and each abnormal user in each abnormal user group.
[0143] The second determining unit is configured to determine an average similarity between the first target user and the abnormal users in each abnormal user group according to the similarity between the first target user and each abnormal user in each abnormal user group.
[0144] The third determining unit is configured to determine the average similarity between the first target user and the abnormal users in each abnormal user group as the similarity between the first target user and the corresponding abnormal user group.
[0145] Optionally, the scoring module 304 may include:
[0146] A scoring unit is configured to substitute the probability and the maximum similarity into a first preset formula to obtain a user score of the first target user.
[0147] The first preset formula is a functional relationship between the maximum similarity, the probability that the user is an abnormal user, and the user score.
[0148] Optionally, the first preset formula is:
[0149]
[0150] Among them, score userrepresents the user score of the first target user; odds represents the boundary probability that the user is an abnormal user; score odds Indicates the user's basic score when the boundary probability of the user being an abnormal user is odds; P 异常 represents the probability that the first target user is an abnormal user, P 正常 represents the probability that the first target user is a normal user, P 正常 =1-P 异常 ;simility max represents the maximum similarity; m represents the maximum value of the user rating value range.
[0151] Optionally, the device may further include:
[0152] The setting module is used to set N preset abnormal users in the abnormal user set as N initial cluster centers respectively.
[0153] Each of the preset abnormal users corresponds to an abnormal user type, and in the abnormal user set, the preset abnormal user has the highest matching degree with the abnormal user type corresponding to it compared with other abnormal users; N is determined according to the number of abnormal user types.
[0154] The group division module is used to divide the abnormal users in the abnormal user set into N abnormal user groups according to the initial cluster center and a preset clustering algorithm.
[0155] Optionally, the group division module includes:
[0156] The grouping unit is configured to group the other abnormal users according to the first similarity between the other abnormal users and each initial cluster center to obtain N first abnormal user groups.
[0157] The other abnormal users are abnormal users in the abnormal user set except the abnormal user serving as the initial cluster center.
[0158] The fourth determining unit is configured to respectively determine a new cluster center for each of the first abnormal user groups.
[0159] The first processing unit is configured to group the other abnormal users according to a weighted sum of a first similarity between the other abnormal users and each of the initial cluster centers and a second similarity between the other abnormal users and each of the new cluster centers, to obtain N second abnormal user groups, and to determine a new cluster center for each of the second abnormal user groups.
[0160] A control unit is used to control the first processing unit to repeatedly execute the steps of "determining a new cluster center for each second abnormal user group separately" and "grouping the other abnormal users according to the weighted sum of the first similarity of the other abnormal users with each of the initial cluster centers and the second similarity with each new cluster center to obtain N second abnormal user groups, and determining a new cluster center for each of the second abnormal user groups separately" until the number of groupings reaches a preset number or until the difference between the cluster center determined after each grouping and the cluster center determined after the previous grouping is less than or equal to a preset value.
[0161] Optionally, the third determining module 305 may include:
[0162] The fifth determining unit is configured to determine that the first target user is a normal user if the user score is greater than or equal to a score threshold.
[0163] A sixth determining unit is configured to determine that the first target user is an abnormal user if the user score is less than a score threshold.
[0164] Optionally, the device may further include:
[0165] The sending module is used to send a reminder message to the second target user when it is determined that the first target user is an abnormal user.
[0166] The second target user is the communication partner when the first target user is the communication initiator, and the reminder information is used to inform the second target user that the first target user is an abnormal user.
[0167] Optionally, the device may further include:
[0168] The receiving module is configured to receive feedback information from the second target user indicating that the first target user has been marked abnormal.
[0169] A statistics module is used to count the number of users who mark the first target user as abnormal based on the feedback information.
[0170] The data processing module is configured to add the first target user to an abnormal user set when the number of users who mark the first target user as abnormal is greater than a first threshold.
[0171] Optionally, the data processing module may include:
[0172] a second processing unit, configured to add the first target user to an abnormal marking set if the number of users who have marked the first target user as abnormal is greater than a first threshold;
[0173] The third processing unit is configured to add abnormal users including the first target user in the abnormal mark set to an abnormal user set and clear the abnormal mark set when the number of users in the abnormal mark set is greater than a second threshold.
[0174] The abnormal user detection device provided in the embodiment of the present application can achieve Figure 1 To avoid repetition, the various processes implemented by the abnormal user detection device in the illustrated method embodiment will not be described again here.
[0175] In an embodiment of the present application, for abnormal user detection, the first target user is scored based on the model output result (i.e., the probability that the first target user predicted by the model is an abnormal user) and the similarity between the first target user and the abnormal user group. First, by scoring the first target user through multi-dimensional information (i.e., probability and similarity), the user score can be made more accurate, thereby making the judgment of abnormal users more accurate. Moreover, the abnormal user group contains abnormal users who have been detected. Using the behavioral patterns (i.e., behavioral characteristics) of known abnormal users as prior knowledge for abnormal behavior detection can also improve the accuracy of user scores, which is conducive to overcoming the problem of low accuracy of model detection results. Secondly, compared with probability, user scores are more intuitive and have higher quantitative interpretability.
[0176] An embodiment of the present application also provides an electronic device, including a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor. When the program or instruction is executed by the processor, the various steps of the above-mentioned abnormal user detection method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, they are not repeated here.
[0177] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned abnormal user detection method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0178] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0179] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM, RAM, magnetic disk, optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting abnormal users, characterized in that: The method comprises: Predicting the probability that the first target user is an abnormal user through a preset detection model; N preset abnormal users in the abnormal user set are set as N initial cluster centers; wherein each of the preset abnormal users corresponds to an abnormal user type, and in the abnormal user set, the preset abnormal user has the highest matching degree with the abnormal user type corresponding to it compared with other abnormal users; According to the initial cluster center and the preset clustering algorithm, the abnormal users in the abnormal user set are divided into N abnormal user groups; Determine the similarity between the first target user and N abnormal user groups respectively; wherein each abnormal user group corresponds to an abnormal user type, and N is an integer greater than or equal to 1; Determining a maximum similarity among similarities between the first target user and the N abnormal user groups; Obtaining a user score for the first target user based on the probability that the first target user is an abnormal user and the maximum similarity predicted by the preset detection model; Determining whether the first target user is an abnormal user based on the user score; The method of dividing the abnormal users in the abnormal user set into N abnormal user groups according to the initial cluster centers and the preset clustering algorithm includes: Grouping the other abnormal users according to the first similarity between the other abnormal users and each initial cluster center to obtain N first abnormal user groups; wherein the other abnormal users are abnormal users in the abnormal user set other than the preset abnormal user serving as the initial cluster center; Determining a new cluster center for each of the first abnormal user groups respectively; Grouping the other abnormal users according to a weighted sum of a first similarity between the other abnormal users and each of the initial cluster centers and a second similarity between the other abnormal users and each of the new cluster centers to obtain N second abnormal user groups, and determining a new cluster center for each of the second abnormal user groups; Repeat the steps of "determining a new cluster center for each of the second abnormal user groups separately" and "grouping the other abnormal users according to the weighted sum of the first similarity of the other abnormal users with each of the initial cluster centers and the second similarity with each new cluster center to obtain N second abnormal user groups, and determining a new cluster center for each of the second abnormal user groups separately" in sequence until the number of groupings reaches a preset number or the difference between the cluster center determined after each grouping and the cluster center determined after the previous grouping is less than or equal to a preset value.
2. The abnormal user detection method according to claim 1, characterized in that: The respectively determining similarities between the first target user and N abnormal user groups includes: respectively determining a similarity between the first target user and each abnormal user in each abnormal user group; determining an average similarity between the first target user and the abnormal users in each abnormal user group according to the similarity between the first target user and each abnormal user in each abnormal user group; The average similarity between the first target user and the abnormal users in each abnormal user group is determined as the similarity between the first target user and the corresponding abnormal user group.
3. The abnormal user detection method according to claim 1 or 2, characterized in that: Obtaining a user score of the first target user based on the probability of the first target user being an abnormal user and the maximum similarity predicted by the preset detection model includes: Substituting the probability and the maximum similarity into a first preset formula to obtain a user score of the first target user; wherein the first preset formula is a functional relationship between the maximum similarity, the probability that the user is an abnormal user, and the user score.
4. The abnormal user detection method according to claim 3, characterized in that: The first preset formula is: ; in, represents the user rating of the first target user; Indicates the boundary probability that the user is an abnormal user; The boundary probability that the user is an abnormal user is User base rating at the time; represents the probability that the first target user is an abnormal user, represents the probability that the first target user is a normal user, ;simility max represents the maximum similarity; m represents the maximum value of the user rating value range.
5. The abnormal user detection method according to claim 1, characterized in that: The determining whether the first target user is an abnormal user according to the user score includes: If the user score is greater than or equal to a score threshold, determining that the first target user is a normal user; When the user score is less than a score threshold, the first target user is determined to be an abnormal user.
6. The abnormal user detection method according to claim 1 or 5, characterized in that: After determining whether the first target user is an abnormal user based on the user score, the method further includes: When it is determined that the first target user is an abnormal user, a reminder message is sent to the second target user; wherein, the second target user is the communication object when the first target user is the communication initiator, and the reminder message is used to inform the second target user that the first target user is an abnormal user.
7. The abnormal user detection method according to claim 6, characterized in that: After sending the reminder information to the second target user, the method further includes: receiving feedback information from the second target user that the first target user is marked abnormal; Counting the number of users who have marked the first target user as abnormal based on the feedback information; When the number of users who mark the first target user as abnormal is greater than a first threshold, the first target user is added to an abnormal user set.
8. The abnormal user detection method according to claim 7, characterized in that: When the number of users who mark the first target user as abnormal is greater than a first threshold, adding the first target user to the abnormal user set includes: When the number of users who mark the first target user as abnormal is greater than a first threshold, adding the first target user to an abnormal mark set; When the number of users in the abnormal mark set is greater than a second threshold, abnormal users in the abnormal mark set including the first target user are added to the abnormal user set, and the abnormal mark set is cleared.
9. An abnormal user detection device, characterized in that: The device comprises: A prediction module, configured to predict the probability that the first target user is an abnormal user through a preset detection model; A setting module, configured to set N preset abnormal users in the abnormal user set as N initial cluster centers; wherein each of the preset abnormal users corresponds to an abnormal user type, and in the abnormal user set, the preset abnormal user has the highest matching degree with the abnormal user type corresponding to it compared with other abnormal users; A group division module, configured to divide abnormal users in the abnormal user set into N abnormal user groups according to the initial cluster centers and a preset clustering algorithm; A first determination module is configured to determine similarities between the first target user and N abnormal user groups, respectively; wherein each abnormal user group corresponds to an abnormal user type, and N is an integer greater than or equal to 1; A second determining module is configured to determine a maximum similarity among similarities between the first target user and the N abnormal user groups; A scoring module, configured to obtain a user score of the first target user based on the probability that the first target user is an abnormal user and the maximum similarity predicted by the preset detection model; a third determining module, configured to determine whether the first target user is an abnormal user based on the user score; The group division module includes: a grouping unit, configured to group the other abnormal users according to a first similarity between the other abnormal users and each initial cluster center, to obtain N first abnormal user groups; wherein the other abnormal users are abnormal users in the abnormal user set other than the preset abnormal user serving as the initial cluster center; a fourth determining unit, configured to respectively determine a new cluster center of each of the first abnormal user groups; a first processing unit, configured to group the other abnormal users according to a weighted sum of a first similarity between the other abnormal users and each of the initial cluster centers and a second similarity between the other abnormal users and each of the new cluster centers, to obtain N second abnormal user groups, and to determine a new cluster center for each of the second abnormal user groups; A control unit is used to control the first processing unit to repeatedly execute the steps of "respectively determining a new cluster center for each second abnormal user group" and "grouping the other abnormal users according to the weighted sum of the first similarity of the other abnormal users with each of the initial cluster centers and the second similarity with each new cluster center to obtain N second abnormal user groups, and respectively determining a new cluster center for each of the second abnormal user groups" until the number of groupings reaches a preset number or until the difference between the cluster center determined after each grouping and the cluster center determined after the previous grouping is less than or equal to a preset value.
10. The abnormal user detection device according to claim 9, characterized in that: The first determining module includes: a first determining unit, configured to respectively determine a similarity between the first target user and each abnormal user in each abnormal user group; a second determining unit, configured to determine an average similarity between the first target user and the abnormal users in each abnormal user group according to the similarity between the first target user and each abnormal user in each abnormal user group; The third determining unit is configured to determine the average similarity between the first target user and the abnormal users in each abnormal user group as the similarity between the first target user and the corresponding abnormal user group.
11. The abnormal user detection device according to claim 9 or 10, characterized in that: The scoring module includes: A scoring unit is configured to substitute the probability and the maximum similarity into a first preset formula to obtain a user score of the first target user; wherein the first preset formula is a functional relationship between the maximum similarity, the probability that the user is an abnormal user, and the user score.
12. The abnormal user detection device according to claim 11, characterized in that: The first preset formula is: ; in, represents the user rating of the first target user; Indicates the boundary probability that the user is an abnormal user; The boundary probability that the user is an abnormal user is User base rating at the time; represents the probability that the first target user is an abnormal user, represents the probability that the first target user is a normal user, ;simility max represents the maximum similarity; m represents the maximum value of the user rating value range.
13. The abnormal user detection device according to claim 9, characterized in that: The third determining module includes: a fifth determining unit, configured to determine that the first target user is a normal user if the user score is greater than or equal to a score threshold; A sixth determining unit is configured to determine that the first target user is an abnormal user if the user score is less than a score threshold.
14. The abnormal user detection device according to claim 9 or 13, characterized in that: The device further comprises: A sending module is used to send a reminder message to a second target user when it is determined that the first target user is an abnormal user; wherein, the second target user is the communication object when the first target user is the communication initiator, and the reminder message is used to inform the second target user that the first target user is an abnormal user.
15. The abnormal user detection device according to claim 14, characterized in that: The device further comprises: A receiving module, configured to receive feedback information from the second target user regarding abnormal marking of the first target user; a statistics module, configured to count the number of users who have marked the first target user as abnormal based on the feedback information; The data processing module is configured to add the first target user to an abnormal user set when the number of users who mark the first target user as abnormal is greater than a first threshold.
16. The abnormal user detection device according to claim 15, characterized in that: The data processing module includes: a second processing unit, configured to add the first target user to an abnormal marking set if the number of users who have marked the first target user as abnormal is greater than a first threshold; The third processing unit is configured to add abnormal users including the first target user in the abnormal mark set to an abnormal user set and clear the abnormal mark set when the number of users in the abnormal mark set is greater than a second threshold.
17. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the abnormal user detection method according to any one of claims 1 to 8 are implemented.
18. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the abnormal user detection method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Anti-telecommunication fraud system, method and terminal
CN107094291A
Abnormal account identification method and device
CN111507470A
Abnormal user detection method and device based on clustering analysis, equipment and medium
CN111783875A