Risk User Identification Method, Device, Electronic Device and Storage Medium
By calculating the similarity of user feature vectors and generating risk information, the problem of lagging risk user identification in the prior art is solved, and the rapid identification of new fraud patterns is achieved and the risk of fraud is effectively reduced.
Patent Information
- Application Number
- CN202210933371.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-04
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-08-04
AI Technical Summary
The risk user identification method in the prior art has a problem of identification lag, and it is impossible to identify new fraud patterns in a short period of time.
By obtaining the user feature vector, calculating the similarity between each two user feature vectors, dividing it into multiple groups, determining the risk score of each group, and obtaining the frequent item set, and generating risk information to identify the risk user.
It realizes the identification of new fraud modes in a short period of time, reduces the risk of fraud, improves the identification efficiency and accuracy, and meets the needs of real-time analysis of massive data.
Smart Images

Figure CN115392351B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of risk control technologies, and in particular, to a method, device, electronic device, and storage medium for identifying risk users. Background Art
[0002] With the development of Internet technologies, users are accustomed to registering accounts in the systems of network service providers that run relevant services, and then using the accounts as representatives of their identities to execute relevant business logics. However, risk users usually use the registered accounts to carry out criminal acts such as fraud, which not only harms the interests of enterprises but also endangers the security of users' personal information.
[0003] In the prior art, risk users are mainly identified by methods based on rule engines and methods based on supervised machine classification models. Among them, the method based on the rule engine includes converting the experience knowledge of risk control experts into fraud prevention business rules, or establishing black and white list rules for matching through the rule engine. The method based on the supervised machine classification model is to collect risk user samples, extract corresponding features, and then use supervised machine learning methods to construct a classification model to identify risk users through the classification model.
[0004] However, the method based on the rule engine relies too much on manual work and has a high cost, while the method based on the supervised machine classification model requires a large amount of labeled data, is limited by the accumulation and timeliness of labels, and both of the above methods can only identify existing fraud patterns and cannot identify new fraud patterns in a short time, resulting in the problem of lagging risk identification. Summary of the Invention
[0005] The main technical problem to be solved by this application is a method, device, electronic device, and storage medium for identifying risk users, which can solve the problem of lagging risk identification existing in the prior art.
[0006] To solve the above technical problem, the first technical solution adopted by this application is to provide a method for identifying risk users, including: obtaining a first set; where the first set includes multiple user feature vectors; calculating the similarity between every two user feature vectors in the first set; dividing the first set into multiple groups based on each user feature vector and the similarity between every two user feature vectors; determining the risk score of each group, and obtaining at least one frequent item set in each group; where each frequent item set corresponds to at least one risk identification rule; generating corresponding risk information according to the risk score of each group and the frequent item set to identify risk users through the risk information.
[0007] Among them, the steps of obtaining the first set include: collecting a preset number of multiple user data; where the user data includes structured data and unstructured data; preprocessing the multiple user data to obtain a second set; where the second set includes multiple user initial feature vectors; using information entropy to represent the weight value of each user initial feature vector; obtaining each user feature vector based on each user initial feature vector and the corresponding weight value, so as to form the first set based on the multiple user feature vectors.
[0008] Among them, the steps of preprocessing the multiple user data to obtain a second set include: statistically analyzing each structured data and unstructured data, and cleaning abnormal data based on the statistical results; sorting the numerical data based on the statistical results, and bucketing the sorted numerical data to convert the sorted numerical data into user initial feature vectors, and constructing the second set based on the multiple user initial feature vectors.
[0009] Among them, the steps of obtaining each user feature vector based on each user initial feature vector and the corresponding weight value, so as to form the first set based on the multiple user feature vectors include: encoding each user initial feature vector based on the one-hot encoding mechanism to obtain the high-dimensional vector corresponding to each user initial feature vector; where the same user corresponds to multiple user initial feature vectors; multiplying each high-dimensional vector by the corresponding weight value to obtain multiple concatenated feature vectors; concatenating the multiple concatenated feature vectors belonging to the same user to obtain each user feature vector, and forming the first set based on the multiple user feature vectors.
[0010] Among them, the steps of calculating the similarity between every two user feature vectors in the first set include: using at least one similarity algorithm to determine the similarity between every two user feature vectors in the first set.
[0011] Among them, the step of dividing the first set into multiple groups based on the similarity between each user feature vector and between every two user feature vectors includes: determining a neighborhood parameter; where the neighborhood parameter includes a clustering radius and the minimum number of each clustering sample; calculating based on the similarity between every two user feature vectors to obtain the distance between each user feature vector and the remaining user vector features; counting the multiple distances corresponding to the same user feature vector to determine the number of values less than the clustering radius among the multiple distances; in response to the number being greater than the minimum number, determining the corresponding user feature vector as a core vector and adding a corresponding label to the core vector; traversing the multiple user feature vectors in the first set to determine all the core vectors; determining multiple user feature vectors within the neighborhood of each core vector based on the neighborhood parameter and adding the same label as the corresponding core vector to the multiple user feature vectors within the neighborhood of each core vector; dividing the multiple user feature vectors with the same label into the same group.
[0012] Among them, the step of determining the risk score of each group and obtaining at least one frequent item set in each group includes: determining the aggregation degree of the group based on the similarity between every two user feature vectors in each group; using the multiple user initial feature vectors corresponding to the multiple user feature vectors in each group to determine the risk degree of the group; calculating the aggregation degree and the risk degree to determine the risk score of each group based on the calculation result; mining the multiple user initial feature vectors corresponding to the multiple user feature vectors in each group based on the risk identification rule to obtain at least one frequent item set in each group.
[0013] Among them, the step of using the multiple user initial feature vectors corresponding to the multiple user feature vectors in each group to determine the risk degree of the group includes: obtaining the value corresponding to each user initial feature vector in each group; sequentially determining whether each value is in the risk value set; in response to the value corresponding to the user initial feature vector being in the risk value set, setting the parameter corresponding to the user initial feature vector to 1; or, in response to the value corresponding to the user initial feature vector not being in the risk value set, setting the parameter corresponding to the user initial feature vector to 0; calculating the risk degree of the group using the parameter corresponding to each user initial feature vector.
[0014] Among them, the step of mining multiple user initial feature vectors corresponding to multiple user feature vectors in each group based on risk recognition rules to obtain at least one frequent item set in each group includes: dividing the values corresponding to each user initial feature vector in the group into multiple item sets; where each item set includes feature data of the same type; determining the corresponding risk recognition rule according to the data type in each item set; matching the feature data in each item set with the corresponding risk recognition rule; in response to the similarity between the feature data and the preset value in the risk recognition rule being not less than the similarity threshold, accumulating the occurrence times of the feature data; in response to the occurrence times being not less than the set threshold, determining the item set as a frequent item set.
[0015] To solve the above technical problems, the second technical solution adopted by this application is to provide a risk user identification device, including: an acquisition module, configured to acquire a first set; where the first set includes multiple user feature vectors; a calculation module, configured to calculate the similarity between every two user feature vectors in the first set; a classification module, configured to divide the first set into multiple groups by using a clustering algorithm based on the similarity between each user feature vector and every two user feature vectors; a determination module, configured to determine the risk score of each group and obtain at least one frequent item set in each group; where each frequent item set corresponds to at least one risk recognition rule; an identification module, configured to generate corresponding risk information according to the risk score of each group and the frequent item set, so as to identify risk users through the risk information.
[0016] To solve the above technical problems, the third technical solution adopted by this application is to provide an electronic device, including: a memory, configured to store program data, and when the program data is executed, it implements the steps in the risk user identification method described in any one of the above; a processor, configured to execute the program instructions stored in the memory to implement the steps in the risk user identification method described in any one of the above.
[0017] To solve the above technical problems, the fourth technical solution adopted by this application is to provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps in the risk user identification method described in any one of the above.
[0018] The beneficial effects of the present application are as follows: Different from the prior art, the present application provides a risk user identification method, device, electronic device and storage medium. By calculating the similarity between every two user feature vectors in the first set, the association between users can be effectively constructed. Then, by using each user feature vector and the similarity between every two of the user feature vectors, the first set is divided into multiple groups, and the user feature vectors with close similarity can be clustered. Further, by determining the risk score of each group and obtaining at least one corresponding frequent item set, the risk identification rule corresponding to each group can be obtained according to the frequent item set. Furthermore, the corresponding risk information can be generated based on the risk score and frequent item set of each group, and the risk users can be identified based on the risk score and risk identification rule. By constructing user feature vectors and generating risk information through the association between user feature vectors, the present application can identify new fraud patterns in a short time and identify risk users based on the new fraud patterns, thus not only meeting the need for diversified risk identification, but also achieving early prevention and control, and effectively reducing the fraud risk. In addition, since the present application does not need to rely on labeled samples and a large amount of manual work, it not only reduces the labor cost, but also improves the identification efficiency and accuracy, thus meeting the need for real-time analysis of massive data. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a flowchart of the first implementation manner of the risk user identification method of the present application;
[0021] Figure 2 It is a flowchart of the second implementation manner of the risk user identification method of the present application;
[0022] Figure 3 It is a flowchart of the third implementation manner of the risk user identification method of the present application;
[0023] Figure 4 It is a flowchart of the fourth implementation manner of the risk user identification method of the present application;
[0024] Figure 5 is Figure 4 a flowchart of a specific implementation manner of S45 in
[0025] Figure 6 is Figure 4Flow chart of a specific implementation of S47;
[0026] Figure 7 It is a work flow chart of an application scenario of the risk user identification method of the present application;
[0027] Figure 8 It is a structural schematic diagram of an implementation of the risk user identification device of the present application;
[0028] Figure 9 It is a structural schematic diagram of an implementation of the electronic device of the present application;
[0029] Figure 10 It is a structural schematic diagram of an implementation of the computer-readable storage medium of the present application. Specific implementation
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0031] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless clearly stated otherwise in the context. "Plural" generally includes at least two, but does not exclude the case of including at least one.
[0032] It should be understood that the term " / and / " used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0033] It should be understood that the term "including", "comprising" or any other variant used herein is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements.
[0034] Please refer to Figure 1 ,Figure 1 It is a schematic flowchart of the first implementation mode of the risk user identification method of this application. In this implementation mode, the risk user identification method includes:
[0035] S11: Obtain the first set; where the first set includes multiple user feature vectors.
[0036] In this implementation mode, the user feature vector is obtained by vectorizing the acquired user data.
[0037] Among them, the user feature vector includes multiple concatenated feature vectors.
[0038] Among them, each user corresponds to user data in multiple dimensions, and each dimension of user data corresponds to a feature vector. Since the importance of user data in different dimensions is different, the feature vector formed by each dimension of user data includes its corresponding weight value to represent the importance of user data through the weight value.
[0039] It can be understood that based on the same user, multi-dimensional user data can be obtained. Constructing user feature vectors based on multi-dimensional user data and constructing the first set can effectively establish the association between users.
[0040] S12: Calculate the similarity between every two user feature vectors in the first set.
[0041] Among them, similarity is a measure for comprehensively evaluating the proximity between two things. The closer two things are, the greater their similarity measure, and the more distant two things are, the smaller their similarity measure.
[0042] In this implementation mode, at least one similarity algorithm is used to determine the similarity between every two user feature vectors in the first set.
[0043] Specifically, any one of the Jaccard Coefficient, Cosine Similarity, Euclidean Distance, Pearson Correlation Coefficient, Kullback-Leibler Divergence, Tanimoto coefficient (generalized Jaccard similarity coefficient), and Mutual Information can be used to determine the similarity between every two user feature vectors in the first set. This application does not make any limitations in this regard.
[0044] S13: Divide the first set into multiple groups based on the similarity between each user feature vector and between every two user feature vectors.
[0045] In this embodiment, a clustering algorithm is used to divide multiple user feature vectors in the first set into multiple groups.
[0046] Specifically, the similarity between user feature vectors in the same group is relatively high, and the similarity between user feature vectors in different groups is relatively low. Therefore, risk users can be identified from multiple users to be identified according to the distribution information of user feature vectors in the multiple groups obtained by the division.
[0047] It can be understood that determining risk users based on the groups of user feature vectors can greatly reduce the probability that the corresponding user accounts are associated with risk groups, making the identification of risk users more accurate.
[0048] S14: Determine the risk score of each group, and obtain at least one frequent item set in each group; wherein, each frequent item set corresponds to at least one risk identification rule.
[0049] In this embodiment, the risk score is calculated based on the similarity and risk level between each user vector feature.
[0050] In a specific implementation scenario, in response to the calculated risk score not exceeding the preset risk score threshold, determine that the users in the corresponding group are low-risk or medium-risk users. In another specific implementation scenario, in response to the calculated risk score exceeding the preset risk score threshold, determine that the users in the corresponding group are high-risk users.
[0051] In this embodiment, a frequent item set refers to a set whose support is greater than or equal to the minimum support. Support refers to the frequency of a certain set appearing in all services.
[0052] Among them, each frequent item set corresponds to at least one risk identification rule, and the risk identification rules corresponding to different frequent item sets are not repeated. Among them, the user feature vectors divided into the same group all have corresponding risk identification rules.
[0053] The risk identification rule can be considered as the reason (clustering reason) why the user feature vectors in the group are clustered. For example, if more than 50% of the user feature vectors in a certain group include the same IP address data marked as risky, then using the IP address data is the risk identification rule corresponding to this frequent item set, and the reason why this group is clustered is that most users use the same risky IP address data.
[0054] Understandably, by covering as many groups as possible with frequent itemsets, it is possible to ensure that the risk identification rules corresponding to the frequent itemsets are more accurate, thereby comprehensively reflecting the relationship between user characteristic data and risks.
[0055] S15: Generate corresponding risk information based on the risk scores of each group and the frequent itemsets to identify risk users through the risk information.
[0056] In this embodiment, not only the reasons for clustering are mined from the frequent itemsets for the groups corresponding to high-risk users, but also the reasons for clustering can be mined from the frequent itemsets for low-risk users or medium-risk users to generate more risk information.
[0057] Among them, different groups correspond to different frequent itemsets, and different frequent itemsets correspond to different risk identification rules. After the subsequent client obtains the risk information generated based on the risk scores and the frequent itemsets, it can use different risk identification rules to identify different risk operations, and then perform different risk treatments for different risk operations.
[0058] In a specific implementation scenario, in response to the risk score in the risk information of a certain group being a value exceeding the preset risk score threshold, it indicates that the users in this group are high-risk users. Directly use the risk identification rules included in the risk information to obtain the corresponding risk treatment method, and control the user account based on the risk treatment method, thereby realizing the risk control of the client.
[0059] Understandably, by performing risk control through the generated risk information, the efficiency and real-time performance of risk control can be greatly improved, thereby solving the problem of lagging risk identification.
[0060] Different from the prior art, this embodiment generates risk information by constructing user feature vectors and through the association between user feature vectors, can identify new fraud patterns in a short time, and identify risk users based on the new fraud patterns, thus not only meeting the need for diversified risk identification, but also achieving early prevention and control, and then effectively reducing the fraud risk. In addition, since this embodiment does not need to rely on labeled samples and a large amount of manual work, it not only reduces the labor cost, but also can improve the identification efficiency and accuracy, thus meeting the need for real-time analysis of massive data.
[0061] Please refer to Figure 2 , Figure 2 is a schematic flowchart of the second embodiment of the risk user identification method of this application. In this embodiment, the information entropy is used to represent the weight value of each user's initial feature vector, and the user feature vector is constructed. The risk user identification method includes:
[0062] S21: Collect a preset number of multiple user data; among them, the user data includes structured data and unstructured data.
[0063] In this embodiment, the user data is collected in the way of a sliding window.
[0064] In a specific implementation scenario, set the interval of each movement as T and the time window as 2T, so that there is an overlapping area between two windows, thereby ensuring that the user association between different time windows will not be lost.
[0065] It can be understood that in actual business, the data volume of the whole day is too large. If the user data at any moment is collected and analyzed, it will lead to too long calculation time. Collecting data in the way of a sliding window can effectively improve the calculation efficiency without losing the user association information.
[0066] In this embodiment, the structured data is the application service data of the user, and the unstructured data is the device data, operation data and social data of the user.
[0067] Specifically, the device data can be the user's mobile phone number data, GPS (Global Positioning System) positioning data, MAC (Device Identification Information) address data, IP (Internet Protocol) address data, etc. The operation data can be the data when the user uses the application program, such as inputting the ID number, mobile phone number, bank card number, etc. The social data can be the data when the user communicates with other users on the application program.
[0068] S22: Preprocess the multiple user data to obtain a second set; among them, the second set includes multiple user initial feature vectors.
[0069] In this embodiment, first, each piece of structured data and unstructured data is statistically analyzed, and the abnormal data is cleaned based on the statistical results.
[0070] Among them, since the unstructured data cannot be used for subsequent calculations, after the unstructured data is obtained, it is usually structured and then statistically analyzed.
[0071] Among them, the abnormal data is the data not within the preset interval. For example, the normal value range corresponding to a certain type of feature is 10% - 90%. If the value of a certain feature in this type obtained is 95%, then this feature is abnormal data.
[0072] Further, the numerical data is sorted based on the statistical results, and the sorted numerical data is binned (hive) to convert the sorted numerical data into user initial feature vectors.
[0073] Among them, numerical data refers to the observed values measured on a numerical scale, and the results are expressed as specific numerical values. Most of the user data obtained in this embodiment are numerical data.
[0074] In a specific implementation scenario, the numerical data can be sorted in ascending order. In another specific implementation scenario, the numerical data can also be sorted in descending order. This application does not make any restrictions in this regard.
[0075] Among them, hive refers to mapping structured data into a database table and providing SQL (Structured Query Language) - like functions.
[0076] The reasonable partitions formed in Hive provide a convenient way to isolate data and optimize queries. When dealing with large - scale data sets, a part of the entire data set can be used for sampling test queries and modifications, thus making development more efficient.
[0077] Furthermore, a second set is constructed based on multiple user initial feature vectors.
[0078] Since the same user corresponds to user data in multiple dimensions, and the user data in the same dimension corresponds to a type feature, the same user corresponds to multiple user initial feature vectors.
[0079] In this embodiment, the second set is represented as X = {x 1 [1], x 2 [1], …, x i [j], …, x m [n]}, x i ∈R n , where x i [j] represents the value of the i - th user on the j - th type of feature.
[0080] S23: Use information entropy to characterize the weight value of each user initial feature vector.
[0081] Among them, information entropy is a rather abstract concept in mathematics. Information entropy can be understood as the occurrence probability of a certain specific information (random variable), or as the "average value" of the amount of information of each event in a random variable, that is, the mathematical expectation of the amount of information.
[0082] Considering that different user initial feature vectors have different degrees of importance, in this embodiment, based on information theory, information entropy is used to characterize the weight value of each user initial feature vector.
[0083] Specifically, the information entropy is calculated through the following formula:
[0084]
[0085] Among them, p k represents the probability distribution of the k-th user's initial feature vector, and H represents the information entropy.
[0086] Since the probability distribution of the random variable in the information entropy calculation formula is already given, and its sample space and the probability values of each sample point in the sample space are also given, that is, the probability of the random variable is given. Therefore, the weight value calculated through the information entropy can quantitatively describe the importance degree of each user's initial feature vector.
[0087] S24: Obtain each user feature vector based on each user's initial feature vector and the corresponding weight value, so as to form a first set based on multiple user feature vectors.
[0088] In this embodiment, first, each user's initial feature vector is encoded based on the one-hot encoding mechanism to obtain the high-dimensional vector corresponding to each user's initial feature vector.
[0089] Among them, since the same user corresponds to multiple user initial feature vectors, the same user corresponds to multiple high-dimensional vectors.
[0090] Among them, one-hot encoding is also called one-hot effective encoding. Its method is to use an N-bit status register to encode N states. Each state has its independent register bit, and at any time, only one of them is valid. That is, only one bit is 1, and the rest are all zero values.
[0091] Furthermore, multiply each high-dimensional vector by the corresponding weight value to obtain multiple concatenated feature vectors.
[0092] Among them, the corresponding weight value is the information entropy calculated above.
[0093] Furthermore, concatenate the multiple concatenated feature vectors belonging to the same user to obtain each user feature vector, and form a first set based on multiple user feature vectors.
[0094] In this embodiment, the user feature vector is represented as Among them, H i represents the weight value corresponding to the i-th user's initial feature vector, v i represents the high-dimensional vector corresponding to the i-th user's initial feature vector, H i v i represents the concatenated vector corresponding to the i-th user's initial feature vector, and U represents the user feature vector.
[0095] Among them, each user corresponds to a user feature vector, and the first set includes the user feature vectors of multiple users.
[0096] S25: Calculate the similarity between every two user feature vectors in the first set.
[0097] In this embodiment, the Jaccard similarity coefficient is used to calculate the similarity between every two user feature vectors in the first set to characterize the association between users.
[0098] Specifically, the similarity calculation formula is as follows:
[0099]
[0100] Among them, U i and U j respectively represent the i-th user feature vector and the j-th user feature vector in the first set, represents the intersection elements of the i-th user feature vector and the j-th user feature vector, represents the union elements of the i-th user feature vector and the j-th user feature vector, Sim(U i ,U j ) represents the similarity between the i-th user feature vector and the j-th user feature vector.
[0101] Among them, the greater the similarity, the more common features the two user feature vectors have.
[0102] S26: Based on the similarity between each user feature vector and every two user feature vectors, divide the first set into multiple groups.
[0103] For the specific process, please refer to the description in S13 and will not be elaborated here.
[0104] S27: Determine the risk score of each group and obtain at least one frequent item set in each group; among them, each frequent item set corresponds to at least one risk identification rule.
[0105] For the specific process, please refer to the description in S14 and will not be elaborated here.
[0106] S28: Generate corresponding risk information according to the risk score of each group and the frequent item set to identify risk users through the risk information.
[0107] For the specific process, please refer to the description in S15 and will not be elaborated here.
[0108] Different from the prior art, in this embodiment, information entropy is used to characterize the weight value of each user initial feature vector, which can quantitatively describe the importance of each user initial feature vector, thereby effectively constructing the association between users.
[0109] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of the third embodiment of the risk user identification method of this application. In this embodiment, a clustering algorithm is used to divide the first set into multiple groups. The risk user identification method includes:
[0110] S301: Obtain the first set; wherein, the first set includes multiple user feature vectors.
[0111] For the specific process, please refer to the descriptions in S13 and S21 - S24, which will not be elaborated here.
[0112] S302: Calculate the similarity between every two user feature vectors in the first set.
[0113] For the specific process, please refer to the description in S25, which will not be elaborated here.
[0114] S303: Determine the neighborhood parameters; wherein, the neighborhood parameters include the clustering radius and the minimum number of each clustering sample.
[0115] In this embodiment, the DBSCAN (Density - Based Spatial Clustering of Applications with Noise) clustering algorithm is used to cluster multiple user feature vectors in the first set.
[0116] Among them, DBSCAN is a relatively representative density - based clustering algorithm. It defines a cluster as the largest set of density - connected points, can divide regions with sufficient high density into clusters, and can discover clusters of any shape in a spatial database with noise.
[0117] Among them, the clustering radius refers to the radius of the ε - neighborhood in DBSCAN. Among them, the ε - neighborhood refers to the region within a radius of ε for a given object.
[0118] Among them, in DBSCAN, if the number of sample points within the ε - neighborhood of a given object is greater than or equal to the minimum number of each clustering sample (MinPts), then the object is called a core object. For the convenience of narration in this embodiment, the core object is called the core vector.
[0119] In this embodiment, the neighborhood parameters are obtained through fine - tuning after analyzing a large amount of data.
[0120] S304: Calculate based on the similarity between every two user feature vectors to obtain the distance between each user feature vector and the feature vectors of the remaining users.
[0121] In this embodiment, the distance between each user feature vector and the feature vectors of the remaining users is calculated by the following formula:
[0122] r = 1 / Sim(U i , U j )
[0123] where r is the distance between the i-th user feature vector and the j-th user feature vector, and Sim(U i , U j ) represents the similarity between the i-th user feature vector and the j-th user feature vector.
[0124] S305: Count the multiple distances corresponding to the same user feature vector and determine the number of distances whose values are less than the clustering radius.
[0125] In this embodiment, by counting the multiple distances corresponding to the same user feature vector and determining the number of distances whose values are less than the clustering radius, it is judged whether the user feature vector is a core vector.
[0126] S306: In response to the number being greater than the minimum number, determine the corresponding user feature vector as a core vector and add a corresponding label to the core vector.
[0127] In this embodiment, in response to the number being greater than the minimum number, it indicates that the number of sample points within the ε-neighborhood centered on a certain user feature vector is greater than or equal to MinPts, meeting the requirements for a core object in DBSCAN, and determine the user feature vector as a core vector.
[0128] Further, add a corresponding label to the core vector.
[0129] S307: Traverse the multiple user feature vectors in the first set to determine all core vectors.
[0130] In this embodiment, use the above method to traverse the multiple user feature vectors in the first set to determine all core vectors and add corresponding labels.
[0131] Among them, each core vector has a different label.
[0132] S308: Based on the neighborhood parameter, determine the multiple user feature vectors located within the neighborhood of each core vector, and add the same label as the corresponding core vector to the multiple user feature vectors located within the neighborhood of each core vector.
[0133] In this embodiment, user feature vectors of multiple non-core vectors located within the ε-neighborhood of each core vector are determined based on the clustering radius.
[0134] Specifically, if a certain core vector is located within the ε-neighborhood of another core vector, this core vector will not be assigned to the cluster to which the other core vector belongs. That is, only non-core vectors will be assigned to the cluster to which the corresponding core vector belongs and will be added with the same label as the corresponding core vector.
[0135] S309: Partition multiple user feature vectors with the same label into the same group.
[0136] In this embodiment, a set of groups is obtained through the above method, and the set of groups is represented as C = {C 1 , C 2 , …, C i …, C k}, where C i represents the user group with label i, and k represents that there are a total of k user groups.
[0137] It can be understood that the similarity between user feature vectors with the same label is relatively large and the correlation is relatively strong.
[0138] S310: Determine the risk score of each group and obtain at least one frequent item set in each group; where each frequent item set corresponds to at least one risk identification rule.
[0139] For the specific process, please refer to the description in S14 and will not be elaborated here.
[0140] S311: Generate corresponding risk information based on the risk score of each group and the frequent item set, so as to identify risk users through the risk information.
[0141] For the specific process, please refer to the description in S15 and will not be elaborated here.
[0142] Different from the prior art, in this embodiment, a clustering algorithm is used to partition multiple user feature vectors in the first set into multiple groups, which can discover groups with different characteristics to improve the accuracy of the user groups to which users belong, thereby greatly reducing the probability that the corresponding user accounts are associated with risk groups and making the identification of risk users more accurate.
[0143] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of the fourth embodiment of the risk user identification method of this application. In this embodiment, the risk score of the group is calculated using the similarity and the user initial feature vectors, and the frequent item set of the group is obtained based on the risk identification rules. The risk identification method includes:
[0144] S41: Obtain the first set; wherein, the first set includes multiple user feature vectors.
[0145] For the specific process, please refer to the descriptions in S13 and S21 - S24, which will not be elaborated here.
[0146] S42: Calculate the similarity between every two user feature vectors in the first set.
[0147] For the specific process, please refer to the description in S25, which will not be elaborated here.
[0148] S43: Based on each user feature vector and the similarity between every two user feature vectors, divide the first set into multiple groups.
[0149] For the specific process, please refer to the descriptions in S303 - S309, which will not be elaborated here.
[0150] S44: Determine the aggregation degree of each group based on the similarity between every two user feature vectors in each group.
[0151] In this embodiment, the aggregation degree is calculated by the following formula:
[0152]
[0153] wherein, I is the aggregation degree, C k represents the user group with label k, U i and U j respectively represent the i-th user feature vector and the j-th user feature vector in C k , and Sim(U i , U j ) represents the similarity between the i-th user feature vector and the j-th user feature vector.
[0154] Among them, dividing by 2 is to eliminate the influence of duplicate calculations to improve the calculation accuracy.
[0155] S45: Determine the risk degree of each group by using multiple user initial feature vectors corresponding to multiple user feature vectors in each group.
[0156] Specifically, please refer to Figure 5 , Figure 5 which Figure 4 is the flowchart of a specific implementation manner of S45 in
[0157] S451: Obtain the values corresponding to each user initial feature vector in each group.
[0158] In this embodiment, the value is the value corresponding to the user's initial feature vector.
[0159] S452: Determine whether each value is in the risk value set in sequence.
[0160] In this embodiment, the risk values in the risk value set are determined. The value is compared with the risk values in the risk value set to determine whether the value is in the risk value set.
[0161] For example, the risk values in the risk value set are from 10 to 90. If a certain value is 80, it is in the risk value set.
[0162] For another example, a certain range of IP address data is in the risk value set. If the IP values in a certain group are all in this risk value set, the risk level of this group is relatively high.
[0163] S453: In response to the value corresponding to the user's initial feature vector being in the risk value set, set the parameter corresponding to the user's initial feature vector to 1; or, in response to the value corresponding to the user's initial feature vector not being in the risk value set, set the parameter corresponding to the user's initial feature vector to 0.
[0164] S454: Calculate the risk level of the group using the parameter corresponding to each user's initial feature vector.
[0165] In this embodiment, the risk level of the group is calculated in the following manner:
[0166]
[0167] Among them, R is the risk level of the group, C k represents the user group with label k, N is the number of user feature vectors in C k , and x i [p]inF p is the parameter corresponding to the p-th user initial feature vector of the i-th user.
[0168] S46: Calculate the aggregation degree and the risk level to determine the risk score of each group based on the calculation results.
[0169] In this embodiment, the risk score of the group is calculated in the following manner:
[0170] RiskScore = I + R
[0171] Among them, RiskScore is the risk score of the group, I is the aggregation degree of the group, and R is the risk level of the group.
[0172] S47: Mining the multiple initial user feature vectors corresponding to the multiple user feature vectors in each group based on the risk identification rules to obtain at least one frequent item set in each group.
[0173] Specifically, please refer to Figure 6 , Figure 6 which Figure 4 is the schematic flowchart of a specific implementation manner of S47 in . In this implementation manner, the steps of mining the multiple initial user feature vectors corresponding to the multiple user feature vectors in each group based on the risk identification rules to obtain at least one frequent item set in each group specifically include:
[0174] S471: Divide the values corresponding to each initial user feature vector in the group into multiple item sets; where each item set includes feature data of the same type.
[0175] In this implementation manner, item set classification can be performed based on application service data, device data, operation data, and social data. For example, the IP address data in the device data can be divided into one item set, or the GPS positioning data in the device data can be divided into one item set.
[0176] S472: Determine the corresponding risk identification rules according to the data type in each item set.
[0177] In a specific implementation scenario, if the feature data in the item set is IP address data, the corresponding risk identification rule can be to determine the IP address data belonging to a specific range as a risky IP.
[0178] In another specific implementation scenario, if the feature data in the item set is GPS positioning data, the corresponding identification rule can be to determine the GPS positioning data belonging to a specific location as a risky location.
[0179] S473: Match the feature data in each item set with the corresponding risk identification rules.
[0180] In a specific implementation scenario, if the feature data in the item set is IP address data and the corresponding risk identification rule is to determine the IP address data belonging to a specific range as a risky IP, then compare each IP address data in the item set with the IP address data in the specific range.
[0181] In another specific implementation scenario, if the feature data in the item set is GPS positioning data and the corresponding identification rule is to determine the GPS positioning data belonging to a specific location as a risky location, then compare each GPS positioning data in the item set with the GPS positioning data in the specific location.
[0182] S474: In response to the similarity between the feature data and the preset value in the risk identification rule being not less than the similarity threshold, accumulate the occurrence times of the feature data.
[0183] In a specific implementation scenario, if the similarity between a certain IP address data in the item set and the IP address data within a specific range is not less than the preset similarity threshold, increment the accumulated times by 1.
[0184] In another specific implementation scenario, if the similarity between a certain GPS positioning data in the item set and the GPS positioning data of a specific location is not less than the preset similarity threshold, increment the accumulated times by 1.
[0185] S475: In response to the occurrence times being not less than the set threshold, determine the item set as a frequent item set.
[0186] In this embodiment, the value obtained by multiplying the total number of feature data in the item set by the set ratio can be determined as the preset threshold. Here, the set ratio can be 50%, 60% or other ratios, and this application does not limit it.
[0187] It can be understood that by mining frequent item sets, at least one reason for the formation of clusters in the group can be determined, that is, due to the same value of a certain type of data in the group, the similarity between them is relatively large.
[0188] It can be understood that by covering as many groups as possible with frequent item sets, it can be ensured that the risk identification rules corresponding to the frequent item sets are more accurate, so as to comprehensively reflect the relationship between user feature data and risks.
[0189] S48: Generate corresponding risk information according to the risk scores of each group and the frequent item sets, so as to identify risk users through the risk information.
[0190] In this embodiment, by refining the frequent item sets and the corresponding risk identification rules in the risk information, a business rule (rule) with a specific format can be generated. For example, the format of the business rule can be rule = {feature data 1 : value 1 , feature data 2 : value 2 , …, feature data i : value i , …, feature data m : value m}, where feature data i : value i refers to the value of the i-th feature data in the i-th type of feature data, and m refers to a total of m types of feature data.
[0191] For example, if a group corresponds to two frequent itemsets, one of which is IP address data and the other is GPS positioning data, then rule = {IP address data: specific value of IP address data, GPS positioning data: specific value of GPS positioning data}. Based on this rule, the reason for the formation of the group can be obtained.
[0192] In this embodiment, risk information can be generated in a specific format. For example, risk information = {group: C k , group risk score: RiskScore, reason for group formation: rule, group users: {user 1 , user 2 , …, user i , …, user m}, where C k represents the user group with label k, RiskScore is the risk score of the group, rule is the business rule, user i is the i-th user in the group, and m is the total number of users in the group.
[0193] Furthermore, the generated risk information is provided to the risk control business personnel so that they can perform risk identification on the user to be identified based on the risk information.
[0194] In a specific implementation scenario, automated risk control can be performed on the risk information based on a rule engine, such as intercepting the identified risk users. In another specific implementation scenario, manual sampling and evaluation can be performed on the automatically intercepted risk users to detect whether there are miskill situations, and the results are fed back to the corresponding algorithm of the above method for iterative optimization. In yet another specific implementation scenario, the risk information can also be used as a portrait factor for different risk users and provided to the subsequent supervised scoring model.
[0195] Please refer to Figure 7 , Figure 7It is a workflow diagram of an application scenario of the risk user identification method of this application. In this embodiment, after obtaining user data, first, the ruleless learning engine calculates and analyzes the user data to generate multiple user feature vectors. Then, calculate the similarity between every two user feature vectors in the first set, and based on each user feature vector and the similarity between every two user feature vectors, divide the multiple user feature vectors into multiple groups. Then determine the risk score of each group, and obtain at least one frequent item set in each group, so as to generate corresponding risk information according to the risk score and frequent item set of each group, and identify risk users through the risk information. Furthermore, input the risk information into the rule engine, and based on the rule engine, use the risk information for automated risk control, and conduct manual sampling and evaluation on the risk users automatically intercepted to detect whether there is a mis-kill situation, and feedback the result to the corresponding algorithm of the above method for iterative optimization.
[0196] The risk user identification method of this embodiment can be applied to multiple scenarios such as social media registration, login, new user activation, and bullet screens. The inventors of this application have tested and found that using the risk user identification method provided by this application in the registration scenario can detect and control more than 90% of the risk users within 1 hour.
[0197] Different from the prior art, this embodiment constructs user feature vectors through information entropy, and generates risk information through the association between user feature vectors, which can identify new fraud patterns in a short time, and identify risk users based on the new fraud patterns, thus not only meeting the need for diversified risk identification, but also achieving early prevention and control, and then effectively reducing the fraud risk. In addition, since this embodiment does not need to rely on labeled samples and a large amount of manual work, it not only reduces the labor cost, but also improves the identification efficiency and accuracy, so as to meet the need for real-time analysis of massive data.
[0198] Correspondingly, this application provides a risk user identification device.
[0199] Please refer to Figure 8 , Figure 8 It is a schematic structural diagram of an embodiment of the risk user identification device of this application. As Figure 8 shown, the risk user identification device 80 includes an acquisition module 81, a calculation module 82, a classification module 83, a determination module 84, and an identification module 85.
[0200] The acquisition module 81 is used to obtain the first set; wherein, the first set includes multiple user feature vectors.
[0201] The calculation module 82 is used to calculate the similarity between every two user feature vectors in the first set.
[0202] A classification module 83, configured to divide the first set into multiple groups by using a clustering algorithm based on the similarity between each user feature vector and between every two user feature vectors.
[0203] A determination module 84, configured to determine the risk score of each group and obtain at least one frequent item set in each group; wherein each frequent item set corresponds to at least one risk identification rule.
[0204] An identification module 85, configured to generate corresponding risk information according to the risk score of each group and the frequent item set, so as to identify risk users through the risk information.
[0205] Among them, for the specific process, please refer to the relevant textual descriptions in S11 - S15, S21 - S28, S301 - S311, and S41 - S48, which will not be elaborated here.
[0206] Different from the prior art, in this embodiment, the acquisition module 81 is used to construct user feature vectors, and the determination module 84 is used to generate risk information by using the association between user feature vectors. The identification module 85 can identify new fraud patterns in a short time and identify risk users based on the new fraud patterns, thereby not only meeting the requirement of diversified risk identification but also achieving early prevention and control, and then effectively reducing the fraud risk. In addition, since this embodiment does not need to rely on labeled samples and a large amount of manual work, it not only reduces the labor cost but also improves the identification efficiency and accuracy, so as to meet the requirement of real-time analysis of massive data.
[0207] Correspondingly, the present application provides an electronic device.
[0208] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of an embodiment of the electronic device of the present application. As Figure 9 shown, in this embodiment, the electronic device 90 includes a memory 91 and a processor 92.
[0209] In this embodiment, the memory 91 is used to store program data, and when the program data is executed, it implements the steps in the risk user identification method described in any one of the above; the processor 92 is used to execute the program instructions stored in the memory 91 to implement the steps in the risk user identification method described in any one of the above.
[0210] Specifically, the processor 92 is used to control itself and the memory 91 to implement the steps in any of the above risk user identification methods. The processor 92 can also be referred to as a CPU (Central Processing Unit). The processor 92 may be an integrated circuit chip with signal processing capabilities. The processor 92 can also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Additionally, the processor 92 can be implemented jointly by multiple integrated circuit chips.
[0211] Different from the prior art, in this embodiment, the processor 92 constructs user feature vectors and generates risk information through the association between user feature vectors, which can identify new fraud patterns in a short time and identify risk users based on the new fraud patterns, thus not only meeting the requirement for diversified risk identification but also achieving early prevention and control, and then effectively reducing the fraud risk. In addition, since this embodiment does not need to rely on labeled samples and a large amount of manual work, it not only reduces the labor cost but also improves the identification efficiency and accuracy, thus meeting the requirement for real-time analysis of massive data.
[0212] Correspondingly, the present application provides a computer-readable storage medium.
[0213] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application.
[0214] The computer-readable storage medium 100 includes a computer program 1001 stored on the computer-readable storage medium 100. When the computer program 1001 is executed by the above-mentioned processor, it implements the steps in the risk user identification method as described in any of the above.
[0215] Specifically, when the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium 100. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium 100 and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned computer-readable storage medium 100 includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0216] In several embodiments provided in this application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical, or other forms.
[0217] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0218] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0219] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods according to various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0220] The above are only the embodiments of this application, and do not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A risk user identification method, characterized in that, it includes: Obtain a first set; wherein, the first set includes multiple user feature vectors; Calculate the similarity between every two of the user feature vectors in the first set; Based on each user feature vector and the similarity between every two of the user feature vectors, divide the first set into multiple groups, including: determining a neighborhood parameter; wherein, the neighborhood parameter includes a clustering radius and the minimum number of each clustering sample; Calculate based on the similarity between every two of the user feature vectors to obtain the distance between each user feature vector and the remaining user vector features; Statistically analyze the multiple distances corresponding to the same user feature vector, and determine the number of values in the multiple distances that are less than the clustering radius; In response to the number being greater than the minimum number, determine the corresponding user feature vector as a core vector, and add a corresponding label to the core vector; Traverse the multiple user feature vectors in the first set to determine all core vectors; Based on the neighborhood parameter, determine multiple user feature vectors within the neighborhood of each core vector, and add the same label as the corresponding core vector to the multiple user feature vectors within the neighborhood of each core vector; Divide the multiple user feature vectors with the same label into the same group; Determine the risk score of each group, and obtain at least one frequent item set in each group; wherein, each frequent item set corresponds to at least one risk identification rule; Generate corresponding risk information according to the risk score of each group and the frequent item set, so as to identify risk users through the risk information.
2. The risk user identification method according to claim 1, characterized in that, The step of obtaining the first set includes: Collect a preset number of multiple user data; wherein, the user data includes structured data and unstructured data; Preprocess the multiple user data to obtain a second set; wherein, the second set includes multiple user initial feature vectors; Use information entropy to characterize the weight value of each user initial feature vector; Obtain each user feature vector based on each user initial feature vector and the corresponding weight value, so as to form the first set based on the multiple user feature vectors.
3. The risk user identification method according to claim 2, characterized in that, The step of preprocessing the multiple user data to obtain a second set includes: Statistically analyze each structured data and the unstructured data, and clean abnormal data based on the statistical results; Sort the numerical data based on the statistical results, and bucket the sorted numerical data to convert the sorted numerical data into the user initial feature vectors, and construct the second set based on the multiple user initial feature vectors.
4. The risk user identification method according to claim 3, wherein, the step of obtaining each of the user feature vectors based on each of the user initial feature vectors and the corresponding weight values, and forming the first set based on the multiple user feature vectors, includes: encoding each of the user initial feature vectors based on a one-hot encoding mechanism to obtain a high-dimensional vector corresponding to each of the user initial feature vectors; wherein, the same user corresponds to multiple user initial feature vectors; multiplying each of the high-dimensional vectors by the corresponding weight value to obtain a plurality of spliced feature vectors; splicing the multiple spliced feature vectors belonging to the same user to obtain each of the user feature vectors, and forming the first set based on the multiple user feature vectors.
5. The risk user identification method according to claim 4, wherein, the step of calculating the similarity between every two of the user feature vectors in the first set includes: using at least one similarity algorithm to determine the similarity between every two of the user feature vectors in the first set.
6. The risk user identification method according to claim 1 or 5, wherein, the step of determining the risk score of each of the groups and obtaining at least one frequent item set in each of the groups includes: determining the aggregation degree of the group based on the similarity between every two of the user feature vectors in each of the groups; determining the risk level of the group by using the multiple user initial feature vectors corresponding to the multiple user feature vectors in each of the groups; calculating the aggregation degree and the risk level to determine the risk score of each of the groups based on the calculation result; mining the multiple user initial feature vectors corresponding to the multiple user feature vectors in each of the groups based on the risk identification rule to obtain at least one frequent item set in each of the groups.
7. The risk user identification method according to claim 6, wherein, the step of determining the risk level of the group by using the multiple user initial feature vectors corresponding to the multiple user feature vectors in each of the groups includes: obtaining the value corresponding to each of the user initial feature vectors in each of the groups; sequentially determining whether each of the values is in the risk value set; in response to the value corresponding to the user initial feature vector being in the risk value set, setting the parameter corresponding to the user initial feature vector to 1; or, in response to the value corresponding to the user initial feature vector not being in the risk value set, setting the parameter corresponding to the user initial feature vector to 0; calculating the risk level of the group by using the parameter corresponding to each of the user initial feature vectors.
8. The risk user identification method according to claim 7, wherein, The step of mining the multiple user initial feature vectors corresponding to the multiple user feature vectors in each of the groups based on the risk identification rule to obtain at least one frequent item set in each of the groups includes: Dividing the values corresponding to each user initial feature vector in the group into multiple item sets; wherein, each item set includes feature data of the same type; Determining a corresponding risk identification rule according to the data type in each item set; Matching the feature data in each item set with the corresponding risk identification rule; In response to the similarity between the feature data and a preset value in the risk identification rule being not less than a similarity threshold, accumulating the occurrence times of the feature data; In response to the occurrence times being not less than a set threshold, determining the item set as the frequent item set.
9. A risk user identification device, characterized in that, it includes: An acquisition module, configured to acquire a first set; wherein, the first set includes multiple user feature vectors; A calculation module, configured to calculate the similarity between every two user feature vectors in the first set; A classification module, configured to divide the first set into multiple groups by using a clustering algorithm based on the similarity between each user feature vector and the similarity between every two user feature vectors, including: determining a neighborhood parameter; wherein, the neighborhood parameter includes a clustering radius and the minimum number of each clustering sample; calculating based on the similarity between every two user feature vectors to obtain the distance between each user feature vector and the remaining user vector features; counting the number of values less than the clustering radius among the multiple distances corresponding to the same user feature vector; in response to the number being greater than the minimum number, determining the corresponding user feature vector as a core vector and adding a corresponding label to the core vector; traversing the multiple user feature vectors in the first set to determine all core vectors; determining multiple user feature vectors located in the neighborhood of each core vector based on the neighborhood parameter and adding the same label as the corresponding core vector to the multiple user feature vectors located in the neighborhood of each core vector; dividing the multiple user feature vectors with the same label into the same group; A determination module, configured to determine the risk score of each group and obtain at least one frequent item set in each group; wherein, each frequent item set corresponds to at least one risk identification rule; An identification module, configured to generate corresponding risk information according to the risk score of each group and the frequent item set, so as to identify risk users through the risk information.
10. An electronic device, characterized in that, it includes: A memory, configured to store program data, and when the program data is executed, it implements the steps in the risk user identification method according to any one of claims 1 to 8; A processor, configured to execute the program data stored in the memory.
11. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the risk user identification method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Abnormal group detection method based on unsupervised algorithm
CN113919415A