Abnormal data identification method and device, computer device, and storage medium

By constructing social network graphs and analyzing user communication frequency, login information, transaction information, and account cancellation behavior, financial fraud risks can be identified. This solves the problem of accurately identifying fraudulent behavior of abnormal user groups in existing technologies and achieves more efficient abnormal data identification.

CN120163585BActive Publication Date: 2025-11-18CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510272188.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-11-18
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

Existing user behavior analysis methods and information communication content correlation mining are insufficient to accurately identify fraudulent behavior by abnormal user groups, especially those who exploit potential loopholes in platform risk control rules to commit financial fraud.

Method used

By constructing a social network graph, risk groups are segmented based on user communication frequency, login information, and transaction information. Combined with user account information and account cancellation behavior, potential abnormal users are identified.

Benefits of technology

It improves the accuracy and reliability of identifying abnormal data among users that pose a risk of fraud, and effectively identifies potential financial fraud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163585B_ABST
    Figure CN120163585B_ABST
Patent Text Reader

Abstract

The application relates to an abnormal data identification method and device, computer equipment and a storage medium. The method comprises the following steps: constructing a user social network graph according to user communication information; dividing a high-risk group according to user login information and user transaction information; marking the high-risk group in the user social network graph to obtain a potential abnormal group; marking a target user to be identified in the user login information and user account information of the potential abnormal group; obtaining the number of logouts of the target user to be identified, and marking the target user to be identified as target abnormal data according to the number of logouts. The application can be applied to a financial anti-fraud application scene, can effectively realize accurate identification of target abnormal data with fraud risks in users, and improves the reliability of identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to an abnormal data identification method, apparatus, computer equipment, and storage medium. Background Technology

[0002] With the rapid development of technology, certain abnormal user groups exist in the financial sector, using various technologies and methods to commit financial fraud. These abnormal user groups use encrypted communication software or specific coding languages ​​to transmit covert instructions and conduct illegal fund settlements. They conduct multi-layered nested transactions through shell company accounts and encrypted wallets, and their fund chains intersect with the daily financial activities of ordinary users in multiple ways. For example, high-frequency and legitimate transaction scenarios such as automatic salary payments for wealth management, automatic deduction of utility bills, automatic installment repayments, and medical insurance can all be implanted with fraudulent instructions disguised as "wealth management redemption" or "insurance premium deduction," thus resulting in financial fraud.

[0003] Traditional user behavior analysis methods and information communication content correlation mining can, to some extent, conduct risk assessments based on user behavior patterns and information exchange characteristics. However, because the boundaries between the behavior patterns of abnormal user groups and normal business operations are relatively blurred, it is difficult to accurately capture abnormal signals based solely on user behavior fit and communication content analysis.

[0004] Although existing user flow monitoring and anomaly detection systems can detect abnormal behavior through preset rules (such as account registration and cancellation frequency, transaction amount thresholds, and transaction frequency limits), abnormal user groups not only exploit potential loopholes in the platform's risk control rules, but also constantly adjust their strategies to adapt to rule updates, making their activities appear to be completely in line with normal user behavior patterns. This poses certain difficulties for the accurate identification and investigation of abnormal data in financial security and medical data. Summary of the Invention

[0005] The purpose of this application is to provide an abnormal data identification method, apparatus, computer equipment, and storage medium to solve the problem of being unable to accurately identify target abnormal data of users that pose a risk of fraud.

[0006] In a first aspect, embodiments of this application provide an abnormal data identification method, which adopts the following technical solution:

[0007] Extract user communication information within a preset time range from the database, and perform frequency statistics on the user communication information to obtain the user's communication frequency;

[0008] Determine whether the user's communication frequency is greater than a preset communication frequency threshold;

[0009] If the frequency of communication between users is greater than the preset communication frequency threshold, then the user objects corresponding to the frequency of communication between users are determined to have a target relationship, and the set of all user objects with the target relationship is used as nodes of the social network graph. The communication relationships of user objects with the target relationship are used as edges of the social network graph to construct the user social network graph.

[0010] Obtain user login information and user transaction information, and classify risk groups based on the user login information and user transaction information to obtain high-risk groups;

[0011] Based on the association and labeling of the high-risk groups in the user's social network graph, potential abnormal groups are obtained;

[0012] The login status information and user account information of the potential abnormal groups are compared and analyzed, and the target users to be identified are determined based on the analysis results;

[0013] The number of times the target user to be identified has logged out is obtained, and when the number of logouts exceeds a preset logout threshold, the target user to be identified is marked as target abnormal data.

[0014] Secondly, embodiments of this application also provide an abnormal data identification device, which adopts the following technical solution:

[0015] The frequency acquisition module is used to extract user communication information within a preset time range from the database, and to perform frequency statistics on the user communication information to obtain the user's communication frequency.

[0016] The threshold determination module is used to determine whether the user's communication frequency is greater than a preset communication frequency threshold.

[0017] The graph construction module is used to determine the user object corresponding to the user pair communication frequency as having a target relationship if the user pair communication frequency is greater than the preset communication frequency threshold, and to use the set of all user objects with the target relationship as nodes of the social network graph, and to construct the user social network graph using the communication relationships of the user objects with the target relationship as edges of the social network graph.

[0018] The user clustering module is used to obtain user login information and user transaction information, and to classify risk groups based on the user login information and user transaction information to obtain high-risk groups;

[0019] The group identification module is used to associate and mark the high-risk groups in the user's social network graph to obtain potential abnormal groups;

[0020] The target labeling module is used to compare and analyze the login status information and user account information of the potential abnormal groups, and determine the target users to be identified based on the analysis results.

[0021] An abnormal data identification module is used to obtain the number of times the target user to be identified has logged out, and when the number of logins exceeds a preset login count threshold, the target user to be identified is marked as target abnormal data.

[0022] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below:

[0023] A computer device includes a memory and a processor, the memory storing computer-readable instructions, wherein the processor, when executing the computer-readable instructions, implements the steps of the abnormal data identification method as described in any of the preceding claims.

[0024] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below:

[0025] A computer-readable storage medium storing computer-readable instructions that, when executed by a processor, implement the steps of the abnormal data identification method as described in any of the preceding claims.

[0026] Compared with the prior art, the embodiments of this application have the following advantages: This embodiment extracts user communication information within a preset time range from the database and performs frequency statistics on the user communication information to obtain the user-to-user communication frequency; it determines whether the user-to-user communication frequency is greater than a preset communication frequency threshold; if the user-to-user communication frequency is greater than the preset communication frequency threshold, the user object corresponding to the user-to-user communication frequency is identified as having a target relationship, and the set of all user objects having the target relationship is used as nodes of the social network graph, and the communication relationship of the user objects having the target relationship is used as the edge of the social network graph to construct the user social network graph; it obtains user login information and user transaction information, and classifies risk groups according to the user login information and user transaction information to obtain high-risk groups; it performs association marking on the user social network graph according to the high-risk groups to obtain potential abnormal groups; it compares and analyzes the login status information and user account information of the potential abnormal groups, and determines the target users to be identified according to the analysis results; it obtains the number of times the target users to be identified have logged out, and when the number of times they have logged out is greater than a preset number of times they have logged out, it marks the target users to be identified as target abnormal data. This effectively enables accurate identification of fraudulent target data among users and improves the reliability of the identification. Attached Figure Description

[0027] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;

[0029] Figure 2 A flowchart of an embodiment of the abnormal data identification method according to this application;

[0030] Figure 3 This is a schematic diagram of the structure of one embodiment of the abnormal data identification device according to this application;

[0031] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0033] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a non-related or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0035] like Figure 1As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0036] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0037] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0038] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0039] It should be noted that the abnormal data identification method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the abnormal data identification device is generally set in the server / terminal device.

[0040] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0041] Continue to refer to Figure 2 A flowchart of an embodiment of the abnormal data identification method according to this application is shown. The abnormal data identification method includes the following steps:

[0042] Step S10: Extract user communication information within a preset time range from the database, and perform frequency statistics on the user communication information to obtain the user's communication frequency.

[0043] In this embodiment, user communication information refers to the information data generated during user interactions such as text and voice. User communication information includes user communication frequency and user communication partners. User communication frequency refers to the number of times a user interacts with other users within a certain period, and user communication partners refer to other users involved in the communication process. For example, in financial applications, user communication information originates from consultations between users and financial experts or investment advisors, or from conversations between users and other users regarding investment, financial management, and other business matters. In healthcare applications, user communication information originates from consultations between users and doctors or health experts, or from conversations between users and other users regarding medical diagnosis, medical insurance, and other related matters. The user social network graph is graph data constructed based on the user social relationships contained in the user communication information, used to display the social scope and social attributes during the user communication process. User communication information can be extracted from the database using communication information extraction identifiers, and then filtered using a preset time range as a filtering condition to obtain user communication information within the preset time range. User-to-user communication frequency refers to the frequency with which a user communicates with other users in pairs. It is calculated by counting the number of times each user communicates with other users within a specified time frame to obtain the frequency of user communication information. In this embodiment, the preset time frame can be three months prior to the current time, and can be adjusted accordingly based on actual circumstances.

[0044] Step S20: Determine whether the user's communication frequency is greater than a preset communication frequency threshold;

[0045] In this embodiment, the preset communication frequency threshold is a pre-set standard used to determine whether communication between two users is frequent enough to establish a specific relationship (i.e., a target relationship). In this embodiment, the preset communication frequency threshold can be preset to 5, and can be adjusted accordingly based on actual conditions.

[0046] Step S30: If the communication frequency of the user pair is greater than the preset communication frequency threshold, then the user object corresponding to the communication frequency of the user pair is determined to have a target relationship, and the set of all user objects with the target relationship is used as the node of the social network graph. The communication relationship of the user objects with the target relationship is used as the edge of the social network graph to construct the user social network graph.

[0047] In this embodiment, user objects can be identified as having a target relationship by adding tag pairs. For example, if the communication frequency between two users exceeds a preset communication frequency threshold, they can be assigned a tag pair indicating "having a target relationship." All user objects with target relationships are treated as nodes in the social network graph, and the communication relationships between these user objects are used as edges in the social network graph. For example, if two users have a target relationship, they will be connected by an edge. This effectively constructs the user social network graph based on nodes and edges.

[0048] Step S40: If the communication frequency between the users is less than or equal to the preset communication frequency threshold, then the user object corresponding to the communication frequency between the users is determined to have no target relationship.

[0049] In this embodiment, user objects can be identified as having no target relationship by adding tag pairs to user objects. For example, if the communication frequency between two users does not reach a preset communication frequency threshold, they can be assigned a tag pair indicating "no target relationship" for identification.

[0050] Step S50: Obtain user login information and user transaction information, and classify risk groups based on the user login information and user transaction information to obtain high-risk groups;

[0051] In this embodiment, both user login information and user transaction information originate from a database. For example, in financial applications, user login information refers to the login information of a user accessing a financial trading platform, while user transaction information refers to the information of transactions such as transfers, loans, and payments made by the user on the financial trading platform. In healthcare applications, user login information refers to the login information of a patient accessing a medical institution's system, while user transaction information refers to the information of transactions such as medical expense settlement, medical insurance enrollment, and medical insurance reimbursement made by the user on the medical institution's system. Risk group segmentation includes user clustering and abnormal behavior judgment. User clustering effectively clusters the user objects corresponding to user login information and user transaction information, and abnormal behavior judgment identifies the risk of the user clusters generated after clustering, thereby effectively identifying high-risk groups.

[0052] Step S60: Based on the high-risk group, perform association marking in the user's social network graph to obtain potential abnormal groups;

[0053] In this embodiment, nodes and edges associated with users in high-risk groups are found in the user's social network graph, and then association labeling is performed based on the relationship between nodes and edges and the characteristics of users' historical behavior, so as to identify a larger range of potential abnormal groups in the high-risk groups.

[0054] Step S70: Compare and analyze the login status information and user account information of the potential abnormal group, and determine the target users to be identified based on the analysis results;

[0055] In this embodiment, user login information includes login status information, and user account information includes username information and real name information. By comparing and analyzing the login status information and user account information of potential abnormal groups, it is possible to determine whether the user has suspicious login behavior and account interaction behavior, so as to effectively identify the target user to be identified from the potential abnormal group.

[0056] Step S80: Obtain the number of times the target user to be identified has logged out, and when the number of times logged out exceeds a preset threshold, mark the target user to be identified as target abnormal data.

[0057] In this embodiment, "target anomalous data" refers to target users identified as having fraudulent intentions. "Account cancellation count" refers to the number of times a user's account has been cancelled. By detecting and judging the number of account cancellations by the target user to be identified, the reliability and accuracy of the identification can be further improved.

[0058] This embodiment extracts user communication information within a preset time range from a database and performs frequency statistics on the user communication information to obtain the user-to-user communication frequency. It then determines whether the user-to-user communication frequency is greater than a preset communication frequency threshold. If the user-to-user communication frequency is greater than the preset communication frequency threshold, the user object corresponding to the user-to-user communication frequency is identified as having a target relationship. All user objects with the target relationship are used as nodes in a social network graph, and the communication relationships of these user objects are used as edges to construct the user social network graph. User login information and user transaction information are obtained, and risk groups are segmented based on these information to obtain high-risk groups. These high-risk groups are then associated and marked in the user social network graph to obtain potential abnormal groups. The login status information and user account information of these potential abnormal groups are compared and analyzed, and target users to be identified are determined based on the analysis results. The number of times a target user cancels their account is obtained, and when the number of cancellations exceeds a preset cancellation threshold, the target user is marked as target abnormal data. This effectively achieves accurate identification of target abnormal data with fraud risk among users and improves the reliability of the identification.

[0059] This method can be applied to anomaly data identification in financial business systems. These systems include wealth management systems, lending systems, and consumer business systems. The database refers to the database of the aforementioned financial business systems. User communication information refers to communication information between logged-in users within these systems. User login information and user transaction information refer to user login events and user transaction events occurring within these systems. The target user to be identified refers to a confirmed suspicious user. All of the aforementioned user communication information, user login information, and user transaction information are recorded in and extracted from the database of the aforementioned financial business systems.

[0060] In this embodiment, communication information between users, between users and financial experts, and between users and investment advisors within the wealth management system is acquired to determine whether a target relationship exists between users. If a target relationship exists, a user social network graph related to that user is constructed. Login information and transaction information of users within the wealth management system are acquired to assess whether their current login and transaction behaviors are high-risk, identifying high-risk groups. These high-risk groups are then linked to the user social network graph to identify a wider range of potential abnormal groups. Further analysis of the potential abnormal groups' historical login status (login endpoint, location, time, etc.) within the wealth management system, along with usernames and whether they are verified, is used to identify suspicious target users. Finally, the number of times a target user logs out of the wealth management system is used to determine if they belong to the target abnormal data category. If the number of logouts exceeds a preset threshold, the target user is identified as fraudulent target abnormal data. This will provide reliable and effective data support for subsequent anti-fraud efforts.

[0061] In some optional implementations of this embodiment, obtaining user login information and user transaction information, and classifying risk groups based on the user login information and user transaction information to obtain high-risk groups includes the following steps:

[0062] Obtain user login information and user transaction information, and perform user clustering based on the user login information and user transaction information to obtain user cluster groups;

[0063] User login information includes login frequency, and user transaction information includes transaction amount and transaction frequency. A user feature vector matrix is ​​constructed using this user login information and user transaction information. Clustering is then performed on the user feature vector matrix to obtain user clusters based on user characteristics.

[0064] Abnormal behavior is identified from the user clusters to determine high-risk groups.

[0065] Abnormal behaviors include unusual login frequency, unusual transaction amounts, unusual transaction times, and unusual browsing behavior. By identifying abnormal behaviors within user clusters and calculating the proportion of such behaviors, it is possible to effectively determine whether a user cluster is a high-risk group.

[0066] This embodiment obtains user login information and user transaction information, and then performs user clustering based on these information to obtain user cluster groups. Abnormal behavior is then assessed within these user cluster groups to identify high-risk groups. This effectively achieves the segmentation of high-risk groups based on user clustering and abnormal behavior, facilitating the subsequent identification of potential abnormal groups.

[0067] In some optional implementations of this embodiment, obtaining user login information and user transaction information, and performing user clustering based on the user login information and user transaction information to obtain user cluster groups includes the following steps:

[0068] Construct a user feature vector matrix based on the user login information and the user transaction information;

[0069] In this embodiment, for each user, the specific values ​​of their login frequency, transaction amount, and transaction frequency are extracted. These three feature values ​​for each user are combined into a feature vector. For example, for user A, their feature vector might be [login frequency A, transaction amount A, transaction frequency A]. All user feature vectors are arranged in rows to form a two-dimensional matrix. Each row represents a user, and each column represents a feature. After the above steps, the user feature vector matrix is ​​obtained. If the dimensions of the features differ significantly, the feature vector matrix is ​​standardized to eliminate the impact of these differences on the clustering results. This standardization can be based on the mean and standard deviation of each feature.

[0070] The user feature vector matrix is ​​subjected to data standardization processing to obtain a standard user feature vector matrix;

[0071] In this embodiment, the numerical features in the user feature vector matrix are standardized to eliminate the dimensional differences between different features. In this embodiment, Z-score standardization is used. Z-score standardization converts the feature values ​​into a distribution with a mean of 0 and a standard deviation of 1, thereby obtaining a standard user feature vector matrix.

[0072] The standard user feature vector matrix is ​​clustered according to a preset clustering algorithm to obtain the user cluster group.

[0073] In this embodiment, the preset clustering algorithm can be the K-means clustering algorithm. The steps for clustering calculation based on K-means clustering include: determining the number of clusters K, where K represents the number of groups into which users are to be divided; randomly selecting K users from the standard user feature vector matrix as initial cluster centers; calculating the distance from each user to the cluster center using metrics such as Euclidean distance and Manhattan distance; assigning users to the nearest cluster center; assigning each user to the group to which the nearest cluster center belongs based on the distance metric results; recalculating the cluster center of each group, i.e., calculating the mean of the feature vectors of all users within the group; and checking whether the cluster centers have changed or whether the preset number of iterations has been reached. If the convergence condition is met, the iteration stops; otherwise, the iteration continues. After completing the above clustering calculation steps, the user cluster groups are obtained.

[0074] This embodiment constructs a user feature vector matrix based on the user login information and the user transaction information; it then performs data standardization on the user feature vector matrix to obtain a standard user feature vector matrix; finally, it performs clustering calculations on the standard user feature vector matrix according to a preset clustering algorithm to obtain the user cluster groups. This effectively achieves user clustering based on user characteristics, facilitating subsequent abnormal behavior detection and processing.

[0075] In some optional implementations of this embodiment, the step of judging abnormal behavior in the user cluster to obtain a high-risk group includes the following steps:

[0076] Obtain user behavior information of the user cluster;

[0077] In this embodiment, user behavior information related to each user in a user cluster can be extracted from the database based on the user identifier corresponding to that cluster. This user behavior information includes login behavior, transaction behavior, browsing behavior, etc.

[0078] The abnormal behavior information of the user is identified by performing abnormal behavior identification on the abnormal behavior rules to obtain abnormal user behavior information.

[0079] In this embodiment, abnormal behavior rules refer to rule information that defines abnormal behaviors. These rules may include rules for abnormal login frequency, abnormal transaction amount, abnormal transaction time, and abnormal browsing behavior. User behavior information is matched against these defined rules to identify user behavior information that conforms to them as abnormal user behavior information.

[0080] The proportion of abnormal behavior in a user group is calculated based on the user abnormal behavior information and the total number of users in the user cluster.

[0081] In this embodiment, for each user cluster, the number of users exhibiting abnormal behavior is counted, i.e., the number of users with abnormal behavior. Then, based on the total number of users in the user cluster and the number of users with abnormal behavior, the abnormal behavior ratio of the cluster is calculated. The formula for calculating the abnormal behavior ratio of the cluster is: Abnormal behavior ratio of the cluster = Number of users with abnormal behavior / Total number of users.

[0082] Determine whether the proportion of abnormal behavior in the group is greater than a preset behavior proportion threshold;

[0083] In this embodiment, a preset behavior ratio threshold is used to determine whether the proportion of abnormal behavior in a group has reached a high-risk level. In this embodiment, the preset behavior ratio threshold is set to 60%, but it can be set and adjusted according to actual conditions.

[0084] If the proportion of abnormal behavior in the group is greater than the preset behavior proportion threshold, then the user cluster corresponding to the proportion of abnormal behavior in the group is identified as the high-risk group.

[0085] In this embodiment, a user cluster can be identified as a high-risk group by adding a label corresponding to a high-risk group. For example, adding the label "high-risk group" to a user cluster indicates that the user cluster is a high-risk group.

[0086] If the proportion of abnormal behavior in the group is less than or equal to the preset behavior proportion threshold, then the user cluster corresponding to the proportion of abnormal behavior in the group is determined as a non-high-risk group.

[0087] In this embodiment, user clusters can be identified as non-high-risk groups by adding tags corresponding to non-high-risk groups. For example, adding the tag "non-high-risk group" to a user cluster indicates that the user cluster is a non-high-risk group.

[0088] This embodiment obtains user behavior information from the user cluster; identifies abnormal behavior based on abnormal behavior rules to obtain abnormal user behavior information; calculates the abnormal behavior ratio of the cluster based on the abnormal behavior information and the total number of users in the cluster; determines whether the abnormal behavior ratio is greater than a preset behavior ratio threshold; if the abnormal behavior ratio is greater than the preset behavior ratio threshold, the user cluster corresponding to the abnormal behavior ratio is identified as a high-risk group; if the abnormal behavior ratio is less than or equal to the preset behavior ratio threshold, the user cluster corresponding to the abnormal behavior ratio is identified as a non-high-risk group. This effectively enables the determination of whether a user cluster belongs to a high-risk group based on its abnormal behavior, providing a reliable basis for subsequent identification of potential abnormal groups.

[0089] In some optional implementations of this embodiment, the step of identifying potential anomalous groups by associating them with the high-risk group in the user's social network graph includes the following steps:

[0090] Based on the high-risk groups, corresponding related subgraphs are extracted from the user's social network graph;

[0091] In this embodiment, all nodes and edges associated with users in high-risk groups are found in the user's social network graph to form an association subgraph, which includes high-risk group users and their direct and indirect associated users in the network.

[0092] Determine whether there are edge connections between users in the high-risk group based on the edge set of the associated subgraph;

[0093] In this embodiment, by traversing the edge set of the associated subgraph, it is determined whether there are edge connections between users in the high-risk group. If there are edge connections, it means that these users have a direct connection on the social network.

[0094] If there are edge connections between users in the high-risk group, then the users in the high-risk group with edge connections are marked as potential users, thus obtaining a set of potential users;

[0095] In this embodiment, users within a high-risk group who are connected by edges are labeled with tags corresponding to potential users for effective marking. For example, a tag "identified as a potential user" is added between users with connected edges. Once marking is complete, the marked potential users are recorded in a dataset to form a potential user set.

[0096] If there are no edge connections between users in the high-risk group, the process of determining whether there are edge connections between users in the high-risk group will continue until all users in the high-risk group have been determined.

[0097] In this embodiment, the judgment can only be completed after traversing all users in the entire high-risk group. Once the judgment of the existence of edge connections for all users in the high-risk group has been completed, the current judgment step is stopped and the result is output.

[0098] Calculate the connection degree between each user and other users in the potential user set to obtain the user connection centrality.

[0099] In this embodiment, an adjacency matrix or adjacency list can be used to represent the connections between users in the potential user set. The adjacency matrix is ​​a two-dimensional array where rows and columns represent users, and the values ​​of the array elements represent the connection status between users (1 for connected, 0 for not connected). The adjacency list is a dictionary or hash table where the key is the user ID and the value is a list of other user IDs directly connected to that user. By iterating through each user in the potential user set and counting the number of direct connections between each user and other users, this can be achieved by iterating through the rows of the adjacency matrix or the values ​​of the adjacency list. The connection degree of each user is recorded in a data structure, such as a list or dictionary, with the user ID as the key and the connection degree as the value. This effectively yields the user connection centrality of each user in the potential user set, which reflects the user's central position and influence in the social network.

[0100] Extract historical behavior information corresponding to the potential user set, and perform feature extraction on the historical behavior information to obtain user behavior feature vectors;

[0101] In this embodiment, historical behavior information corresponding to a set of potential users can be extracted from the database using historical behavior information extraction identifiers. Then, feature extraction is performed on the historical behavior information to obtain a user behavior feature vector, which may include login frequency, transaction amount, browsing behavior, social interaction, etc.

[0102] The user risk assessment value is obtained by fusing the user connectivity centrality and the user behavior feature vector using a graph neural network algorithm.

[0103] In this embodiment, the graph neural network algorithm can employ a GCN (Graph Convolutional Network). A GCN learns node embeddings by propagating information through nodes (users) and edges (connections between users) in a graph structure. By using user behavior feature vectors and connection centrality as input to the GCN model, the model applies one or more graph convolutional layers to aggregate information from neighboring nodes. These layers update the embedding of each node to include information from its neighbors. A fully connected layer (or other types of layers, such as a softmax layer, depending on the task) is then applied to produce the final output, the user risk assessment value. The training steps of the GCN model include: acquiring users with known risk levels as training data and labeling them (high risk, medium risk, low risk, etc.); selecting an appropriate loss function (such as cross-entropy loss) to measure the difference between the model's predicted risk level and the true label; and selecting an optimizer (such as Adam) to update the model's weights to minimize the loss function, resulting in a GCN model capable of effectively generating user risk assessment values.

[0104] The potential user set is clustered based on the user risk assessment value to obtain the potential abnormal group.

[0105] In this embodiment, the potential user set can be clustered according to the user risk assessment value using the K-means clustering algorithm to obtain potential abnormal groups with different risk levels.

[0106] This embodiment extracts corresponding association subgraphs from the user social network graph based on the high-risk group; determines whether there are edge connections between users in the high-risk group based on the edge set of the association subgraph; if there are edge connections between users in the high-risk group, then the users with edge connections in the high-risk group are marked as potential users, obtaining a potential user set; calculates the connectivity degree between each user in the potential user set and other users, obtaining the user connectivity centrality; extracts historical behavior information corresponding to the potential user set, and performs feature extraction on the historical behavior information to obtain a user behavior feature vector; fuses the user connectivity centrality and the user behavior feature vector according to a graph neural network algorithm to obtain a user risk assessment value; and clusters the potential user set according to the user risk assessment value to obtain the potential abnormal group. This effectively achieves clustering based on user relationships and historical behavior characteristics in the high-risk group, thereby improving the accuracy of identifying potential abnormal groups.

[0107] In some optional implementations of this embodiment, the step of comparing and analyzing the login status information and user account information of the potential abnormal group, and determining the target user to be identified based on the analysis results, includes the following steps:

[0108] Obtain the login status information corresponding to the potential abnormal group;

[0109] In this embodiment, login status information can be extracted from the database based on the identifier extracted from the login information. This login status information includes the login endpoint, login time, login location, login status, and login records.

[0110] Based on the login status information, identify users with the same device ID;

[0111] In this embodiment, by comparing the login terminal information in the login status information, if there are identical login terminal information, users with identical login terminal information are initially determined to be users from the same device source; the registration information of the users initially determined to be from the same device source is obtained, the registration information including registration terminal information and registrant information; by comparing the registration terminal information and the registrant information, if the registration terminal information and the registrant information are identical, users with identical registration terminal information and the registrant information are determined to be users with the same device number.

[0112] Identify users with the same device information based on the account information of users with the same device number;

[0113] In this embodiment, account information of users with the same device ID is obtained, wherein the account information includes username information and real-name information; by comparing the username information and the real-name information, if the username information and the real-name information are the same, then users with the same username information and the real-name information are determined to be users of the same sub-device; payment information of users determined to be users of the same sub-device is obtained, wherein the payment information includes payment number information and binding number information; by comparing the payment number information and the binding number information, if the payment number information and the binding number information are the same, then users with the same payment number information and the binding number information are determined to be users with the same IP address; device ID information of users determined to be users with the same IP address is obtained; by comparing the device ID information, if there is identical device ID information, then users with identical device ID information are determined to be users with the same device information.

[0114] The user account information of the users with the same device information is obtained, and the potential abnormal group is marked according to the user account information to obtain the target user to be identified.

[0115] In this embodiment, user account information of users identified as having the same device information is obtained; by comparing the user account information, if there is identical user account information, the users with identical user account information are identified as having the same associated account; device information of users identified as having the same associated account is obtained; by comparing the device information, if the device information is the same, the users with the same device information and the users with the same associated account are identified as target users to be identified.

[0116] This embodiment obtains the login status information corresponding to the potential abnormal group; identifies users with the same device ID based on the login status information; identifies users with the same device information based on the account information of the users with the same device ID; obtains the user account information of the users with the same device information; and marks the potential abnormal group based on the user account information to obtain the target users to be identified. This effectively achieves the identification of suspicious target users based on user login status and associated accounts, facilitating the subsequent identification of potential abnormal groups.

[0117] In some optional implementations of this embodiment, the step of obtaining the number of logouts of the target user to be identified, and marking the team to which the target user belongs as a potentially abnormal team when the number of logouts is greater than or equal to a preset logout threshold, includes the following steps:

[0118] Obtain the account registration time of the target user to be identified, and count the number of account cancellations based on the account registration time;

[0119] In this embodiment, registration information identifiers can be used to retrieve the account registration time and account cancellation records of target users to be identified from the database. For each target user to be identified, the number of cancellations is calculated based on their account cancellation records. In this embodiment, when calculating the number of cancellations, a time range (such as the past three months, six months, etc.) can be set to consider the cancellation behavior of users within that time range, resulting in a more effective number of cancellations.

[0120] Determine whether the number of cancellations is greater than or equal to a preset cancellation number threshold;

[0121] In this embodiment, the preset cancellation count threshold is a pre-defined threshold used to determine whether the cancellation count is abnormal. This preset cancellation count threshold can be determined based on historical data, industry standards, and business logic. In this embodiment, the preset cancellation count threshold is set to 3 times, but it can be set and adjusted according to actual circumstances.

[0122] If the number of cancellations is greater than or equal to the preset cancellation threshold, then the target user to be identified is determined to be a risk user.

[0123] In this embodiment, a target user can be identified as a risky user by adding a tag corresponding to a risky user to the target user ID. For example, the target user ID can be labeled with the tag "risky user" to facilitate subsequent identification.

[0124] If the number of cancellations is less than the preset cancellation threshold, the target user to be identified is determined to be a non-risk user.

[0125] In this embodiment, the target user to be identified can be identified as a non-risk user by adding a tag corresponding to non-risk users to the target user ID. For example, the target user ID can be labeled as a "non-risk user" to facilitate subsequent identification.

[0126] Obtain the associated transaction information and transaction amount of the risky user, and determine the frequent transaction accounts and high-risk accounts based on the associated transaction information and transaction amount;

[0127] In this embodiment, a graph algorithm is used to obtain the transaction frequency of all accounts through the associated transaction information of high-risk users. If the transaction frequency is greater than a preset transaction frequency threshold, it is identified as a frequent transaction account. Based on the transaction amount of the frequent transaction account, the cumulative transaction amount of the frequent transaction account is obtained. If the cumulative transaction amount is greater than a preset transaction amount threshold, the frequent transaction account is identified as a high-risk account.

[0128] The target abnormal data is obtained by confirming abnormal data based on the frequently traded accounts and the high-risk accounts.

[0129] In this embodiment, a hierarchical clustering algorithm is used to obtain the associated accounts of risky users based on the frequently traded accounts and the high-risk accounts. If there is a high-risk account among the associated accounts, then the account is determined to be the target abnormal data.

[0130] This embodiment obtains the account registration time of the target user to be identified and counts the number of account cancellations based on the registration time; it determines whether the number of cancellations is greater than or equal to a preset cancellation threshold; if the number of cancellations is greater than or equal to the preset cancellation threshold, the target user to be identified is identified as a risk user; it obtains the associated transaction information and transaction amount of the risk user, and determines frequent transaction accounts and high-risk accounts based on the associated transaction information and transaction amount; it confirms abnormal data based on the frequent transaction accounts and the high-risk accounts to obtain the target abnormal data. This effectively achieves accurate identification of potential abnormal teams among users from multiple dimensions, facilitating subsequent investigation.

[0131] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0132] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0133] Further reference Figure 3 As a response to the above Figure 1 To implement the method shown, this application provides an embodiment of an abnormal data identification device, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0134] like Figure 3 As shown, the abnormal data identification device 900 described in this embodiment includes: a frequency acquisition module 901, a threshold judgment module 902, a graph construction module 903, a relationship determination module 904, a user clustering module 905, a group identification module 906, a target labeling module 907, and an abnormal data identification module 908. Wherein:

[0135] The frequency acquisition module 901 is used to extract user communication information within a preset time range from the database, and to perform frequency statistics on the user communication information to obtain the user's communication frequency.

[0136] The threshold judgment module 902 is used to determine whether the user's communication frequency is greater than a preset communication frequency threshold.

[0137] The graph construction module 903 is used to determine the user object corresponding to the user pair communication frequency as having a target relationship if the user pair communication frequency is greater than the preset communication frequency threshold, and to use the set of all user objects with the target relationship as nodes of the social network graph, and to construct the user social network graph using the communication relationship of the user objects with the target relationship as the edge of the social network graph.

[0138] The relationship determination module 904 is used to determine the user object corresponding to the user communication frequency as having no target relationship if the user communication frequency is less than or equal to the preset communication frequency threshold.

[0139] User clustering module 905 is used to obtain user login information and user transaction information, and to classify risk groups based on the user login information and user transaction information to obtain high-risk groups;

[0140] The group identification module 906 is used to perform association marking on the user's social network graph based on the high-risk group to obtain potential abnormal groups;

[0141] The target marking module 907 is used to compare and analyze the login status information and user account information of the potential abnormal group, and determine the target users to be identified based on the analysis results;

[0142] The abnormal data identification module 908 is used to obtain the number of times the target user to be identified has logged out, and when the number of logged out is greater than a preset number of logged out threshold, the target user to be identified is marked as target abnormal data.

[0143] This embodiment, by employing the aforementioned abnormal data identification device, can extract user communication information within a preset time range from the database, and perform frequency statistics on the user communication information to obtain the user-to-user communication frequency; determine whether the user-to-user communication frequency is greater than a preset communication frequency threshold; if the user-to-user communication frequency is greater than the preset communication frequency threshold, then the user object corresponding to the user-to-user communication frequency is identified as having a target relationship, and the set of all user objects having the target relationship is used as nodes in the social network graph, and the communication relationships of user objects having the target relationship are used as edges in the social network graph to construct the user social network graph; obtain user login information and user transaction information, and classify risk groups based on the user login information and user transaction information to obtain high-risk groups; perform association marking on the user social network graph based on the high-risk groups to obtain potential abnormal groups; compare and analyze the login status information and user account information of the potential abnormal groups, and determine the target users to be identified based on the analysis results; obtain the number of times the target users to be identified have logged out, and when the number of times logged out is greater than a preset number of times logged out, mark the target users to be identified as target abnormal data. This effectively enables accurate identification of fraudulent target data among users and improves the reliability of the identification.

[0144] In some optional implementations of this embodiment, the user clustering module 905 includes: a clustering partitioning unit and a behavior judgment unit. Wherein:

[0145] The clustering division unit is used to obtain user login information and user transaction information, and to perform user clustering based on the user login information and user transaction information to obtain user cluster groups;

[0146] The behavior judgment unit is used to judge abnormal behavior of the user cluster and obtain high-risk groups.

[0147] This embodiment effectively achieves the classification of high-risk groups based on user clustering and abnormal behavior by setting up a user clustering module 905 that includes a clustering division unit and a behavior judgment unit, so as to facilitate the subsequent identification of potential abnormal groups.

[0148] In some optional implementations of this embodiment, the clustering partitioning unit includes: a matrix construction subunit, a matrix normalization subunit, and a population clustering subunit. Wherein:

[0149] The matrix construction subunit is used to construct a user feature vector matrix based on the user login information and the user transaction information.

[0150] The matrix standardization subunit is used to perform data standardization processing on the user feature vector matrix to obtain a standard user feature vector matrix.

[0151] The group clustering subunit is used to perform clustering calculations on the standard user feature vector matrix according to a preset clustering algorithm to obtain the user cluster group.

[0152] This embodiment effectively achieves user clustering based on user characteristics by setting up clustering partitioning units including matrix construction subunits, matrix standardization subunits, and group clustering subunits, so as to facilitate subsequent abnormal behavior judgment and processing.

[0153] In some optional implementations of this embodiment, the behavior judgment unit includes: a behavior information acquisition subunit, an abnormal behavior identification subunit, a behavior ratio calculation subunit, a ratio judgment subunit, a first group determination subunit, and a second group determination subunit. Wherein:

[0154] The behavior information acquisition subunit is used to acquire user behavior information of the user cluster group;

[0155] The abnormal behavior identification subunit is used to identify abnormal behavior in the user behavior information according to abnormal behavior rules, and obtain abnormal user behavior information.

[0156] The behavior ratio calculation subunit is used to calculate the group abnormal behavior ratio based on the user abnormal behavior information and the total number of users in the user cluster group.

[0157] The ratio judgment subunit is used to determine whether the ratio of abnormal behavior of the group is greater than a preset behavior ratio threshold.

[0158] The first group determination subunit is used to determine the user cluster corresponding to the abnormal behavior ratio of the group as the high-risk group if the abnormal behavior ratio of the group is greater than the preset behavior ratio threshold.

[0159] The second group determination subunit is used to determine the user cluster corresponding to the abnormal behavior ratio of the group as a non-high-risk group if the abnormal behavior ratio of the group is less than or equal to the preset behavior ratio threshold.

[0160] This embodiment effectively achieves the determination of whether a user belongs to a high-risk group based on the abnormal behavior of the user cluster group by setting up a behavior judgment unit including a behavior information acquisition subunit, an abnormal behavior recognition subunit, a behavior ratio calculation subunit, a ratio judgment subunit, a first group determination subunit, and a second group determination subunit, so as to provide a reliable basis for subsequent identification of potential abnormal groups.

[0161] In some optional implementations of this embodiment, the group identification module 906 includes: a subgraph extraction unit, an edge connection judgment unit, a user marking unit, a continuation judgment unit, a connectivity calculation unit, a user feature extraction unit, a risk degree fusion unit, and a group segmentation unit. Wherein:

[0162] The subgraph extraction unit is used to extract corresponding related subgraphs from the user's social network graph based on the high-risk group.

[0163] The edge connection determination unit is used to determine whether there are edge connections between users in the high-risk group based on the edge set of the associated subgraph.

[0164] The user marking unit is used to mark users in the high-risk group who have edge connections as potential users if there are edge connections between users in the high-risk group, thereby obtaining a set of potential users.

[0165] The judgment continuation unit is used to continue judging whether there is an edge connection between users in the high-risk group if there is no edge connection between them, until all users in the high-risk group have been judged.

[0166] The connectivity calculation unit is used to calculate the connectivity between each user and other users in the potential user set to obtain the user connectivity centrality.

[0167] The user feature extraction unit is used to extract historical behavior information corresponding to the potential user set, and to extract features from the historical behavior information to obtain a user behavior feature vector.

[0168] The risk fusion unit is used to fuse the user connectivity centrality and the user behavior feature vector according to the graph neural network algorithm to obtain the user risk assessment value.

[0169] The group segmentation unit is used to cluster the potential user set according to the user risk assessment value to obtain the potential abnormal group.

[0170] This embodiment establishes a group identification module 906, which includes a subgraph extraction unit, an edge connection judgment unit, a user marking unit, a continuation judgment unit, a connectivity calculation unit, a user feature extraction unit, a risk degree fusion unit, and a group segmentation unit. This effectively enables clustering based on user relationships and historical behavioral characteristics within high-risk groups, thereby improving the accuracy of identifying potential abnormal groups.

[0171] In some optional implementations of this embodiment, the target marking module 907 includes: a login information acquisition unit, a device user judgment unit, a same device judgment unit, and a target user marking unit. Wherein:

[0172] The login information acquisition unit is used to acquire the login status information corresponding to the potential abnormal group;

[0173] The device user determination unit is used to determine the same device number user based on the login status information;

[0174] The same device determination unit is used to determine users with the same device information based on the account information of the users with the same device number.

[0175] The target user marking unit is used to obtain the user account information of the users with the same device information, and to mark the potential abnormal group according to the user account information to obtain the target user to be identified.

[0176] This embodiment effectively identifies suspicious target users based on their login status and associated accounts by setting up a target marking module 907, which includes a login information acquisition unit, a device user judgment unit, a same device judgment unit, and a target user marking unit, so as to facilitate the subsequent identification of potential abnormal teams.

[0177] In some optional implementations of this embodiment, the abnormal data identification module 908 includes: a cancellation count acquisition unit, a count judgment unit, a first user determination unit, a second user determination unit, an account determination unit, and an abnormal data determination unit. Wherein:

[0178] The cancellation count acquisition unit is used to acquire the account registration time of the target user to be identified, and to count the cancellation count based on the account registration time;

[0179] The count determination unit is used to determine whether the number of cancellations is greater than or equal to a preset cancellation count threshold;

[0180] The first user determination unit is used to determine the target user to be identified as a risk user if the number of cancellations is greater than or equal to the preset number of cancellations threshold.

[0181] The second user determination unit is used to determine the target user to be identified as a non-risk user if the number of cancellations is less than the preset number of cancellations threshold.

[0182] The account determination unit is used to obtain the related transaction information and transaction amount of the risk user, and determine the frequent transaction account and the high-risk account based on the related transaction information and the transaction amount;

[0183] The abnormal data determination unit is used to confirm abnormal data based on the frequent transaction account and the high-risk account to obtain the target abnormal data.

[0184] This embodiment effectively identifies potential abnormal teams among users from multiple dimensions by setting up an abnormal data identification module 908, which includes a cancellation count acquisition unit, a count judgment unit, a first user determination unit, a second user determination unit, an account determination unit, and an abnormal data determination unit, so as to facilitate subsequent investigation.

[0185] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0186] The computer device 11 includes a memory 111, a processor 112, and a network interface 113 that are interconnected via a system bus. It should be noted that only the computer device 11 with components 111-113 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0187] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0188] The memory 111 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 111 may be an internal storage unit of the computer device 11, such as the hard disk or memory of the computer device 11. In other embodiments, the memory 111 may also be an external storage device of the computer device 11, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Of course, the memory 111 may include both internal storage units and external storage devices of the computer device 11. In this embodiment, the memory 111 is typically used to store the operating system and various application software installed on the computer device 11, such as computer-readable instructions for abnormal data identification methods. In addition, the memory 111 can also be used to temporarily store various types of data that have been output or will be output.

[0189] In some embodiments, the processor 112 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 112 is typically used to control the overall operation of the computer device 11. In this embodiment, the processor 112 is used to execute computer-readable instructions stored in the memory 111 or to process data, such as executing computer-readable instructions for the abnormal data identification method.

[0190] The network interface 113 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 11 and other electronic devices.

[0191] This embodiment, by employing the aforementioned computer equipment, can extract user communication information within a preset time range from the database, and perform frequency statistics on the user communication information to obtain user-to-user communication frequency; determine whether the user-to-user communication frequency is greater than a preset communication frequency threshold; if the user-to-user communication frequency is greater than the preset communication frequency threshold, then the user object corresponding to the user-to-user communication frequency is identified as having a target relationship, and the set of all user objects having the target relationship is used as nodes in the social network graph, and the communication relationships of user objects having the target relationship are used as edges in the social network graph to construct the user social network graph; obtain user login information and user transaction information, and classify risk groups based on the user login information and user transaction information to obtain high-risk groups; perform association marking on the user social network graph based on the high-risk groups to obtain potential abnormal groups; compare and analyze the login status information and user account information of the potential abnormal groups, and determine the target users to be identified based on the analysis results; obtain the number of times the target users to be identified have logged out, and when the number of times they have logged out is greater than a preset number of times they have logged out, mark the target users to be identified as target abnormal data. This effectively enables accurate identification of fraudulent target data among users and improves the reliability of the identification.

[0192] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the abnormal data identification method described above.

[0193] This embodiment, by employing the aforementioned computer-readable storage medium, can extract user communication information within a preset time range from a database, perform frequency statistics on the user communication information to obtain user-to-user communication frequency; determine whether the user-to-user communication frequency is greater than a preset communication frequency threshold; if the user-to-user communication frequency is greater than the preset communication frequency threshold, then the user object corresponding to the user-to-user communication frequency is identified as having a target relationship, and the set of all user objects having the target relationship is used as nodes in a social network graph, with the communication relationships of user objects having the target relationship used as edges in the social network graph to construct a user social network graph; obtain user login information and user transaction information, and classify risk groups based on the user login information and user transaction information to obtain high-risk groups; associate and mark high-risk groups in the user social network graph to obtain potential abnormal groups; compare and analyze the login status information and user account information of the potential abnormal groups, and determine the target users to be identified based on the analysis results; obtain the number of times the target users to be identified have logged out, and when the number of times they have logged out is greater than a preset number of times they have logged out, mark the target users to be identified as target abnormal data. This effectively enables accurate identification of fraudulent target data among users and improves the reliability of the identification.

[0194] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0195] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

[0196] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.

Claims

1. A method for identifying abnormal data, characterized in that, Includes the following steps: Extract user communication information within a preset time range from the database, and perform frequency statistics on the user communication information to obtain the user's communication frequency; Determine whether the user's communication frequency is greater than a preset communication frequency threshold; If the frequency of communication between users is greater than the preset communication frequency threshold, then the user objects corresponding to the frequency of communication between users are determined to have a target relationship, and the set of all user objects with the target relationship is used as nodes of the social network graph. The communication relationships of user objects with the target relationship are used as edges of the social network graph to construct the user social network graph. Obtain user login information and user transaction information, and classify risk groups based on the user login information and user transaction information to obtain high-risk groups; Based on the association and labeling of the high-risk groups in the user's social network graph, potential abnormal groups are obtained; The login status information and user account information of the potential abnormal groups are compared and analyzed, and the target users to be identified are determined based on the analysis results; The number of times the target user to be identified has logged out is obtained, and when the number of logouts exceeds a preset logout threshold, the target user to be identified is marked as target abnormal data.

2. The abnormal data identification method according to claim 1, characterized in that, The steps of obtaining user login information and user transaction information, and classifying risk groups based on the user login information and user transaction information to obtain high-risk groups, specifically include: Obtain user login information and user transaction information, and perform user clustering based on the user login information and user transaction information to obtain user cluster groups; Abnormal behavior is identified from the user clusters to determine high-risk groups.

3. The abnormal data identification method according to claim 2, characterized in that, The steps of obtaining user login information and user transaction information, and performing user clustering based on the user login information and user transaction information to obtain user cluster groups, specifically include: Construct a user feature vector matrix based on the user login information and the user transaction information; The user feature vector matrix is ​​subjected to data standardization processing to obtain a standard user feature vector matrix; The standard user feature vector matrix is ​​clustered according to a preset clustering algorithm to obtain the user cluster group.

4. The abnormal data identification method according to claim 2, characterized in that, The step of judging abnormal behavior in the user cluster to obtain a high-risk group specifically includes: Obtain user behavior information of the user cluster; The abnormal behavior information of the user is identified by performing abnormal behavior identification on the abnormal behavior rules to obtain abnormal user behavior information. The proportion of abnormal behavior in a user group is calculated based on the user abnormal behavior information and the total number of users in the user cluster. Determine whether the proportion of abnormal behavior in the group is greater than a preset behavior proportion threshold; If the proportion of abnormal behavior in the group is greater than the preset behavior proportion threshold, then the user cluster corresponding to the proportion of abnormal behavior in the group is determined as the high-risk group.

5. The abnormal data identification method according to claim 1, characterized in that, The step of identifying potential anomalous groups by associating them with the high-risk groups in the user's social network graph specifically includes: Based on the high-risk groups, corresponding related subgraphs are extracted from the user's social network graph; Determine whether there are edge connections between users in the high-risk group based on the edge set of the associated subgraph; If there are edge connections between users in the high-risk group, then the users in the high-risk group with edge connections are marked as potential users, thus obtaining a set of potential users; Calculate the connection degree between each user and other users in the potential user set to obtain the user connection centrality. Extract historical behavior information corresponding to the potential user set, and perform feature extraction on the historical behavior information to obtain user behavior feature vectors; The user risk assessment value is obtained by fusing the user connectivity centrality and the user behavior feature vector using a graph neural network algorithm. The potential user set is clustered based on the user risk assessment value to obtain the potential abnormal group.

6. The abnormal data identification method according to claim 1, characterized in that, The step of comparing and analyzing the login status information and user account information of the potential abnormal groups, and determining the target users to be identified based on the analysis results, specifically includes: Obtain the login status information corresponding to the potential abnormal group; Based on the login status information, identify users with the same device ID; Identify users with the same device information based on the account information of users with the same device number; The user account information of the users with the same device information is obtained, and the potential abnormal group is marked according to the user account information to obtain the target user to be identified.

7. The abnormal data identification method according to claim 1, characterized in that, The step of obtaining the number of logouts of the target user to be identified, and marking the team to which the target user belongs as a potentially abnormal team when the number of logouts is greater than or equal to a preset logout threshold, specifically includes: Obtain the account registration time of the target user to be identified, and count the number of account cancellations based on the account registration time; Determine whether the number of cancellations is greater than or equal to a preset cancellation number threshold; If the number of cancellations is greater than or equal to the preset cancellation threshold, then the target user to be identified is determined to be a risk user. Obtain the associated transaction information and transaction amount of the risky user, and determine the frequent transaction accounts and high-risk accounts based on the associated transaction information and transaction amount; The target abnormal data is obtained by confirming abnormal data based on the frequently traded accounts and the high-risk accounts.

8. An abnormal data identification device, characterized in that, include: The frequency acquisition module is used to extract user communication information within a preset time range from the database, and to perform frequency statistics on the user communication information to obtain the user's communication frequency. The threshold determination module is used to determine whether the user's communication frequency is greater than a preset communication frequency threshold. The graph construction module is used to determine the user object corresponding to the user pair communication frequency as having a target relationship if the user pair communication frequency is greater than the preset communication frequency threshold, and to use the set of all user objects with the target relationship as nodes of the social network graph, and to construct the user social network graph using the communication relationships of the user objects with the target relationship as edges of the social network graph. The user clustering module is used to obtain user login information and user transaction information, and to classify risk groups based on the user login information and user transaction information to obtain high-risk groups; The group identification module is used to associate and mark the high-risk groups in the user's social network graph to obtain potential abnormal groups; The target labeling module is used to compare and analyze the login status information and user account information of the potential abnormal groups, and determine the target users to be identified based on the analysis results. An abnormal data identification module is used to obtain the number of times the target user to be identified has logged out, and when the number of logins exceeds a preset login count threshold, the target user to be identified is marked as target abnormal data.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the abnormal data identification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the abnormal data identification method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Gang fraud risk identification method based on knowledge graph and related equipment

    CN116308824A

  • Three-party payment platform risk account assessment method and system based on atlas model

    CN119005987A