Abnormal data identification method and device, computer equipment and storage medium

By building a user social network map and dividing high-risk groups, and combining user login and transaction information for association marking and comparative analysis, the problem of difficult to identify abnormal data in financial fraud in the prior art is solved, and the accurate identification of target abnormal data is achieved and the recognition reliability is improved.

CN120163585AActive Publication Date: 2025-06-17CHINA PING AN PROPERTY INSURANCE CO LTD

Patent Information

Application Number
CN202510272188.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-17
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately identify target anomaly data with fraud risks, especially in financial fraud behaviors. The behavioral patterns of abnormal user groups are blurred from normal user behavior, resulting in difficulty in risk assessment and abnormal signal capture.

Method used

By extracting user communication information within the preset time range from the database, performing frequency statistics and determining whether it is greater than the preset communication frequency threshold, and building a user social network map; combining user login information and transaction information, divide high-risk groups, and performing correlation marks in the social network map to determine potential abnormal groups; comparing and analyzing the login status and account information of potential abnormal groups, obtaining the number of cancellations of the target user to be identified, and if the threshold is exceeded, it is marked as target abnormal data.

Benefits of technology

It realizes accurate identification of target abnormal data with fraud risks among users, improves the reliability and efficiency of identification, and can more effectively capture abnormal signals in financial fraud behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163585A_ABST
    Figure CN120163585A_ABST
Patent Text Reader

Abstract

The invention relates to an abnormal data identification method and device, computer equipment and a storage medium. The method comprises the following steps: constructing a user social network graph according to user communication information; dividing risk groups according to the user login information and the user transaction information to obtain high-risk groups; performing association marking in the user social network graph according to the high-risk group to obtain a potential abnormal group; marking a target user to be identified for the user login information and the user account information of the potential abnormal group; and obtaining the logout times of the to-be-identified target user, and marking the to-be-identified target user as target abnormal data according to the logout times. The method and the device can be applied to financial anti-fraud application scenes, target abnormal data with fraud risks in users can be effectively and accurately identified, and the identification reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular to an abnormal data recognition method, device, computer device, and storage medium. Background Art

[0002] With the rapid development of technology, there are certain abnormal user groups in the financial field who use various technologies and means to conduct financial fraud. These abnormal user groups use encrypted communication software or specific coding languages to transmit concealed instructions and conduct illegal fund settlements. They conduct multi-layer nested transactions through shell company accounts and encrypted wallets, and their fund chains intersect with the daily financial behaviors of ordinary users in multiple ways - for example, high-frequency legal transaction scenarios such as automatic salary payment for wealth management, automatic deduction of living expenses, automatic installment repayment, and medical insurance premiums may all be implanted with fraud instructions disguised as "wealth management redemption" or "insurance premium withholding", thus resulting in financial fraud phenomena.

[0003] Traditional user behavior analysis methods and information communication content correlation mining can, to a certain extent, conduct risk assessments based on user behavior patterns and information exchange characteristics. However, since the behavior patterns of abnormal user groups are relatively blurred with the boundaries of normal business operations, it is difficult to accurately capture abnormal signals based solely on user behavior fit and communication content analysis.

[0004] Although existing user process monitoring and anomaly exclusion systems can detect abnormal behaviors through preset rules (such as account registration and cancellation frequencies, transaction amount thresholds, transaction frequency limits, etc.), abnormal user groups not only take advantage of potential loopholes in platform risk control rules but also continuously adjust their strategies to adapt to rule updates, making their activities seemingly completely conform to the normal user behavior pattern on the surface, bringing certain difficulties to the accurate identification and investigation of abnormal data in financial security and medical data. Summary of the Invention

[0005] The purpose of the embodiments of this application is to propose an abnormal data recognition method, device, computer device, and storage medium to solve the problem of being unable to accurately identify target abnormal data with fraud risks among users.

[0006] In a first aspect, the embodiments of this application provide an abnormal data recognition method, which adopts the following technical solutions:

[0007] Extract user communication information within a preset time range from a database, and perform frequency statistics on the user communication information to obtain the user communication frequency;

[0008] Determine whether the user communication frequency is greater than a preset communication frequency threshold;

[0009] If the communication frequency of the user is greater than the preset communication frequency threshold, determine the user object corresponding to the user's communication frequency as having a target relationship, and use the set of all user objects having the target relationship as the nodes of the social network graph, and use the communication relationships of the user objects having the target relationship as the edges of the social network graph to construct a user social network graph;

[0010] Obtain user login information and user transaction information, and perform risk group division according to the user login information and the user transaction information to obtain a high-risk group;

[0011] Perform association marking on the high-risk group in the user social network graph to obtain a potential abnormal group;

[0012] Perform comparative analysis on the login status information and user account information of the potential abnormal group, and determine the target user to be identified according to the analysis result;

[0013] Obtain the number of cancellations of the target user to be identified, and when the number of cancellations is greater than the preset cancellation number threshold, mark the target user to be identified as target abnormal data.

[0014] In a second aspect, an abnormal data identification device according to an embodiment of the present application also adopts the following technical solution:

[0015] A frequency acquisition module, configured to extract user communication information within a preset time range from a database, and perform frequency statistics on the user communication information to obtain the communication frequency of user pairs;

[0016] A threshold judgment module, configured to judge whether the communication frequency of user pairs is greater than a preset communication frequency threshold;

[0017] A graph construction module, configured to, if the communication frequency of user pairs is greater than the preset communication frequency threshold, determine the user object corresponding to the communication frequency of user pairs as having a target relationship, and use the set of all user objects having the target relationship as the nodes of the social network graph, and use the communication relationships of the user objects having the target relationship as the edges of the social network graph to construct a user social network graph;

[0018] A user clustering module, configured to obtain user login information and user transaction information, and perform risk group division according to the user login information and the user transaction information to obtain a high-risk group;

[0019] A group identification module, configured to perform association marking on the high-risk group in the user social network graph to obtain a potential abnormal group;

[0020] A target marking module, configured to perform comparative analysis on the login state information and user account information of the potential abnormal group, and determine the target user to be identified according to the analysis result;

[0021] An abnormal data identification module, configured to obtain the number of logouts of the target user to be identified, and mark the target user to be identified as target abnormal data when the number of logouts is greater than a preset logout number threshold.

[0022] In a third aspect, an embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0023] A computer device includes a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the abnormal data identification method described in any one of the above are implemented.

[0024] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0025] A computer-readable storage medium has computer-readable instructions stored thereon, and when the computer-readable instructions are executed by a processor, the steps of the abnormal data identification method described in any one of the above are implemented.

[0026] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: In this embodiment, by extracting user communication information within a preset time range from a database and performing frequency statistics on the user communication information, the user communication frequency is obtained; it is determined whether the user communication frequency is greater than a preset communication frequency threshold; if the user communication frequency is greater than the preset communication frequency threshold, the user object corresponding to the user communication frequency is determined to have a target relationship, and the set of all user objects having the target relationship is used as the nodes of the social network graph, and the communication relationship between the user objects having the target relationship is used as the edges of the social network graph to construct a user social network graph; user login information and user transaction information are obtained, and risk group division is performed according to the user login information and the user transaction information to obtain a high-risk group; the high-risk group is associated and marked in the user social network graph to obtain a potential abnormal group; comparative analysis is performed on the login state information and user account information of the potential abnormal group, and the target user to be identified is determined according to the analysis result; the number of logouts of the target user to be identified is obtained, and when the number of logouts is greater than a preset logout number threshold, the target user to be identified is marked as target abnormal data. Thus, it effectively realizes the accurate identification of target abnormal data with fraud risks among users and improves the reliability of identification. Description of the Drawings

[0027] To more clearly illustrate the solutions in this application, the following will briefly introduce the accompanying drawings required for the description of the embodiments of this application. Obviously, the accompanying drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0028] Figure 1 is an exemplary system architecture diagram to which this application can be applied;

[0029] Figure 2 Flowchart of an embodiment of the abnormal data recognition method according to this application;

[0030] Figure 3 is a schematic structural diagram of an embodiment of the abnormal data recognition device according to this application;

[0031] Figure 4 is a schematic structural diagram of an embodiment of the computer device according to this application. Detailed implementation manners

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the description of this application in the specification are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above accompanying drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above accompanying drawings are used to distinguish different objects and not to describe a specific order.

[0033] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears at various positions in the specification and does not necessarily refer to the same embodiment, nor is it a non-related or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0034] To enable those in the technical field of this application to better understand the solutions of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings.

[0035] Such as Figure 1As shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0036] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0037] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 may also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop portable computer, a desktop computer, etc.

[0038] The server 103 may be a server providing various services, such as a background server supporting the pages displayed on the terminal device 101.

[0039] It should be noted that the abnormal data recognition method provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the abnormal data recognition device is generally set in the server / terminal device.

[0040] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0041] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 2 Continuing to refer to

[0042] Step S10, extract the user communication information within a preset time range from the database, and perform frequency statistics on the user communication information to obtain the user communication frequency;

[0043] In this embodiment, the user communication information refers to the information data during the user's text, voice and other interactive communications. The user communication information includes the user communication frequency and the user communication object. Among them, the user communication frequency refers to the number or frequency of times the user communicates with other users within a certain period of time, and the user communication object refers to other user objects involved by the user during the communication process. For example, in the application in the financial field, the user communication information comes from the communication and consultation between the user and financial experts or investment advisors, or the dialogue communication between the user and other users regarding investment, financial management and other business aspects. In the application in the field of medical and health, the user communication information comes from the communication and consultation between the user and doctors or health experts, or the dialogue communication between the user and other users regarding medical diagnosis, medical insurance and other aspects. The user social network graph is the graph data constructed based on the user social relationships included in the user communication information, and is used to display the social scope and social attributes during the user communication process. The user communication information can be extracted from the database through the communication information extraction identifier, and then the user communication information is filtered with a preset time range as the filtering condition to obtain the user communication information within the preset time range. The user pair communication frequency is the frequency at which the user communicates with other users in pairs. By calculating the number of times each user communicates with other users within the specified time range, the frequency statistics of the user communication information is realized to obtain the user pair communication frequency. In this embodiment, the preset time range can be set to the three months before the current time point, and can be adjusted accordingly according to the actual situation.

[0044] Step S20, determine whether the user pair communication frequency is greater than the preset communication frequency threshold;

[0045] In this embodiment, the preset communication frequency threshold is a preset standard used to determine whether the communication between two users is frequent enough to constitute a certain specific relationship (i.e., the target relationship). In this embodiment, the preset communication frequency threshold can be preset to 5, and can be adjusted accordingly according to the actual situation.

[0046] Step S30, if the user pair communication frequency is greater than the preset communication frequency threshold, then determine the user object corresponding to the user pair communication frequency as having the target relationship, and use the set of all user objects having the target relationship as the nodes of the social network graph, and use the communication relationship of the user objects having the target relationship as the edges of the social network graph to construct the user social network graph;

[0047] In this embodiment, a user object can be determined to have a target relationship by adding a tag pair to the user object. For example, if the communication frequency between two users exceeds a preset communication frequency threshold, a tag pair indicating "having a target relationship" can be assigned to them for marking and determination. The set of all user objects having a target relationship is used as the nodes of the social network graph, and the communication relationships of the user objects having a target relationship are used as the edges of the social network graph. For example, if there is a target relationship between two users, there will be an edge connecting them. Thus, a user social network graph can be effectively constructed based on the nodes and edges.

[0048] Step S40, if the communication frequency of the user pair is less than or equal to the preset communication frequency threshold, the user objects corresponding to the communication frequency of the user pair are determined to have no target relationship;

[0049] In this embodiment, a user object can be determined to have no target relationship by adding a tag pair to the user object. For example, if the communication frequency between two users does not reach the preset communication frequency threshold, a tag pair indicating "having no target relationship" can be assigned to them for marking and determination.

[0050] Step S50, obtain user login information and user transaction information, and perform risk group division according to the user login information and the user transaction information to obtain a high-risk group;

[0051] In this embodiment, both the user login information and the user transaction information are sourced from a database. For example, in a financial field application, the user login information is the login information for a user to access a financial trading platform, and the user transaction information is the information on transactions such as transfers, loans, and payments made by the user on the financial trading platform. In a medical and health field application, the user login information is the login information for a patient to access a medical institution system, and the user transaction information is the information on transactions such as medical expense settlement, medical insurance enrollment, and medical insurance reimbursement made by the user on the medical institution system. The risk group division includes user clustering division and abnormal behavior judgment. Among them, through user clustering division, the user objects corresponding to the user login information and the user transaction information are effectively clustered, and then through abnormal behavior judgment, the riskiness of the user clustering groups generated after clustering is identified, so as to effectively identify the high-risk group.

[0052] Step S60, perform associated marking on the high-risk group in the user social network graph to obtain a potential abnormal group;

[0053] In this embodiment, by searching for the nodes and edges associated with the users in the high-risk group in the user social network graph, and then performing association marking processing according to the relationships between the nodes and edges and the characteristics of the user's historical behavior, a potentially abnormal group with a larger scope is identified from the high-risk group.

[0054] Step S70: Compare and analyze the login state information and user account information of the potentially abnormal group, and determine the target user to be identified according to the analysis result;

[0055] In this embodiment, the user login information includes the login state information, and the user account information includes the user name information and the real-name information. By comparing and analyzing the login state information and user account information of the potentially abnormal group, it is determined whether the user has suspicious login behaviors and account interaction behaviors, so as to effectively determine the target user to be identified from the potentially abnormal group.

[0056] Step S80: Obtain the number of cancellations of the target user to be identified, and when the number of cancellations is greater than the preset cancellation number threshold, mark the target user to be identified as target abnormal data.

[0057] In this embodiment, the target abnormal data refers to the target user determined to have fraudulent nature, and the target user with fraudulent nature can specifically refer to a member of a black production gang. The number of cancellations refers to the number of cancellation behaviors of the user account. By detecting and judging the number of cancellations of the target user to be identified, the reliability and accuracy of the identification are further improved.

[0058] In this embodiment, the user communication information within a preset time range is extracted from the database, and the frequency of the user communication information is counted to obtain the user's communication frequency. It is determined whether the user's communication frequency is greater than a preset communication frequency threshold. If the user's communication frequency is greater than the preset communication frequency threshold, the user object corresponding to the user's communication frequency is determined to have a target relationship, and the set of all user objects having the target relationship is used as the nodes of the social network graph, and the communication relationship between the user objects having the target relationship is used as the edges of the social network graph to construct a user social network graph. The user login information and user transaction information are obtained, and risk group division is performed according to the user login information and the user transaction information to obtain a high-risk group. Association marking is performed on the high-risk group in the user social network graph to obtain a potential abnormal group. The login status information and user account information of the potential abnormal group are compared and analyzed, and a target user to be identified is determined according to the analysis result. The number of cancellations of the target user to be identified is obtained, and when the number of cancellations is greater than a preset cancellation number threshold, the target user to be identified is marked as target abnormal data. Thus, it effectively realizes the accurate identification of target abnormal data with fraud risks among users and improves the reliability of the identification.

[0059] The method of this embodiment can be applied to the identification of abnormal data in a financial business system. Among them, the financial business system includes a wealth management business system, a lending business system, a consumption business system, etc. The database refers to the database of the above financial business system. The user communication information refers to the communication information between logged-in users in the above financial business system. The user login information and user transaction information refer to the user login events and user transaction events that occur in the above financial business system. The target user to be identified refers to a determined suspicious target user, and the target abnormal data refers to a target user identified as a member of a black production gang. The above user communication information, user login information, and user transaction information are all recorded in the database of the above financial business system and are extracted from the database of the above financial business system.

[0060] In this embodiment, by obtaining the communication information among users, between users and financial experts, and between users and investment advisors in the financial management business system, the communication information is used to determine whether there is a target relationship between the user and the user object. In the case of the existence of the target relationship, a user social network graph related to the user is constructed. By obtaining the login information of the user logging in to the financial management business system and the transaction information of the user on the financial management business system, the login information and the transaction information are used to determine whether there are high risks in the current login behavior and transaction behavior of the user, and a high-risk group belonging to high risks is divided. The high-risk group and the user social network graph are associated and marked to find a larger range of potential abnormal groups related to the high-risk group. Then, by comprehensively analyzing and comparing the historical login status (login terminal, location, time, etc.), user name, and whether it is real-name of the potential abnormal group in the financial management business system, a suspicious target user to be identified is determined. Finally, whether the target user to be identified belongs to the target abnormal data is ultimately determined by the number of cancellations of the target user to be identified in the financial management business system. If the number of cancellations of the target user to be identified in the financial management business system is greater than the preset cancellation number threshold, the target user to be identified is determined as the target abnormal data with fraud nature. This provides reliable and effective data support for subsequent anti-fraud.

[0061] In some alternative implementation manners of this embodiment, the steps of obtaining the user login information and the user transaction information, and dividing the risk group according to the user login information and the user transaction information to obtain the high-risk group include the following steps:

[0062] Obtain the user login information and the user transaction information, and perform user clustering division according to the user login information and the user transaction information to obtain a user clustering group;

[0063] The user login information includes the login frequency, and the user transaction information includes the transaction amount and the transaction frequency. The user feature vector matrix is constructed based on the user login information and the user transaction information, and clustering is performed on the basis of the user feature vector matrix to obtain a user clustering group clustered according to user features.

[0064] Judge the abnormal behavior of the user clustering group to obtain the high-risk group.

[0065] The abnormal behaviors include abnormal login frequency, abnormal transaction amount, abnormal transaction time, abnormal browsing behavior, etc. By identifying the abnormal behaviors of the user clustering group and calculating the abnormal behavior ratio, it is effectively determined whether the user clustering group is a high-risk group.

[0066] In this embodiment, by obtaining user login information and user transaction information, user clustering division is performed according to the user login information and the user transaction information to obtain user clustering groups; abnormal behavior judgment is performed on the user clustering groups to obtain high-risk groups. Thus, it effectively realizes the division of high-risk groups according to the clustering situation and abnormal behavior of users, so as to facilitate the subsequent identification of potential abnormal groups.

[0067] In some optional implementation manners of this embodiment, the obtaining user login information and user transaction information, and performing user clustering division according to the user login information and the user transaction information to obtain user clustering groups include the following steps:

[0068] Construct a user feature vector matrix according to the user login information and the user transaction information;

[0069] In this embodiment, for each user, specific values of their login frequency, transaction amount, and transaction frequency are extracted. The three feature values of each user are combined into a feature vector. For example, for user A, its feature vector may be [login frequency A, transaction amount A, transaction frequency A]. The feature vectors of all users are arranged row by row to form a two-dimensional matrix. Each row represents a user, and each column represents a feature. After performing the above steps of processing, the user feature vector matrix is obtained. Among them, if the dimensionality differences between features are large, consider performing standardization processing on the feature vector matrix to eliminate the influence of dimensionality differences on the clustering result, and this standardization can be based on the mean and standard deviation of each feature.

[0070] Perform data standardization processing on the user feature vector matrix to obtain a standard user feature vector matrix;

[0071] In this embodiment, numerical features in the user feature vector matrix are standardized to eliminate the dimensionality differences between different features. In this embodiment, Z-score standardization is used for the standardization processing. Z-score standardization converts the feature values into a distribution with a mean of 0 and a standard deviation of 1, thereby obtaining a standard user feature vector matrix.

[0072] Perform clustering calculation on the standard user feature vector matrix according to a preset clustering algorithm to obtain the user clustering groups.

[0073] In this embodiment, the preset clustering algorithm can adopt the K-means clustering algorithm. The steps of performing clustering calculation according to K-means clustering include: determining the number of clusters K, where the value of K represents the number of groups into which users are expected to be divided; randomly selecting K users from the standard user feature vector matrix as the initial cluster centers; calculating the distance from each user to the cluster centers: using measurement methods such as Euclidean distance and Manhattan distance to calculate the distance from each user to each cluster center; assigning users to the nearest cluster center; according to the distance measurement results, assigning each user to the group to which the nearest cluster center belongs; recalculating the cluster centers of each group, that is, calculating the mean of all user feature vectors within the group; checking whether the cluster centers have changed or reached the preset number of iterations. If the convergence condition is met, stop the iteration; otherwise, continue the iteration. When the above clustering calculation steps are completed, the user clustering groups are obtained.

[0074] In this embodiment, a user feature vector matrix is constructed according to the user login information and the user transaction information; the user feature vector matrix is subjected to data standardization processing to obtain a standard user feature vector matrix; the standard user feature vector matrix is subjected to clustering calculation according to a preset clustering algorithm to obtain the user clustering groups. Thus, it effectively realizes effective user clustering according to the characteristics of users, so as to facilitate subsequent abnormal behavior judgment and processing.

[0075] In some optional implementation manners of this embodiment, the steps of performing abnormal behavior judgment on the user clustering groups to obtain high-risk groups include:

[0076] Obtain the user behavior information of the user clustering groups;

[0077] In this embodiment, the user behavior information related to each user in the user clustering groups can be extracted from the database according to the user identifiers corresponding to the user clustering groups. The user behavior information includes login behavior, transaction behavior, browsing behavior, etc.

[0078] Perform abnormal behavior recognition on the user behavior information according to the abnormal behavior rules to obtain user abnormal behavior information;

[0079] In this embodiment, the abnormal behavior rules refer to the rule information defining abnormal behaviors. The abnormal behavior rules can include abnormal login frequency rules, abnormal transaction amount rules, abnormal transaction time rules, abnormal browsing behavior rules, etc. By using the defined abnormal behavior rules to match the user behavior information one by one, the user behavior information that conforms to the abnormal behavior rules is identified as the user abnormal behavior information.

[0080] Calculate the group abnormal behavior ratio according to the user abnormal behavior information and the total number of users in the user clustering groups;

[0081] In this embodiment, for each user clustering group, the number of identified user abnormal behavior information, that is, the number of users with abnormal behavior, is counted, and then according to the total number of users in the user clustering group and the number of users with abnormal behavior, the group abnormal behavior ratio is calculated. The calculation formula for the group abnormal behavior ratio is: group abnormal behavior ratio = number of users with abnormal behavior / total number of users.

[0082] Determine whether the group abnormal behavior ratio is greater than a preset behavior ratio threshold;

[0083] In this embodiment, the preset behavior ratio threshold is used to determine whether the group abnormal behavior ratio reaches a high-risk level. In this embodiment, the preset behavior ratio threshold is set to 60%, and corresponding settings and adjustments can be made according to the actual situation.

[0084] If the group abnormal behavior ratio is greater than the preset behavior ratio threshold, then the user clustering group corresponding to the group abnormal behavior ratio is determined as the high-risk group;

[0085] In this embodiment, the user clustering group can be determined as the high-risk group by adding label information corresponding to the high-risk group to the user clustering group. For example, add the label "high-risk group" to the user clustering group to indicate that the user clustering group is a high-risk group.

[0086] If the group abnormal behavior ratio is less than or equal to the preset behavior ratio threshold, then the user clustering group corresponding to the group abnormal behavior ratio is determined as a non-high-risk group.

[0087] In this embodiment, the user clustering group can be determined as a non-high-risk group by adding label information corresponding to the non-high-risk group to the user clustering group. For example, add the label "non-high-risk group" to the user clustering group to indicate that the user clustering group is a non-high-risk group.

[0088] In this embodiment, by obtaining the user behavior information of the user clustering group; identifying abnormal behavior of the user behavior information according to the abnormal behavior rule to obtain user abnormal behavior information; calculating the group abnormal behavior ratio according to the user abnormal behavior information and the total number of users in the user clustering group; determining whether the group abnormal behavior ratio is greater than the preset behavior ratio threshold; if the group abnormal behavior ratio is greater than the preset behavior ratio threshold, then the user clustering group corresponding to the group abnormal behavior ratio is determined as the high-risk group; if the group abnormal behavior ratio is less than or equal to the preset behavior ratio threshold, then the user clustering group corresponding to the group abnormal behavior ratio is determined as a non-high-risk group. Thus, it effectively realizes judging whether it belongs to a high-risk group according to the abnormal behavior of the user clustering group, providing a reliable basis for subsequent identification of potential abnormal groups.

[0089] In some alternative implementation manners of this embodiment, the step of obtaining the potential abnormal group by performing associated marking on the high-risk group in the user social network graph includes the following steps:

[0090] Extract a corresponding associated subgraph in the user social network graph according to the high-risk group;

[0091] In this embodiment, by finding all the nodes and edges associated with the users in the high-risk group in the user social network graph, an associated subgraph is formed, where the associated subgraph includes the users in the high-risk group and their directly and indirectly associated users in the network.

[0092] Judge whether there is an edge connection between the users in the high-risk group according to the edge set of the associated subgraph;

[0093] In this embodiment, by traversing the edge set of the associated subgraph, it is judged whether there is an edge connection between the users in the high-risk group. If there is an edge connection, it means that there is a direct connection between these users in the social network.

[0094] If there is an edge connection between the users in the high-risk group, mark the users with edge connections in the high-risk group as potential users to obtain a set of potential users;

[0095] In this embodiment, by adding label information corresponding to potential users to the users with edge connections in the high-risk group for marking processing. For example, if a label of "determined as potential user" is added between the users with edge connections to effectively mark them. After the marking is completed, the marked potential users are recorded in a dataset to form a set of potential users.

[0096] If there is no edge connection between the users in the high-risk group, continue to judge whether there is an edge connection between the users in the high-risk group until all the users in the high-risk group are judged;

[0097] In this embodiment, it is necessary to traverse all the users in the entire high-risk group before it can be determined that the judgment is completed. When the judgment on whether there is an edge connection for all the users in the high-risk group is completed, the current judgment step is stopped and the result is output.

[0098] Calculate the connection degree between each user and other users in the set of potential users to obtain the user connection centrality;

[0099] In this embodiment, an adjacency matrix or an adjacency list can be used to represent the connection relationship between users in the set of potential users. Among them, the adjacency matrix is a two-dimensional array, where the rows and columns represent users respectively, and the value of the array element represents the connection status between users (1 represents connected, 0 represents not connected). The adjacency list is a dictionary or hash table, where the key is the user ID, and the value is a list of other user IDs directly connected to this user. By traversing each user in the set of potential users and then counting the number of direct connections between the user and other users, it can be achieved by traversing the rows of the adjacency matrix or the values of the adjacency list. Record the connection degree of each user in a data structure, such as a list or a dictionary, which uses the user ID as the key and the connection degree as the value. Thus, the user connection centrality of each user in the set of potential users can be effectively obtained, and this user connection centrality is used to reflect the central position and influence of the user in the social network.

[0100] Extract the historical behavior information corresponding to the set of potential users, and perform feature extraction on the historical behavior information to obtain a user behavior feature vector;

[0101] In this embodiment, the historical behavior information corresponding to the set of potential users can be extracted from the database through a historical behavior information extraction identifier. Then, by performing feature extraction on the historical behavior information, a user behavior feature vector can be obtained. Among them, the user behavior feature vector can include login frequency, transaction amount, browsing behavior, social interaction, etc.

[0102] Fuse the user connection centrality and the user behavior feature vector according to the graph neural network algorithm to obtain a user risk degree evaluation value;

[0103] In this embodiment, the graph neural network algorithm can adopt GCN (Graph Convolutional Network). The graph convolutional network GCN propagates information through the nodes (users) and edges (connections between users) in the graph structure, so as to learn the embedding representation of the nodes. By taking the user behavior feature vector and the connection centrality as the input of the GCN model, the GCN model applies one or more graph convolutional layers to aggregate information from neighboring nodes. Among them, these layers will update the embedding representation of each node to make it contain information from its neighbors. Then apply a fully connected layer (or other types of layers, such as the softmax layer, depending on your task) to generate the final output, that is, the user risk degree evaluation value. The training steps of the above GCN model include: obtaining users with known risk degrees as training data and labeling them (high risk, medium risk, low risk, etc.); selecting an appropriate loss function (such as cross-entropy loss) to measure the difference between the risk degree predicted by the model and the true label; selecting an optimizer (such as Adam) to update the weights of the model to minimize the loss function and obtain a GCN model that can effectively generate user risk degree evaluation values.

[0104] Cluster and partition the set of potential users according to the user risk degree evaluation value to obtain the potential abnormal group.

[0105] In this embodiment, the K-means clustering algorithm can be used to cluster and partition the set of potential users according to the user risk degree evaluation value to obtain potential abnormal groups with different risk levels.

[0106] In this embodiment, the corresponding associated subgraph is extracted from the user social network graph according to the high-risk group; it is judged whether there is an edge connection between the users in the high-risk group according to the edge set of the associated subgraph; if there is an edge connection between the users in the high-risk group, the users with edge connections in the high-risk group are marked as potential users to obtain a set of potential users; calculate the connection degree between each user and other users in the set of potential users to obtain the user connection centrality; extract the historical behavior information corresponding to the set of potential users, and perform feature extraction on the historical behavior information to obtain a user behavior feature vector; fuse the user connection centrality and the user behavior feature vector according to the graph neural network algorithm to obtain a user risk degree evaluation value; cluster and partition the set of potential users according to the user risk degree evaluation value to obtain the potential abnormal group. Thus, it effectively realizes clustering and partitioning according to the user association relationship and historical behavior characteristics in the high-risk group to improve the recognition accuracy of the potential abnormal group.

[0107] In some optional implementation manners of this embodiment, the steps of comparing and analyzing the login state information and user account information of the potential abnormal group and determining the target user to be recognized according to the analysis result include:

[0108] Obtain the login state information corresponding to the potential abnormal group;

[0109] In this embodiment, the login state information can be extracted from the database according to the login information extraction identifier. Among them, the login state information includes the login terminal, login time, login location, login status, login record, etc.

[0110] Judge users with the same device number according to the login state information;

[0111] In this embodiment, by comparing the login terminal information in the login status information, if there is the same login terminal information, the users with the same login terminal information are preliminarily determined to be users from the same device source; obtain the registration information of the users preliminarily determined to be from the same device source, where the registration information includes registration terminal information and registrant information; by comparing the registration terminal information and the registrant information, if the registration terminal information and the registrant information are the same, the users with the same registration terminal information and the same registrant information are determined to be users with the same device number.

[0112] Judge users with the same device information according to the account information of the users with the same device number;

[0113] In this embodiment, obtain the account information of the users with the same device number, where the account information includes user name information and real name information; by comparing the user name information and the real name information, if the user name information and the real name information are the same, the users with the same user name information and the same real name information are determined to be users of the same sub-device; obtain the payment information of the users determined to be users of the same sub-device, where the payment information includes payment number information and binding number information; by comparing the payment number information and the binding number information, if the payment number information and the binding number information are the same, the users with the same payment number information and the same binding number information are determined to be users with the same IP address; obtain the device number information of the users determined to be users with the same IP address; by comparing the device number information, if there is the same device number information, the users with the same device number information are determined to be users with the same device information.

[0114] Obtain the user account information of the users with the same device information, and perform user marking on the potential abnormal group according to the user account information to obtain the target user to be identified.

[0115] In this embodiment, obtain the user account information of the users determined to be users with the same device information; by comparing the user account information, if there is the same user account information, the users with the same user account information are determined to be users with the same associated account; obtain the device information of the users determined to be users with the same associated account; by comparing the device information, if the device information is the same, the users with the same device information and the users with the same associated account are determined to be the target users to be identified.

[0116] In this embodiment, the login status information corresponding to the potential abnormal group is obtained; the users with the same device number are judged according to the login status information; the users with the same device information are judged according to the account information of the users with the same device number; the user account information of the users with the same device information is obtained, and the potential abnormal group is marked with users according to the user account information to obtain the target user to be identified. Thus, it is effectively realized to determine the suspicious target user to be identified according to the login situation and associated accounts of the users, so as to facilitate the subsequent determination of the potential abnormal team.

[0117] In some optional implementation manners of this embodiment, the steps of obtaining the cancellation times of the target user to be identified and marking the team where the target user to be identified is located as a potential abnormal team when the cancellation times are greater than or equal to a preset cancellation times threshold include the following steps:

[0118] Obtain the account registration time of the target user to be identified, and count the cancellation times according to the account registration time;

[0119] In this embodiment, an information identifier for registration can be used to obtain the account registration time of the target user to be identified and their account cancellation records from the database. For each target user to be identified, the cancellation times are calculated according to their account cancellation records. In this embodiment, when calculating the cancellation times, a time range (such as the past three months, half a year, etc.) can be set to consider the user's cancellation behavior within this time range to obtain more effective cancellation times.

[0120] Judge whether the cancellation times are greater than or equal to the preset cancellation times threshold;

[0121] In this embodiment, the preset cancellation times threshold is a preset threshold for judging whether the cancellation times are abnormal, and this preset cancellation times threshold can be determined according to historical data, industry standards and business logic. In this embodiment, the preset cancellation times threshold is set to 3 times, and corresponding settings and adjustments can be made according to the actual situation.

[0122] If the cancellation times are greater than or equal to the preset cancellation times threshold, then determine the target user to be identified as a risk user;

[0123] In this embodiment, the target user to be identified can be determined as a risk user by adding label information corresponding to the risk user to the target user to be identified ID. For example, add the label of "risk user" to the target user to be identified ID for marking to facilitate subsequent identification.

[0124] If the cancellation times are less than the preset cancellation times threshold, then determine the target user to be identified as a non-risk user;

[0125] In this embodiment, the target user ID to be identified can be determined as a non-risk user by adding tag information corresponding to non-risk users to the target user ID to be identified. For example, adding a tag of "non-risk user" to the target user ID to be identified for marking to facilitate subsequent identification.

[0126] Obtain the associated transaction information and transaction amount of the risk user, and determine the frequently traded accounts and high-risk accounts according to the associated transaction information and the transaction amount;

[0127] In this embodiment, using a graph algorithm, through the associated transaction information of the risk user, obtain the transaction frequency of all accounts. If the transaction frequency is greater than the preset transaction frequency threshold, it is determined as a frequently traded account; according to the transaction amount of the frequently traded account, obtain the cumulative transaction amount of the frequently traded account. If the cumulative transaction amount is greater than the preset transaction amount threshold, determine the frequently traded account as a high-risk account.

[0128] Perform abnormal data confirmation according to the frequently traded accounts and the high-risk accounts to obtain the target abnormal data.

[0129] In this embodiment, according to the frequently traded accounts and the high-risk accounts, use the hierarchical clustering algorithm to obtain the associated accounts of the risk user. If there is a high-risk account among the associated accounts, determine the account as the target abnormal data.

[0130] In this embodiment, by obtaining the account registration time of the target user to be identified, and counting the number of cancellations according to the account registration time; determining whether the number of cancellations is greater than or equal to the preset number of cancellation threshold; if the number of cancellations is greater than or equal to the preset number of cancellation threshold, determine the target user to be identified as a risk user; obtain the associated transaction information and transaction amount of the risk user, and determine the frequently traded accounts and high-risk accounts according to the associated transaction information and the transaction amount; perform abnormal data confirmation according to the frequently traded accounts and the high-risk accounts to obtain the target abnormal data. Thus, it effectively realizes the accurate identification of potential abnormal teams among users from multiple dimensions to facilitate subsequent investigation.

[0131] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), etc., or a random access memory (RAM), etc.

[0132] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the sequence indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0133] Further reference Figure 3 to Figure 1 As an implementation of the method shown above, an embodiment of an abnormal data recognition device is provided in this application. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0134] As Figure 3 shown, the abnormal data recognition device 900 described in this embodiment includes: a frequency acquisition module 901, a threshold judgment module 902, a graph construction module 903, a relationship determination module 904, a user clustering module 905, a group recognition module 906, a target marking module 907, and an abnormal data recognition module 908. Among them:

[0135] The frequency acquisition module 901 is used to extract user communication information within a preset time range from the database and perform frequency statistics on the user communication information to obtain the user communication frequency.

[0136] The threshold judgment module 902 is used to judge whether the user communication frequency is greater than a preset communication frequency threshold.

[0137] The graph construction module 903 is used to, if the user communication frequency is greater than the preset communication frequency threshold, determine the user object corresponding to the user communication frequency as having a target relationship, and use the set of all user objects having the target relationship as the nodes of the social network graph, and use the communication relationship of the user objects having the target relationship as the edges of the social network graph to construct a user social network graph.

[0138] The relationship determination module 904 is used to, if the user communication frequency is less than or equal to the preset communication frequency threshold, determine that the user object corresponding to the user communication frequency does not have a target relationship.

[0139] The user clustering module 905 is used to obtain user login information and user transaction information, perform risk group division based on the user login information and the user transaction information, and obtain a high-risk group;

[0140] The group identification module 906 is used to perform association marking in the user social network graph according to the high-risk group to obtain a potential abnormal group;

[0141] The target marking module 907 is used to perform a comparative analysis on the login state information and user account information of the potential abnormal group, and determine the target user to be identified according to the analysis result;

[0142] The abnormal data identification module 908 is used to obtain the number of cancellations of the target user to be identified, and when the number of cancellations is greater than a preset cancellation number threshold, mark the target user to be identified as target abnormal data.

[0143] In this embodiment, by adopting the above abnormal data identification device, it is possible to extract user communication information within a preset time range from the database, perform frequency statistics on the user communication information to obtain the user communication frequency; determine whether the user communication frequency is greater than a preset communication frequency threshold; if the user communication frequency is greater than the preset communication frequency threshold, determine the user object corresponding to the user communication frequency as having a target relationship, and use the set of all user objects having the target relationship as the nodes of the social network graph, and use the communication relationship of the user objects having the target relationship as the edges of the social network graph to construct a user social network graph; obtain user login information and user transaction information, perform risk group division based on the user login information and the user transaction information to obtain a high-risk group; perform association marking in the user social network graph according to the high-risk group to obtain a potential abnormal group; perform a comparative analysis on the login state information and user account information of the potential abnormal group, and determine the target user to be identified according to the analysis result; obtain the number of cancellations of the target user to be identified, and when the number of cancellations is greater than a preset cancellation number threshold, mark the target user to be identified as target abnormal data. Thus, it effectively realizes the accurate identification of target abnormal data with fraud risks among users and improves the reliability of the identification.

[0144] In some optional implementation manners of this embodiment, the user clustering module 905 includes: a clustering division unit and a behavior judgment unit. Among them:

[0145] The clustering division unit is used to obtain user login information and user transaction information, perform user clustering division according to the user login information and the user transaction information, and obtain user clustering groups;

[0146] The behavior judgment unit is used to judge the abnormal behaviors of the user clustering groups to obtain high-risk groups.

[0147] In this embodiment, by setting up the user clustering module 905 including a clustering division unit and a behavior judgment unit, the division of high-risk groups can be effectively realized according to the clustering situation and abnormal behaviors of users, so as to facilitate the subsequent identification of potential abnormal groups.

[0148] In some optional implementation manners of this embodiment, the clustering division unit includes: a matrix construction subunit, a matrix standardization subunit, and a group clustering subunit. Among them:

[0149] The matrix construction subunit is used to construct a user feature vector matrix according to the user login information and the user transaction information;

[0150] The matrix standardization subunit is used to perform data standardization processing on the user feature vector matrix to obtain a standard user feature vector matrix;

[0151] The group clustering subunit is used to perform clustering calculation on the standard user feature vector matrix according to a preset clustering algorithm to obtain the user clustering groups.

[0152] In this embodiment, by setting up a clustering division unit including a matrix construction subunit, a matrix standardization subunit, and a group clustering subunit, the effective user clustering can be realized according to the characteristics of users, so as to facilitate the subsequent judgment and processing of abnormal behaviors.

[0153] In some optional implementation manners of this embodiment, the behavior judgment unit includes: a behavior information acquisition subunit, an abnormal behavior recognition subunit, a behavior ratio calculation subunit, a ratio judgment subunit, a first group determination subunit, and a second group determination subunit. Among them:

[0154] The behavior information acquisition subunit is used to acquire the user behavior information of the user clustering groups;

[0155] The abnormal behavior recognition subunit is used to recognize abnormal behaviors according to abnormal behavior rules for the user behavior information to obtain user abnormal behavior information;

[0156] The behavior ratio calculation subunit is used to calculate the group abnormal behavior ratio according to the user abnormal behavior information and the total number of users in the user clustering groups;

[0157] The ratio judgment subunit is used to judge whether the group abnormal behavior ratio is greater than a preset behavior ratio threshold;

[0158] The first group determination subunit is configured to, if the abnormal behavior ratio of the group is greater than the preset behavior ratio threshold, determine the user clustering group corresponding to the abnormal behavior ratio of the group as the high-risk group;

[0159] The second group determination subunit is configured to, if the abnormal behavior ratio of the group is less than or equal to the preset behavior ratio threshold, determine the user clustering group corresponding to the abnormal behavior ratio of the group as a non-high-risk group.

[0160] In this embodiment, by setting a behavior judgment unit including a behavior information acquisition subunit, an abnormal behavior recognition subunit, a behavior ratio calculation subunit, a ratio judgment subunit, a first group determination subunit, and a second group determination subunit, it is effectively realized to judge whether a user belongs to a high-risk group according to the abnormal behavior of the user clustering group, so as to provide a reliable basis for subsequent identification of potential abnormal groups.

[0161] In some alternative implementation manners of this embodiment, the group identification module 906 includes: a sub-graph extraction unit, an edge connection judgment unit, a user marking unit, a judgment continuation unit, a connection degree calculation unit, a user feature extraction unit, a risk degree fusion unit, and a group division unit. Among them:

[0162] The sub-graph extraction unit is configured to extract a corresponding associated sub-graph from the user social network graph according to the high-risk group;

[0163] The edge connection judgment unit is configured to judge whether there is an edge connection between the users of the high-risk group according to the edge set of the associated sub-graph;

[0164] The user marking unit is configured to, if there is an edge connection between the users of the high-risk group, mark the users with edge connections in the high-risk group as potential users to obtain a set of potential users;

[0165] The judgment continuation unit is configured to, if there is no edge connection between the users of the high-risk group, continue to judge whether there is an edge connection between the users of the high-risk group until all users in the high-risk group have been judged;

[0166] The connection degree calculation unit is configured to calculate the connection degree between each user in the set of potential users and other users to obtain the user connection centrality;

[0167] The user feature extraction unit is configured to extract the historical behavior information corresponding to the set of potential users and perform feature extraction on the historical behavior information to obtain a user behavior feature vector;

[0168] The risk degree fusion unit is used to fuse the user connection centrality and the user behavior feature vector according to the graph neural network algorithm to obtain a user risk degree evaluation value;

[0169] The group division unit is used to cluster and divide the set of potential users according to the user risk degree evaluation value to obtain the potential abnormal group.

[0170] In this embodiment, by setting the group identification module 906 including a sub-graph extraction unit, an edge connection judgment unit, a user marking unit, a judgment continuation unit, a connection degree calculation unit, a user feature extraction unit, a risk degree fusion unit, and a group division unit, the clustering division can be effectively realized according to the user association relationship and historical behavior characteristics in the high-risk group, so as to improve the recognition accuracy of the potential abnormal group.

[0171] In some optional implementation manners of this embodiment, the target marking module 907 includes: a login information acquisition unit, a device user judgment unit, a same device judgment unit, and a target user marking unit. Among them:

[0172] The login information acquisition unit is used to acquire the login state information corresponding to the potential abnormal group;

[0173] The device user judgment unit is used to judge the users with the same device number according to the login state information;

[0174] The same device judgment unit is used to judge the users with the same device information according to the account information of the users with the same device number;

[0175] The target user marking unit is used to acquire the user account information of the users with the same device information, and mark the potential abnormal group according to the user account information to obtain the target user to be identified.

[0176] In this embodiment, by setting the target marking module 907 including a login information acquisition unit, a device user judgment unit, a same device judgment unit, and a target user marking unit, the suspicious target user to be identified can be effectively determined according to the login situation and associated accounts of the user, so as to facilitate the subsequent determination of the potential abnormal team.

[0177] In some optional implementation manners of this embodiment, the abnormal data identification module 908 includes: a cancellation times acquisition unit, a times judgment unit, a first user determination unit, a second user determination unit, an account determination unit, and an abnormal data determination unit. Among them:

[0178] The cancellation times acquisition unit is used to acquire the account registration time of the target user to be identified, and count the cancellation times according to the account registration time;

[0179] The number judgment unit is configured to judge whether the cancellation times are greater than or equal to a preset cancellation times threshold value;

[0180] The first user determination unit is configured to, if the cancellation times are greater than or equal to the preset cancellation times threshold value, determine the target user to be identified as a risk user;

[0181] The second user determination unit is configured to, if the cancellation times are less than the preset cancellation times threshold value, determine the target user to be identified as a non-risk user;

[0182] The account determination unit is configured to obtain the associated transaction information and transaction amounts of the risk user, and determine frequent transaction accounts and high-risk accounts according to the associated transaction information and the transaction amounts;

[0183] The abnormal data determination unit is configured to perform abnormal data confirmation according to the frequent transaction accounts and the high-risk accounts to obtain the target abnormal data.

[0184] In this embodiment, by setting the abnormal data identification module 908 including a cancellation times acquisition unit, a number judgment unit, a first user determination unit, a second user determination unit, an account determination unit, and an abnormal data determination unit, it is effectively realized to accurately identify potential abnormal teams among users from multiple dimensions, so as to facilitate subsequent investigation of them.

[0185] To solve the above technical problems, an embodiment of the present application also provides a computer device. For details, please refer to Figure 4 , Figure 4 This is the basic structural block diagram of the computer device in this embodiment.

[0186] The computer device 11 includes a memory 111, a processor 112, and a network interface 113 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 11 with components 111-113 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0187] The computer device may be a computing device such as a desktop computer, a notebook, a palm computer, or a cloud server. The computer device can interact with a user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0188] The memory 111 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 111 may be an internal storage unit of the computer device 11, such as the hard disk or memory of the computer device 11. In other embodiments, the memory 111 may also be an external storage device of the computer device 11, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 11. Of course, the memory 111 may also include both the internal storage unit and the external storage device of the computer device 11. In this embodiment, the memory 111 is generally used to store the operating system installed on the computer device 11 and various application software, such as computer-readable instructions of the abnormal data recognition method. In addition, the memory 111 can also be used to temporarily store various data that have been output or will be output.

[0189] In some embodiments, the processor 112 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 112 is generally used to control the overall operation of the computer device 11. In this embodiment, the processor 112 is used to run the computer-readable instructions stored in the memory 111 or process data, such as running the computer-readable instructions of the abnormal data recognition method.

[0190] The network interface 113 may include a wireless network interface or a wired network interface, and the network interface 113 is generally used to establish a communication connection between the computer device 11 and other electronic devices.

[0191] In this embodiment, by using the above computer device, it is possible to extract user communication information within a preset time range from a database, perform frequency statistics on the user communication information to obtain the user's communication frequency, determine whether the user's communication frequency is greater than a preset communication frequency threshold, if the user's communication frequency is greater than the preset communication frequency threshold, determine the user object corresponding to the user's communication frequency as having a target relationship, and use the set of all user objects having the target relationship as the nodes of the social network graph, and use the communication relationship between the user objects having the target relationship as the edges of the social network graph to construct a user social network graph; obtain user login information and user transaction information, perform risk group division according to the user login information and the user transaction information to obtain a high-risk group; perform associated marking on the high-risk group in the user social network graph to obtain a potential abnormal group; perform comparative analysis on the login status information and user account information of the potential abnormal group, and determine a target user to be identified according to the analysis result; obtain the number of cancellation times of the target user to be identified, and when the number of cancellation times is greater than a preset cancellation times threshold, mark the target user to be identified as target abnormal data. Thus, it effectively realizes the accurate identification of target abnormal data with fraud risks among users and improves the reliability of identification.

[0192] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that the at least one processor executes the steps of the abnormal data identification method as described above.

[0193] In this embodiment, by using the above-mentioned computer-readable storage medium, user communication information within a preset time range can be extracted from the database, and the frequency of the user communication information can be counted to obtain the user's communication frequency. It is determined whether the user's communication frequency is greater than a preset communication frequency threshold. If the user's communication frequency is greater than the preset communication frequency threshold, the user object corresponding to the user's communication frequency is determined to have a target relationship, and the set of all user objects having the target relationship is used as the nodes of the social network graph, and the communication relationship between the user objects having the target relationship is used as the edges of the social network graph to construct a user social network graph. The user login information and user transaction information are obtained, and risk group division is performed according to the user login information and the user transaction information to obtain a high-risk group. Association marking is performed on the high-risk group in the user social network graph to obtain a potential abnormal group. The login status information and user account information of the potential abnormal group are compared and analyzed, and the target user to be identified is determined according to the analysis result. The number of cancellation times of the target user to be identified is obtained, and when the number of cancellation times is greater than a preset cancellation times threshold, the target user to be identified is marked as target abnormal data. Thus, the accurate identification of target abnormal data with fraud risks among users is effectively achieved, and the reliability of the identification is improved.

[0194] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0195] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The preferred embodiments of the present application are given in the drawings, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is equally within the scope of the patent protection of the present application.

[0196] In the embodiments of this application, the non-company software tools or components presented are merely for illustrative purposes and do not represent actual usage.

Claims

1. A method for identifying abnormal data, characterized in that: The steps include: Extracting user communication information within a preset time range from a database, and performing frequency statistics on the user communication information to obtain the user communication frequency; Determining whether the user's communication frequency is greater than a preset communication frequency threshold; If the communication frequency of the user pair is greater than the preset communication frequency threshold, the user object corresponding to the user pair communication frequency is determined to have a target relationship, and the set of all user objects with the target relationship is used as the node of the social network graph, and the communication relationship of the user objects with the target relationship is used as the edge of the social network graph to construct a user social network graph; Obtaining user login information and user transaction information, and dividing risk groups according to the user login information and the user transaction information to obtain a high-risk group; According to the high-risk groups, the high-risk groups are marked in the user social network graph to obtain potential abnormal groups; Comparative analysis is performed on the login status information and user account information of the potential abnormal group, and target users to be identified are determined based on the analysis results; The logout times of the target user to be identified are obtained, and when the logout times are greater than a preset logout times threshold, the target user to be identified is marked as target abnormal data.

2. The abnormal data identification method according to claim 1, characterized in that: The step of obtaining user login information and user transaction information, dividing risk groups according to the user login information and the user transaction information, and obtaining a high-risk group specifically includes: Acquire user login information and user transaction information, and perform user clustering division according to the user login information and the user transaction information to obtain user cluster groups; Abnormal behavior is judged on the user cluster groups to obtain high-risk groups.

3. The abnormal data identification method according to claim 2, characterized in that: The step of obtaining user login information and user transaction information, and performing user clustering division according to the user login information and the user transaction information to obtain user cluster groups specifically includes: Constructing a user feature vector matrix according to the user login information and the user transaction information; Performing data standardization processing on the user feature vector matrix to obtain a standard user feature vector matrix; The standard user feature vector matrix is ​​clustered according to a preset clustering algorithm to obtain the user cluster group.

4. The abnormal data identification method according to claim 2, characterized in that: The step of judging abnormal behavior of the user cluster group to obtain a high-risk group specifically includes: Obtaining user behavior information of the user cluster group; Performing abnormal behavior identification on the user behavior information according to abnormal behavior rules to obtain abnormal behavior information of the user; Calculating a group abnormal behavior ratio according to the user abnormal behavior information and the total number of users in the user cluster group; Determine whether the proportion of abnormal behaviors of the group is greater than a preset behavior proportion threshold; If the abnormal behavior ratio of the group is greater than the preset behavior ratio threshold, the user cluster group corresponding to the abnormal behavior ratio of the group is determined as the high-risk group.

5. The abnormal data identification method according to claim 1, characterized in that: The step of performing association marking in the user social network graph according to the high-risk group to obtain a potential abnormal group specifically includes: Extracting corresponding associated subgraphs from the user social network graph according to the high-risk groups; Determining whether there is an edge connection between users in the high-risk group according to the edge set of the associated subgraph; If there is an edge connection between users in the high-risk group, mark the users in the high-risk group with the edge connection as potential users to obtain a potential user set; Calculate the connection degree between each user in the potential user set and other users to obtain the user connection centrality; Extracting historical behavior information corresponding to the potential user set, and performing feature extraction on the historical behavior information to obtain a user behavior feature vector; According to the graph neural network algorithm, the user connection centrality and the user behavior feature vector are integrated to obtain a user risk assessment value; The potential user set is clustered and divided according to the user risk assessment value to obtain the potential abnormal group.

6. The abnormal data identification method according to claim 1, characterized in that: The step of comparing and analyzing the login status information and user account information of the potential abnormal group and determining the target user to be identified according to the analysis result specifically includes: Obtaining the login status information corresponding to the potential abnormal group; Determine the user with the same device number according to the login status information; Determine users with the same device information based on the account information of users with the same device number; The user account information of the user with the same device information is obtained, and the potential abnormal group is marked as a user according to the user account information to obtain the target user to be identified.

7. The abnormal data identification method according to claim 1, characterized in that: The step of obtaining the number of logouts of the target user to be identified, and marking the team to which the target user to be identified belongs as a potential abnormal team when the number of logouts is greater than or equal to a preset logout threshold, specifically includes: Obtaining the account registration time of the target user to be identified, and counting the number of logouts according to the account registration time; Determine whether the number of logouts is greater than or equal to a preset logout threshold; If the number of logouts is greater than or equal to the preset logout threshold, the target user to be identified is determined as a risky user; Obtaining the associated transaction information and transaction amount of the risky user, and determining frequent transaction accounts and high-risk accounts according to the associated transaction information and the transaction amount; Abnormal data is confirmed based on the frequent transaction accounts and the high-risk accounts to obtain the target abnormal data.

8. An abnormal data identification device, characterized in that: include: A frequency acquisition module is used to extract user communication information within a preset time range from the database, and perform frequency statistics on the user communication information to obtain the user communication frequency; A threshold determination module, used to determine whether the user's communication frequency is greater than a preset communication frequency threshold; A graph construction module, configured to determine that the user objects corresponding to the user pair communication frequency have a target relationship if the user pair communication frequency is greater than the preset communication frequency threshold, and to use the set of all user objects having the target relationship as nodes of the social network graph, and to use the communication relationships of the user objects having the target relationship as edges of the social network graph to construct a user social network graph; A user clustering module, used to obtain user login information and user transaction information, and divide risk groups according to the user login information and the user transaction information to obtain a high-risk group; A group identification module, used to associate and mark the high-risk groups in the user social network graph to obtain potential abnormal groups; A target marking module is used to compare and analyze the login status information and user account information of the potential abnormal group, and determine the target user to be identified based on the analysis results; The abnormal data identification module is used to obtain the logout times of the target user to be identified, and when the logout times are greater than a preset logout times threshold, mark the target user to be identified as target abnormal data.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the abnormal data identification method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the abnormal data identification method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Fee setting method and device based on data processing and computer equipment

    CN113065902A

  • Gang fraud risk identification method based on knowledge graph and related equipment

    CN116308824A

  • Money laundering risk analysis method based on graph calculation

    CN117829994A

  • Bank anti-call fraud data model construction method based on multi-feature fusion

    CN117993919A

  • Three-party payment platform risk account assessment method and system based on atlas model

    CN119005987A

Cited By

  • Electric fraud transaction beforehand detection method, device and equipment and storage medium

    CN121190061A

  • Method and device for identifying illegal user on social platform, equipment and medium

    CN121217693A

  • Abnormal user identification method and device, computer equipment and storage medium

    CN121705914A