Account identification method, device, equipment and storage medium
By building an account relationship network and optimizing node representation vectors using vector representation models, the problem of difficult to identify and recall abnormal accounts in the prior art is solved, and higher recognition accuracy and recall rate are achieved.
Patent Information
- Application Number
- CN202210621428.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-02
AI Technical Summary
It is difficult for the prior art to effectively identify and recall abnormal accounts on the platform, especially when there are multiple relationships and complex relationships between accounts.
By obtaining multiple account attributes of the account to be identified, feature extraction is performed, an account relationship network is constructed, and node representation vectors are optimized through vector representation model, and clustering is performed to identify abnormal accounts.
Without affecting the recognition accuracy, the recall rate of abnormal accounts is improved, and abnormal accounts in complex account relationships can be more accurately identified and processed.
Smart Images

Figure CN115131058B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an account identification method, device, equipment and storage medium. Background Art
[0002] At present, with the continuous development of Internet technology, many platforms carry tens of millions of accounts every day. Users can easily apply for an account on any platform and conduct corresponding activities on the platform through the applied account, which brings convenience to users, but also introduces risks to the platform; for example, there may be many users on a platform who use multiple accounts to gain benefits, and there are often many connections between these accounts, forming abnormal accounts. Based on this, how to identify abnormal accounts from each account to be identified has become a research hotspot. Summary of the invention
[0003] The embodiments of the present application provide an account identification method, apparatus, device and storage medium, which can improve the recall rate of abnormal accounts without affecting the identification accuracy through multiple account attributes of each account to be identified.
[0004] On the one hand, an embodiment of the present application provides an account identification method, the method comprising:
[0005] Acquire multiple account attributes of each account to be identified, and perform feature extraction on each account attribute of each account to obtain multiple target account features;
[0006] An account relationship network is constructed using the multiple target account features, wherein a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature;
[0007] A plurality of triplets are sampled from the account relationship network, wherein a triplet includes a central node, a positive sample node, and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node;
[0008] According to the initial representation vectors of each node in each triple, the vector representation model is trained and optimized, and the target representation vectors of each node in the account relationship network are determined by using the optimized vector representation model;
[0009] Based on the target characterization vector of each node, the accounts corresponding to each node are clustered, and abnormal accounts are identified among the multiple accounts participating in the clustering according to the clustering result.
[0010] On the other hand, an embodiment of the present application provides an account identification device, the device comprising:
[0011] an acquisition unit, configured to acquire multiple account attributes of each account to be identified, and perform feature extraction on each account attribute of each account to obtain multiple target account features;
[0012] A processing unit, configured to construct an account relationship network using the multiple target account features, wherein a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature;
[0013] The processing unit is further used to sample a plurality of triplets from the account relationship network, wherein a triplet includes a central node, a positive sample node, and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node;
[0014] The processing unit is further used to train and optimize the vector representation model according to the initial representation vector of each node in each triple, and use the optimized vector representation model to determine the target representation vector of each node in the account relationship network;
[0015] The processing unit is further used to cluster the accounts corresponding to the nodes based on the target characterization vectors of the nodes, and identify abnormal accounts among the multiple accounts participating in the clustering according to the clustering results.
[0016] In another aspect, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the following steps are implemented:
[0017] Acquire multiple account attributes of each account to be identified, and perform feature extraction on each account attribute of each account to obtain multiple target account features;
[0018] An account relationship network is constructed using the multiple target account features, wherein a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature;
[0019] A plurality of triplets are sampled from the account relationship network, wherein a triplet includes a central node, a positive sample node, and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node;
[0020] According to the initial representation vectors of each node in each triple, the vector representation model is trained and optimized, and the target representation vectors of each node in the account relationship network are determined by using the optimized vector representation model;
[0021] Based on the target characterization vector of each node, the accounts corresponding to each node are clustered, and abnormal accounts are identified among the multiple accounts participating in the clustering according to the clustering result.
[0022] In another aspect, an embodiment of the present application provides a computer storage medium, wherein the computer storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the following steps:
[0023] Acquire multiple account attributes of each account to be identified, and perform feature extraction on each account attribute of each account to obtain multiple target account features;
[0024] An account relationship network is constructed using the multiple target account features, wherein a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature;
[0025] A plurality of triplets are sampled from the account relationship network, wherein a triplet includes a central node, a positive sample node, and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node;
[0026] According to the initial representation vectors of each node in each triple, the vector representation model is trained and optimized, and the target representation vectors of each node in the account relationship network are determined by using the optimized vector representation model;
[0027] Based on the target characterization vector of each node, the accounts corresponding to each node are clustered, and abnormal accounts are identified among the multiple accounts participating in the clustering according to the clustering result.
[0028] On the other hand, an embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned account identification method.
[0029] The embodiment of the present application can obtain multiple account attributes of each account to be identified, that is, it can simultaneously refer to account attributes of more dimensions, and support the expansion of account attributes. Accordingly, each account attribute of each account can be feature extracted to obtain multiple target account features; and multiple target account features are used to construct a more accurate account relationship network, in which a node in the account relationship network records the target account feature of an account, and two nodes connected by an edge record the same target account feature. Then, multiple triples can be sampled from the account relationship network, so that according to the initial representation vector of each node in each triple, the vector representation model is trained and optimized, and the optimized vector representation model is used to determine a more accurate target representation vector of each node in the account relationship network, so as to improve the accuracy of the representation vector of each node, and then the target representation vector of each node is used to better represent the account corresponding to the corresponding node; accordingly, based on the target representation vector of each node, the account corresponding to each node can be clustered to obtain a more accurate clustering result, and the abnormal account can be identified in the multiple accounts participating in the clustering according to the clustering result, so as to improve the recall rate of the abnormal account without affecting the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0031] Figure 1a It is a flowchart of an account identification solution provided in an embodiment of the present application;
[0032] Figure 1b is a schematic diagram of interaction between a terminal and a server provided in an embodiment of the present application;
[0033] Figure 2 It is a flowchart of an account identification method provided in an embodiment of the present application;
[0034] Figure 3 It is a flowchart of another account identification method provided in an embodiment of the present application;
[0035] Figure 4 It is a flowchart of another account identification method provided in an embodiment of the present application;
[0036] Figure 5a It is a schematic diagram of an interface display provided in an embodiment of the present application;
[0037] Figure 5bis a schematic diagram of another interface display provided in an embodiment of the present application;
[0038] Figure 5c It is a schematic diagram of another interface display provided in an embodiment of the present application;
[0039] Figure 6 It is a structural diagram of an account identification device provided in an embodiment of the present application;
[0040] Figure 7 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0042] With the continuous development of Internet technology, artificial intelligence (AI) technology has also been better developed. The so-called artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0043] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0044] Among them, machine learning (ML) is a multi-disciplinary interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Deep learning is a technology that uses deep neural network systems to perform machine learning; machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching.
[0045] Based on machine learning / deep learning technology in AI technology, the embodiment of the present application proposes an account identification solution to improve the recall rate of abnormal accounts. It should be noted that the embodiment of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.
[0046] See also Figure 1a As shown, the general principle of the account identification scheme proposed in the embodiment of the present application is as follows: first, multiple account attributes of each account to be identified can be obtained, and features of each account attribute of each account can be extracted to obtain multiple target account features, thereby using multiple target account features to construct an account relationship network; accordingly, multiple triples can be obtained through the account relationship network to train an optimized vector representation model, thereby using the optimized vector representation model to determine the target representation vector of each node in the account relationship network, and based on the target representation vector of each node, the accounts corresponding to each node are clustered, and then abnormal accounts are identified from multiple accounts participating in the clustering according to the clustering results.
[0047] Practice has shown that the account identification scheme proposed in the embodiment of the present application can have at least the following beneficial effects: ① It can refer to account attributes of more dimensions at the same time and support the expansion of account attributes, so as to build a more accurate account relationship network; ② Obtain the target representation vector of each node to improve the accuracy of the representation vector of each node, and then more accurately represent the account corresponding to each node, which can effectively improve the accuracy of the clustering results; ③ The recall rate of abnormal accounts can be improved without affecting the recognition accuracy.
[0048] In a specific implementation, the account identification scheme mentioned above can be executed by a computer device, which can be a terminal or a server; wherein the terminal mentioned here can include but is not limited to: smart phones, tablet computers, laptops, desktop computers, smart watches, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc.; various clients (applications, APPs) can be run in the terminal, such as video playback clients, social clients, browser clients, information flow clients, education clients, etc. The server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing (cloud computing), cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms, etc.; so-called cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space and information services as needed. Furthermore, the computer device mentioned in the embodiments of the present application may be located outside or inside the blockchain network, without limitation. The so-called blockchain network is a network composed of a peer-to-peer network (P2P network) and a blockchain, and the blockchain refers to a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. It is essentially a decentralized database, a string of data blocks (or blocks) generated by cryptographic methods.
[0049] Alternatively, in other embodiments, the account identification scheme mentioned above may also be jointly executed by the server and the terminal; the terminal and the server may be directly or indirectly connected via wired or wireless communication, and this application does not limit this. For example: the terminal may be responsible for obtaining multiple account attributes of each account to be identified to obtain multiple target account features, and then send the multiple target account features to the server; so that the server can use the multiple target account features to construct an account relationship network, and determine the target feature vector of each node through the account relationship network, and then identify abnormal accounts based on the target feature vector of each node, such as Figure 1bAs shown. For another example, the terminal may be responsible for obtaining multiple account attributes of each account to be identified to obtain multiple target account features, and then send the multiple target account features to the server; so that the server can use multiple target account features to build an account relationship network, and determine the target representation vector of each node based on the account relationship network, so as to send the target representation vector of each node to the terminal; then the terminal clusters the accounts corresponding to each node based on the target representation vector of each node, and identifies abnormal accounts among the multiple accounts participating in the clustering according to the clustering results. It should be understood that only two situations in which the terminal and the server jointly execute the above account identification scheme are exemplified here, and are not exhaustive.
[0050] Based on the relevant description of the above account identification scheme, the embodiment of the present application proposes an account identification method, which can be executed by the computer device (terminal or server) mentioned above; or, the account identification method can be executed by the terminal and the server together. For the convenience of explanation, the following description will take the computer device executing the account identification method as an example; please refer to Figure 2 , the account identification method may include the following steps S201-S205:
[0051] S201, obtaining multiple account attributes of each account to be identified, and performing feature extraction on each account attribute of each account to obtain multiple target account features.
[0052] Among them, each account to be identified can refer to: all accounts in any platform, that is, every account created in any platform; and each account to be identified can also refer to: every account complained in any platform, and this application does not limit this.
[0053] It should be noted that any platform may receive multiple complaints every day. In this case, the computer device may treat all complained accounts as accounts to be identified; or, when it is necessary to perform association detection on each account in any platform, the computer device may treat each account in the platform as an account to be identified, and so on.
[0054] Correspondingly, account attributes may include but are not limited to: account information and reported information, etc. This application does not limit the specific content of the account attributes of any account; among which, account information includes but is not limited to: nickname, avatar, personal signature, login address and other information; and reported information includes but is not limited to: reported description and reported chat information.
[0055] In a specific implementation, when a computer device obtains multiple account attributes of each account to be identified, if multiple accounts and account attributes of each account are stored in the computer device's own storage space, the computer device can select each account to be identified and multiple account attributes of each account to be identified from the stored multiple accounts and corresponding account attributes; or, each account and corresponding account attributes of any platform are stored in the target running device, and the computer device can obtain each account to be identified and multiple account attributes of each account to be identified from the target running device, etc. It should be noted that this application does not limit the method of obtaining multiple account attributes of each account to be identified.
[0056] S202, using multiple target account features to construct an account relationship network, where a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature.
[0057] It should be noted that one account in each account to be identified may have at least one target account feature among the above-mentioned multiple target account features, that is, at least one target account feature among the above-mentioned multiple target account features may be an account feature of the same account; correspondingly, one account in each account to be identified may not have any target account feature among the above-mentioned multiple target account features, and this application does not limit this.
[0058] In a specific implementation, a computer device may use multiple target account features to construct an account relationship network for each account to be identified that has at least one target account feature. In this case, the account corresponding to any node in the account relationship network has at least one target account feature. It can be understood that an account that does not have any target account feature may refer to an account that is not associated with any account other than itself in each account to be identified. Then the computer device may determine that the account that does not have any target account feature is a non-abnormal account, so further account identification may not be performed on these accounts. Optionally, the account relationship network may also include nodes corresponding to accounts that do not have any target account feature. Based on this, the target account features recorded in the nodes corresponding to the accounts that do not have any target account feature may be empty, and these nodes are not connected to any node. This application does not limit the specific construction method of the account relationship network.
[0059] It should be noted that the computer device can filter the account attributes of each account based on the SAN framework (SimilarAttributeNetwork, a general framework for building a relationship graph based on similar attributes of accounts) to obtain multiple target account features (i.e., filtered features); that is, the computer device can extract features from the account attributes of each account based on the framework to obtain multiple initial account features, and filter the multiple initial account features to obtain multiple target account features; further, the multiple target account features are used as relationships to associate the nodes corresponding to each account together, thereby forming an account relationship network, such as Figure 3 It is understandable that the SAN framework proposed in the embodiment of the present application can flexibly use the account attributes of each account to build an account relationship network, that is, this framework is universal for merchant-based account clustering and supports the expansion of account attributes.
[0060] S203, sampling multiple triplets from the account relationship network, a triplet includes a central node, a positive sample node and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node.
[0061] It should be noted that the computer device can traverse the account relationship network and generate the above-mentioned multiple triples by random walk, so as to obtain the training data set corresponding to the above-mentioned multiple triples, so as to facilitate the subsequent training optimization of the vector representation model.
[0062] In one embodiment, for any triple among the above-mentioned multiple triples, the computer device may randomly select a node from the account relationship network as the center node (center) included in the any triple, and select a node from the nodes connected to the center node included in the any triple as the positive sample node (positive) included in the any triple; and select a node from the nodes not connected to the center node included in the any triple as the negative sample node (negativate) included in the any triple. In this case, the computer device may randomly select a node in the account relationship network as the center node in the corresponding triple, so as to sample multiple triples from the account relationship network.
[0063] In another embodiment, the computer device may sequentially use the nodes in the account relationship network as the central nodes included in each triple in at least one triple, so as to sample multiple triples from the account relationship network. Accordingly, after determining the central node in any triple, the computer device may select a node from the nodes connected to the central node in the any triple as the positive sample node in the any triple; and select a node from the nodes not connected to the central node in the any triple as the negative sample node in the any triple.
[0064] For example, assuming that the computer device uses node A in the account relationship network as the central node in any triple, the nodes connected to the central node in any triple include: node D, node E and node H, and the nodes not connected to the central node in any triple include: node B, node C, node F, node G, node I, node J and node K, then the computer device can select a node from node D, node E and node H as the positive sample node in any triple, and select a node from node B, node C, node F, node G, node I, node J and node K as the negative sample node in any triple, such as Figure 3 shown.
[0065] Furthermore, for any triplet among the above-mentioned multiple triples, the methods for determining the positive sample nodes in the any triplet include but are not limited to the following:
[0066] The first determination method: the computer device can randomly select a node from the nodes connected to the central node in any triplet as the positive sample node in any triplet; that is, the computer device can select a node from the nodes connected to the central node in any triplet with equal probability as the positive sample node in any triplet.
[0067] The second determination method: the computer device can determine the weight value of the edge corresponding to each node in the nodes connected to the central node in any triple, and perform sampling processing in the nodes connected to the central node in any triple according to the respective weight values, that is, weighted sampling is performed on the nodes connected to the central node in any triple, so that one of the sampled nodes is used as a positive sample node in any triple; that is, the computer device can select a node from the weighted nodes connected to the central node in any triple as a positive sample node in any triple. The method for determining the weight value of the edge involved in any two connected nodes is as described below and will not be repeated here.
[0068] It can be understood that in the process of weighted sampling, nodes connected by edges with larger weight values (i.e., nodes with larger weight values) are more likely to be sampled. In other words, after multiple samplings, nodes with larger weight values are sampled more times.
[0069] Correspondingly, the method of determining the negative sample node in any triplet may refer to: randomly selecting a node from the nodes that are not connected to the central node in any triplet as the negative sample node in any triplet.
[0070] S204, training and optimizing the vector representation model according to the initial representation vector of each node in each triple, and using the optimized vector representation model to determine the target representation vector of each node in the account relationship network.
[0071] The vector representation model may include, but is not limited to: Graphsage (Graph Sample Aggregate, a graph neural network (GNN) that generates a target representation vector of a central node by learning a function that aggregates neighboring nodes), DeepWalk (an algorithm for learning node embedding), and GCN (Graph Convolutional Network), etc.; this application does not limit this. It is understandable that the vector representation model is an unsupervised model, so the computer device can use the vector representation model for unsupervised training to determine the target representation vector of each node in the account relationship network.
[0072] S205 , clustering the accounts corresponding to the nodes based on the target representation vectors of the nodes, and identifying abnormal accounts among the multiple accounts participating in the clustering according to the clustering results.
[0073] It should be noted that when the computer device clusters the accounts corresponding to each node based on the target representation vector of each node, it can use the K-means clustering algorithm to cluster the target representation vector of each node to achieve clustering of the accounts corresponding to each node; it can also use the mean shift clustering algorithm to cluster the target representation vector of each node to achieve clustering of the accounts corresponding to each node; it can also use the hierarchical clustering algorithm to cluster the target representation vector of each node to achieve clustering of the accounts corresponding to each node, and so on; the present application does not limit the specific implementation method of clustering.
[0074] The embodiment of the present application can obtain multiple account attributes of each account to be identified, that is, it can simultaneously refer to account attributes of more dimensions, and support the expansion of account attributes. Accordingly, each account attribute of each account can be feature extracted to obtain multiple target account features; and multiple target account features are used to construct a more accurate account relationship network, in which a node in the account relationship network records the target account feature of an account, and two nodes connected by an edge record the same target account feature. Then, multiple triples can be sampled from the account relationship network, so that according to the initial representation vector of each node in each triple, the vector representation model is trained and optimized, and the optimized vector representation model is used to determine a more accurate target representation vector of each node in the account relationship network, so as to improve the accuracy of the representation vector of each node, and then the target representation vector of each node is used to better represent the account corresponding to the corresponding node; accordingly, based on the target representation vector of each node, the account corresponding to each node can be clustered to obtain a more accurate clustering result, and the abnormal account can be identified in the multiple accounts participating in the clustering according to the clustering result, so as to improve the recall rate of the abnormal account without affecting the recognition accuracy.
[0075] See also Figure 4 , is a flowchart of another account identification method provided in an embodiment of the present application. The account identification method can be executed by the computer device (terminal or server) mentioned above; or, the account identification method can be executed by the terminal and the server together. For the sake of convenience, the following description will be based on the example of a computer device executing the account identification method; please refer to Figure 4 , the account identification method may include the following steps S401-S409:
[0076] S401, obtaining multiple account attributes of each account to be identified.
[0077] S402, extracting features from each account attribute of each account to obtain a plurality of initial account features, where one account attribute corresponds to one or more initial account features.
[0078] Specifically, for any account attribute, the computer device can determine the feature extraction method corresponding to any account attribute according to the attribute type of any account attribute; wherein, when the attribute type is a text type, the feature extraction method includes at least one of the following: an original string extraction method, a word segmentation extraction method and a keyword extraction method, and the original string extraction method is used to indicate: any account attribute is used as the corresponding initial account feature, the word segmentation extraction method is used to indicate: any account attribute is segmented and the word segmentation result is used as the corresponding initial account feature, and the keyword extraction method is used to indicate: all extracted keywords are used as the corresponding initial account features; it should be understood that the above-mentioned word segmentation extraction method includes but is not limited to a bigram extraction method and a trigram extraction method, that is, the word segmentation result can be a bigram or a trigram, and the present application does not limit this; and when extracting keywords of any account attribute, the computer device can extract valuable phrases in any account attribute through algorithms such as hot word discovery or new word discovery to extract keywords of any account attribute.
[0079] It should be noted that the above-mentioned text types may include short text types and long text types; wherein, the short text type refers to the text type of short text, short text refers to text whose text length is less than a preset length threshold, and the text length may refer to the number of characters included in the text; correspondingly, the long text type refers to the text type of long text, and long text refers to text whose text length is greater than or equal to the preset length threshold; it should be understood that the preset text length may be set according to experience or according to actual needs, and this application does not limit this. Optionally, when the above-mentioned attribute type is a short text type, the feature extraction method may include at least one of the following: an original string extraction method and a word segmentation extraction method; when the above-mentioned attribute type is a long text type, the feature extraction method may include: a keyword extraction method.
[0080] Correspondingly, when the above-mentioned attribute type is an image type, the feature extraction method includes an image quantization extraction method, and the image quantization extraction method is used to indicate: image quantization processing is performed on any account attribute, and the image quantization processing result is used as the corresponding initial account feature; wherein, the above-mentioned image quantization processing can refer to feature extraction for an image, and when the image quantization processing result is a hash algorithm string, image quantization processing on any account attribute can refer to: extracting the hash algorithm string of any account attribute to obtain the image quantization processing result corresponding to any account attribute (i.e., the extraction result of the hash algorithm string); wherein, the above-mentioned hash algorithm can refer to an ahash (mean hash) algorithm, or a dhash (difference hash) algorithm, etc.; this application does not limit this. Alternatively, when the above-mentioned attribute type is a numeric type or a character type, the feature extraction method includes an original string extraction method, such as an encrypted mobile phone number, an encrypted md5 (Message-Digest Algorithm 5, fifth edition of the information digest algorithm) code (i.e., a feature code obtained by mathematically transforming the original information through the md5 algorithm), etc.
[0081] Furthermore, the computer device may perform feature extraction on any account attribute according to the feature extraction method corresponding to any account attribute, and obtain one or more initial account features corresponding to the any account attribute. Specifically, when the feature extraction method corresponding to any account attribute includes the original string extraction method, the computer device may use the any account attribute as the corresponding initial account feature; when the feature extraction method corresponding to any account attribute includes the word segmentation extraction method, the computer device may perform word segmentation processing on the any account attribute, and use the word segmentation processing result as the corresponding initial account feature; when the feature extraction method corresponding to any account attribute includes the keyword extraction method, the computer device may perform keyword extraction processing on the any account attribute, and use the extracted keyword as the corresponding initial account feature; when the feature extraction method corresponding to any account attribute includes the image quantization extraction method, the computer device may perform image quantization processing on the any account attribute, and use the image quantization processing result as the corresponding initial account feature.
[0082] S403, calculating a feature value of each initial account feature according to the feature type of each initial account feature.
[0083] In a specific implementation, the computer device may group multiple initial account features according to the feature type of each initial account feature to obtain multiple account feature groups; wherein one account feature group corresponds to one feature type, that is, the feature type of each initial account feature in the same account feature group is the same. Furthermore, for any initial account feature, the computer device may determine the feature value of any initial account feature according to the proportion of any initial account feature in the corresponding account feature group; wherein the feature value of any initial account feature is negatively correlated with the corresponding proportion. Based on this, the computer device may use formula 1.1 to calculate the feature value of the initial account feature f as follows:
[0084] T f = -log p 2 Formula 1.1
[0085] When calculating the characteristic value of the initial account characteristic f, p is the proportion of the initial account characteristic f in the corresponding account characteristic group.
[0086] It should be noted that the feature type of any initial account feature can be used to indicate: the attribute source of the account attribute corresponding to any initial account feature, that is, the attribute source of the account attribute corresponding to the initial account features of the same feature type is the same, and the attribute source of any account attribute includes but is not limited to nicknames, avatars, personal signatures, login addresses, and reported items, etc. For example, if account attribute A is the nickname "I love you very much", and account attribute B is the nickname "Honest Chinese Medicine", then the attribute source of account attribute A and the attribute source of account attribute B are both nicknames. Correspondingly, the feature type of any initial account feature can also be used to indicate: the attribute source of the account attribute corresponding to any initial account feature and the corresponding feature extraction method, that is, the attribute source and the corresponding feature extraction method of the initial account features of the same feature type are the same; it can also be used to indicate: the attribute type of the account attribute corresponding to any initial account feature, that is, the attribute type of the account attribute corresponding to the initial account features of the same feature type is the same, and so on; this application does not limit this.
[0087] For example, assuming that the account attributes corresponding to the initial account features of the same feature type have the same attribute source, and the feature extraction method corresponding to the initial account features of the same feature type is the same, and the nickname is used as an example for explanation; assuming that the nicknames of each account are short texts, and account attribute A is the nickname "I am an honest Chinese medicine practitioner", and account attribute B is the nickname "Honest Chinese medicine practitioner", then the computer device can perform feature extraction on account attribute A and account attribute B according to the original string extraction method and the word segmentation extraction method respectively, and obtain one or more initial account features of account attribute A and one or more initial account features of account attribute B. Assume that the word segmentation extraction method refers to the binary word extraction method, and the nickname binary word of account attribute A can include "I am" and "Integrity-Traditional Chinese Medicine", and the nickname binary word of account attribute B can include "Integrity-Traditional Chinese Medicine", that is, the initial account feature of account attribute A includes the original nickname string "I am an honest Chinese medicine practitioner", the nickname binary word "I am", and the nickname binary word "Integrity-Traditional Chinese Medicine", and the initial account feature of account attribute B includes the original nickname string "Integrity-Traditional Chinese Medicine practitioner" and the nickname binary word "Integrity-Traditional Chinese Medicine practitioner", then the feature types of the original nickname string "I am an honest Chinese medicine practitioner" and the original nickname string "Integrity-Traditional Chinese Medicine practitioner" in each initial account feature are the same, and the feature types of the nickname binary word "I am" and the nickname binary word "Integrity-Traditional Chinese Medicine practitioner" are the same. In this case, the computer device can divide the above-mentioned initial account features into two account feature groups, one account feature group includes an original nickname string "I am an honest Chinese medicine practitioner" and an original nickname string "Integrity-Traditional Chinese Medicine practitioner", and the other account feature group includes a nickname binary word "I am" and two nickname binary words "Integrity-Traditional Chinese Medicine practitioner".
[0088] Based on this, we will take the example of determining the feature value of the nickname dichotomy word "integrity-traditional Chinese medicine" as an example to further illustrate. Assume that the number of nickname dichotomy words "integrity-traditional Chinese medicine" in the corresponding account feature group is C 诚信-中医 , and the number of all features in the corresponding account feature group (i.e., the number of all nickname dichotomies) is C 昵称-bigram , then the computer device can use formula 1.2 to calculate the feature value of the nickname dichotomy "integrity-traditional Chinese medicine":
[0089]
[0090] S404, detecting common features from a plurality of initial account features according to feature values of each initial account feature, wherein the common features refer to initial account features held by at least K accounts, where K is a positive integer.
[0091] It should be understood that the common feature has a large universality and can be held by at least K accounts, and the universality of an initial account feature is positively correlated with the corresponding proportion, that is, the greater the proportion of any initial account feature in the corresponding account feature group, the greater the universality of any initial account feature; correspondingly, since the characteristic value of an initial account feature is negatively correlated with the corresponding proportion, the greater the characteristic value of any initial account feature, the smaller the universality of any initial account feature, that is, the computer device can use the initial account feature with a smaller characteristic value as a common feature. Among them, the common feature can be a province, a city, etc., and this application does not limit this.
[0092] In a specific implementation, the computer device may arrange each initial account feature in descending order according to the feature value of each initial account feature to obtain a feature sequence; each initial account feature located after the target arrangement position in the feature sequence is taken as a common feature.
[0093] In one embodiment, before the computer device uses each initial account feature located after the target arrangement position in the feature sequence as a common feature, it can determine the target arrangement position according to a preset elimination threshold; wherein the preset elimination threshold can be a specific elimination number or a elimination percentage, which is not limited in this application. It should be understood that when the preset elimination threshold refers to the elimination number, the target arrangement position can refer to: the first position before the reciprocal preset elimination threshold position in the feature sequence; when the preset elimination threshold refers to the elimination percentage, the target arrangement position can refer to: the first position between the reciprocal target number positions in the feature sequence, and the target number is equal to the multiplication result between the number of initial account features in the feature sequence and the elimination percentage.
[0094] For example, assuming that the preset elimination threshold refers to the elimination number, and the elimination number is 20, then the target arrangement position may refer to: the first position before the 20th to last position in the feature sequence, that is, the 21st to last position in the feature sequence, and the computer device may take each initial account feature after the 21st to last position in the feature sequence as a common feature; for another example, assuming that the preset elimination threshold refers to the elimination percentage, and the elimination percentage is 10%, and the number of initial account features in the feature sequence is 100, then the target arrangement position may refer to: the first position before the 10th to last position in the feature sequence, then the computer device may take each initial account feature after the 11th to last position in the feature sequence as a common feature, that is, the initial account features at the tail end of the feature sequence may be taken as common features, and so on.
[0095] In another embodiment, before treating each initial account feature located after the target arrangement position in the feature sequence as a common feature, the computer device may determine the target arrangement position according to a preset feature threshold; specifically, the computer device may use the position of the initial account feature corresponding to the minimum feature value in the feature sequence whose feature value is greater than the preset feature threshold as the target arrangement position. In this case, treating each initial account feature located after the target arrangement position in the feature sequence as a common feature may mean treating each initial account feature in the feature sequence whose feature value is less than or equal to the preset feature threshold as a common feature. The initial feature threshold may be set according to experience or according to actual needs, and this application does not limit this.
[0096] S405, taking the remaining initial account features except the common features among the multiple initial account features as target account features.
[0097] It is understandable that the computer device can use each initial account feature with a larger characteristic value among multiple initial account features as the target account feature. In other words, the computer device can select representative target account features according to the proportion of the initial account features in the group; further, the selected target account features can be used to associate accounts together to form an account relationship network.
[0098] S406, using multiple target account features to construct an account relationship network, where a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature.
[0099] In a specific implementation, a computer device may generate multiple nodes, and in accordance with the principle that one node records the target account features of one account, multiple target account features are recorded in multiple nodes according to the accounts corresponding to each target account feature; for any two nodes, the same target account features recorded by any two nodes are detected; and when the same target account features are detected, the number of features of the same target account features is counted; if the number of features is greater than the number threshold, an edge is used to connect any two nodes; wherein the same target account features may also be referred to as similar account features. In other words, a computer device may use accounts as nodes and the common target account features between accounts as edges to form an account relationship network; based on this, a node in the account relationship network corresponds to an account, and the same target account features exist between two nodes connected by an edge in the account relationship network.
[0100] The quantity threshold may be set based on experience or actual needs, and this application does not limit this. For example, when the quantity threshold is 1, the computer device uses an edge to connect any two nodes if and only if the number of features corresponding to any two nodes is greater than 1.
[0101] Furthermore, the number of features can be represented by M, where M is a positive integer; then if the number of features is greater than the number threshold, the specific implementation method of using an edge to connect any two nodes may include: if the number of features is greater than the number threshold, then in the account attributes corresponding to the M target account features, the number of different account attributes is counted; if the number of different account attributes counted is greater than the attribute number threshold, then an edge is used to connect the above any two nodes.
[0102] Among them, the attribute quantity threshold can be set according to experience or according to actual needs, and this application does not limit this; illustratively, assuming that the attribute quantity threshold is 1, then an edge is used to connect any two of the above nodes when and only when the number of different account attributes counted is greater than 1, that is, if the two accounts have only one type of common target account feature, the nodes corresponding to the two accounts are not connected together; for example, assuming that the same target account features between the account indicated by node A and the account indicated by node B include the nickname (i.e., the original nickname string) "Integrity Chinese Medicine" and the nickname dichotomy "Integrity-Chinese Medicine", and the account attributes corresponding to these two target account features are both the nickname "Integrity Chinese Medicine", then the account indicated by node A and the account indicated by node B have only one type of common target account feature, that is, the number of different account attributes counted is equal to 1, then an edge is not used to connect the two nodes, that is, the accounts indicated by the two nodes are not associated. It should be understood that this situation is excluded mainly because as the amount of data increases, many accounts will have the same target account characteristics by chance. Therefore, a threshold value of the number of attributes can be set to constrain the connection relationship between any two nodes to improve the accuracy of the account relationship network.
[0103] It should be noted that the number of account attributes corresponding to the M target account features is greater than or equal to M; for any target account feature among the M target account features, the computer device may use the two account attributes corresponding to any target account feature under the account indicated by any two nodes as the account attributes corresponding to the any target account feature, or may use the one account attribute corresponding to any target account feature under the account indicated by any two nodes as the account attribute corresponding to the any target account feature, and so on; this application does not limit this.
[0104] Specifically, when the computer device counts the number of different account attributes in the account attributes corresponding to M target account features, it can initialize the account attribute set, and the account attribute set after initialization is an empty set; traverse each target account feature in the M target account features, and determine whether the account attribute corresponding to the currently traversed target account feature is in the account attribute set; if not, add the account attribute corresponding to the currently traversed target account feature to the account attribute set; if yes, continue to traverse the M target account features; then, after each target account feature in the M target account features has been traversed, count the number of account attributes in the account attribute set to obtain the number of different account attributes.
[0105] Furthermore, for any two nodes, the computer device can also determine the account association degree corresponding to the any two nodes (i.e., the account association degree of the accounts corresponding to the two nodes), and use the account association degree as the weight value of the edge connecting the any two nodes; that is, the account relationship network can be a weighted network. Specifically, if the any two nodes have the same target account characteristics, then when determining the account association degree corresponding to the any two nodes, the computer device can calculate the account association degree corresponding to the any two nodes based on the feature values of the same target account characteristics recorded by the any two nodes; if the any two nodes do not have the same target account characteristics, it can be determined that the account association degree corresponding to the any two nodes is zero.
[0106] In one embodiment, when calculating the account association degree corresponding to any two nodes based on the feature values of the same target account feature recorded by the any two nodes, the computer device may sum the feature values of the same target account feature recorded by the any two nodes and use the summation result as the account association degree corresponding to the any two nodes.
[0107] In another embodiment, when calculating the account association degree corresponding to any two nodes based on the feature values of the same target account feature recorded by the any two nodes, the computer device can determine one or more target feature values from the feature values of the same target account feature recorded by the any two nodes, and any target feature value refers to: the largest feature value among the feature values of at least one target account feature under the corresponding account attribute, and the target account feature corresponding to any target feature value refers to: the target account feature with the largest feature value among at least one target account feature under the corresponding account attribute; then, the above one or more target feature values are summed, and the summation result is used as the account association degree corresponding to the any two nodes. Specifically, the computer device can use formula 1.3 to calculate the account association degree corresponding to node A and node B (that is, the account association degree of the account corresponding to node A and the account corresponding to node B) as follows:
[0108] R AB =max(T 昵称bigram-诚信-中医 ,T 昵称原串-诚信中医 …)+T 头像hash +…+T 其他 Formula 1.3
[0109] Among them, the account attribute corresponding to the nickname dichotomy "Integrity-Traditional Chinese Medicine" and the account attribute corresponding to the original nickname string "Integrity Traditional Chinese Medicine" are all the nickname "Integrity Traditional Chinese Medicine". That is to say, these identical target account characteristics correspond to the same account attribute. Then the computer device can select a maximum eigenvalue from the eigenvalues of these identical target account characteristics as the target eigenvalue, and use the obtained target eigenvalue to perform a sum operation to obtain the account association degree corresponding to the two nodes, that is, the account association degree of the corresponding two accounts.
[0110] S407, sampling multiple triplets from the account relationship network, a triplet includes a central node, a positive sample node and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node.
[0111] S408, training and optimizing the vector representation model according to the initial representation vector of each node in each triple, and using the optimized vector representation model to determine the target representation vector of each node in the account relationship network.
[0112] It should be noted that before training and optimizing the vector representation model based on the initial representation vectors of each node in each triple, for any node in the account relationship network, the computer device can determine the initial representation vector of any node according to a preset representation method.
[0113] In one embodiment, the computer device can perform vector representation processing on the target account attributes of the accounts corresponding to each node in the account relationship network, respectively, to obtain the initial representation vector of each node; wherein, the target account attribute of any account can refer to any account attribute among the multiple account attributes of any account, and the attribute sources of the target account attributes of the accounts corresponding to each node are the same, such as account nicknames or avatars, etc. For example, assuming that the target account attribute is an account nickname, the computer device can perform vector representation processing on the nicknames of the accounts corresponding to each node in the account relationship network, to obtain the initial representation vector of each node; in this case, the computer device can use w2v (word2vec, a word vector construction model) or Bert (Bidirectional Encoder Representation from Transformers, a pre-trained language representation model), etc., to perform vector representation processing on the nicknames of the accounts corresponding to each node in the account relationship network, to obtain the initial representation vector of each node, and this application does not limit the specific implementation method of the vector representation processing.
[0114] In another embodiment, the computer device may randomly generate a vector for each node in the account relationship network according to the target dimension, and use each randomly generated vector as the initial characterization vector of the corresponding node. Specifically, the computer device may determine the target dimension and randomly generate a vector according to the target dimension; or, the computer device may randomly generate a vector within a preset vector range, and the dimension of any vector within the preset vector range is the target dimension. The target dimension and the preset vector range may be set according to experience or according to actual needs, and this application does not limit this.
[0115] In another embodiment, the computer device can perform a random walk in the account relationship network with any node as the starting point according to a preset walk length for any node in the account relationship network, and generate an initial characterization vector of any node according to the random walk result. It should be noted that the computer device can perform a random walk in the account relationship network with a preset walk length and equal probability; it can also perform a random walk in the account relationship network according to the weight value of each edge in the account relationship network and the preset walk length, that is, in each random walk, the computer device can randomly select a node from each node connected to the current node according to the weight value of the edge corresponding to each node connected to the current node, and so on; it can be understood that in the process of weighted sampling, the nodes connected to the edges with larger weight values (i.e., the nodes with larger weight values) are more likely to be sampled. In other words, after multiple samplings, the nodes with larger weight values are sampled more times.
[0116] The preset walking length may be set according to experience or according to actual needs, and this application does not limit this. It should be understood that the dimension of the initial characterization vector is equal to the preset walking length, such as 5 or 6.
[0117] For example, Figure 3 As shown, assuming that the account relationship network includes nodes A, B, C, D, E, F, G, H, I, J and K, and assuming that the preset walking length is 5, then when node A is taken as the starting point, since node A is connected to nodes E, E and H, the computer device can randomly walk to any node among nodes E, E and H; assuming that it randomly walks to node D, since node D is connected to nodes A, B, C and G, the computer device can randomly walk from node D to any node among nodes A, B, C and G, and so on; and assuming that it walks to nodes B, C and D in turn, the walking route starting from node A can be nodes A, D, B, C and D, and assuming that the corresponding values of nodes A, B, C and D are 1, 2, 3 and 4 respectively, the initial characterization vector of node A generated by the computer device according to the random walk result can be (1, 4, 2, 3, 4).
[0118] Furthermore, when the computer device trains and optimizes the vector representation model based on the initial representation vector of each node in each triple, the computer device may call the vector representation model to perform vector representation on each node in each triple based on the initial representation vector of each node in each triple, and obtain the intermediate representation vector of each node in each triple; and calculate the model loss value of the vector representation model based on the intermediate representation vector of each node in each triple and the vector difference conditions that the intermediate representation vector of each node in a single triple needs to satisfy; accordingly, the computer device may train and optimize the vector representation model in the direction of reducing the model loss value.
[0119] It should be understood that the representation vectors of connected nodes (i.e., nodes with better relationships) are closer, and the representation vectors of unconnected nodes (nodes with no relationship) are farther. Based on this, the similarity (i.e., distance) between the intermediate representation vector of the central node in any triple and the intermediate representation vector of the corresponding negative sample node is greater than the similarity between the intermediate representation vector of the central node and the intermediate representation vector of the corresponding positive sample node; then correspondingly, the vector difference condition that the intermediate representation vectors of each node in a single triple need to satisfy is: the similarity between the intermediate representation vector of the central node and the intermediate representation vector of the corresponding negative sample node, and the difference between the similarity between the intermediate representation vector of the central node and the intermediate representation vector of the corresponding positive sample node is greater than the preset distance threshold. Among them, the preset distance threshold can be set according to experience or according to actual needs, and this application does not limit this.
[0120] S409, clustering the accounts corresponding to the nodes based on the target representation vectors of the nodes, and identifying abnormal accounts among the multiple accounts participating in the clustering according to the clustering results.
[0121] Wherein, the clustering result includes multiple categories of accounts (i.e., multiple clustering account groups), and one category of accounts includes at least one account; accordingly, when identifying abnormal accounts among multiple accounts participating in the clustering according to the clustering result, the computer device can divide any category of accounts in the clustering result into multiple account pairs; and calculate the account association degree of each account pair according to the characteristic value of the target account feature of each account in each account pair in the multiple account pairs; wherein, if the two accounts in any account pair have the same target account feature, the account association degree of any account pair is obtained by summing the characteristic values of the same target account feature; then, the account association degree of each account pair in the multiple account pairs can be summed to obtain the total account association degree of any category of accounts; accordingly, if the total account association degree of any category of accounts is greater than or equal to the preset association threshold, each account in the any category of accounts is identified as an abnormal account. It should be understood that if the two accounts in any account pair do not have the same target account feature, the account association degree of any account pair is zero. It should be noted that when calculating the account association degree of any account pair, the characteristic values of the same target account characteristics involved in any account pair can be summed up, and the result of the summation can be used as the account association degree of any account pair; or, the target characteristic values among the characteristic values of the same target account characteristics involved in any account pair can be summed up, and the result of the summation can be used as the account association degree of any account pair.
[0122] Specifically, when summing up the account associations of each account pair in the plurality of account pairs to obtain the total account association of any type of account, the computer device may use formula 1.4 to calculate the total account association of any type of account (i.e., the score of the association of any type of account) as follows:
[0123]
[0124] Among them, R ij It refers to the account association degree between the ith account and the jth account in any category of accounts, n is the number of accounts in any category of accounts, and n is a positive integer.
[0125] It should be noted that the above preset correlation threshold can be set according to experience or according to actual needs, and this application does not limit this. t To represent the preset association threshold, the computer device can take S>=S t All kinds of accounts are provided for manual review, and S can be reviewed according to their own circumstances in actual situations. t Make adjustments. For example, Figure 5a A schematic diagram showing the list data provided for manual review. The data of the same type of accounts can form an independent batch with a unique batch number. The main feature is the most important feature of the data in this batch, the secondary feature is other reference features, and the number of accounts is the number of accounts in this batch. For example, click Figure 5a You can view the details of any batch in the batch, including but not limited to the basic information of the account and the information of the complaint, etc. Figure 5b shown.
[0126] It should be understood that through unsupervised learning based on graph neural networks, many types of accounts with "dispersed but interrelated features" can be mined; for example, Figure 5c As shown, the account's signature, avatar, and nickname do not have the same logo, but are related to each other and have similar geographical locations. Figure 5a-5c The contents of each display interface are only exemplarily shown, and this application does not limit this; Figure 5a It may also not include information such as batch status, or Figure 5b Real-name information etc. may also not be included.
[0127] Furthermore, in order to better illustrate the effect of the account identification method proposed in this application, this application also conducted result statistics in actual applications, and the statistical results obtained indicated that the average daily number of abnormal clustered account groups (an abnormal clustered account group is a type of abnormal account) discovered by the account identification method proposed in this application increased by 50%, and the accuracy of the abnormal clustered account group was 60%, where an abnormal clustered account group refers to a type of account whose total account correlation is greater than or equal to the preset correlation threshold; it can be seen that compared with the clustering of account attributes based on a single dimension, the account identification method proposed in this application can simultaneously refer to information on account attributes in more dimensions, thereby improving the recall rate of abnormal clustered account groups without affecting the accuracy, that is, improving the recall rate of abnormal accounts. In addition, the statistical results also indicate that the high-quality accounts (any high-quality account refers to a type of account with similar account attributes in P dimensions, P is a positive integer greater than the preset dimensional threshold) obtained through the account identification method proposed in this application is 93%, which is consistent with the prior art; and the non-high-quality accounts (any non-high-quality account refers to a type of account with similar account attributes in Q dimensions, Q is a positive integer less than or equal to the preset dimensional threshold) obtained have an improvement of 25%. That is to say, the account identification method proposed in this application can dig out more abnormal clustered account groups, that is, it can dig out more types of abnormal accounts.
[0128] In the embodiment of the present application, after obtaining multiple account attributes of each account to be identified, feature extraction is performed on each account attribute of each account to obtain multiple initial account features, and common features are detected from the multiple initial account features based on the feature values of each initial account feature, and the remaining initial account features other than the common features in the multiple initial account features are all used as target account features to build a more accurate account relationship network, and multiple triples are sampled from the account relationship network; then, the vector representation model can be trained and optimized based on the initial representation vector of each node in each triple, and the target representation vector of each node in the account relationship network can be determined using the optimized vector representation model to improve the accuracy of the representation vector of each node, so as to cluster the accounts corresponding to each node based on the target representation vector with higher accuracy of each node, and identify abnormal accounts from the multiple accounts participating in the clustering based on the more accurate clustering results, thereby improving the recall rate of abnormal accounts without affecting the recognition accuracy.
[0129] Based on the description of the above-mentioned account identification method, the present application also proposes an account identification device, which can be a computer program (including program code) running in a computer device. Figure 2 or Figure 4 The account identification method shown in Figure 6, the account identification device can run the following units:
[0130] The acquisition unit 601 is used to acquire multiple account attributes of each account to be identified, and perform feature extraction on each account attribute of each account to obtain multiple target account features;
[0131] Processing unit 602 is configured to construct an account relationship network using the multiple target account features, wherein a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature;
[0132] The processing unit 602 is further configured to sample a plurality of triplets from the account relationship network, wherein a triplet includes a central node, a positive sample node, and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node;
[0133] The processing unit 602 is further used to train and optimize the vector representation model according to the initial representation vector of each node in each triple, and use the optimized vector representation model to determine the target representation vector of each node in the account relationship network;
[0134] The processing unit 602 is further used to cluster the accounts corresponding to the nodes based on the target characterization vectors of the nodes, and identify abnormal accounts among the multiple accounts participating in the clustering according to the clustering results.
[0135] In one implementation, when the acquisition unit 601 extracts features from each account attribute of each account to obtain multiple target account features, it can be specifically used to:
[0136] Extracting features from each account attribute of each account to obtain a plurality of initial account features, wherein one account attribute corresponds to one or more initial account features;
[0137] Calculating the feature value of each initial account feature according to the feature type of each initial account feature;
[0138] Detecting common features from the multiple initial account features according to the feature values of the initial account features, wherein the common features refer to: initial account features held by at least K accounts, where K is a positive integer;
[0139] The remaining initial account features among the multiple initial account features except the common features are all used as target account features.
[0140] In another implementation, when the acquisition unit 601 extracts features from each account attribute of each account to obtain multiple initial account features, it can be specifically used to:
[0141] For any account attribute, determining a feature extraction method corresponding to the any account attribute according to the attribute type of the any account attribute;
[0142] According to the feature extraction method, feature extraction is performed on any of the account attributes to obtain one or more initial account features corresponding to the any of the account attributes.
[0143] In another embodiment, when the attribute type is a text type, the feature extraction method includes at least one of the following: an original string extraction method, a word segmentation extraction method, and a keyword extraction method; the original string extraction method is used to indicate: taking any of the account attributes as the corresponding initial account feature, the word segmentation extraction method is used to indicate: performing word segmentation processing on any of the account attributes, and taking the word segmentation processing result as the corresponding initial account feature, and the keyword extraction method is used to indicate: taking all the extracted keywords as the corresponding initial account feature;
[0144] When the attribute type is an image type, the feature extraction method includes an image quantization extraction method, and the image quantization extraction method is used to indicate: performing image quantization processing on any of the account attributes, and using the image quantization processing result as the corresponding initial account feature;
[0145] When the attribute type is a numeric type or a character type, the feature extraction method includes an original string extraction method.
[0146] In another implementation manner, when the acquisition unit 601 calculates the feature value of each initial account feature according to the feature type of each initial account feature, it can be specifically used to:
[0147] Grouping the multiple initial account features according to the feature type of each initial account feature to obtain multiple account feature groups, where one account feature group corresponds to one feature type;
[0148] For any initial account feature, a feature value of the initial account feature is determined according to the proportion of the initial account feature in the corresponding account feature group; wherein the feature value of the initial account feature is negatively correlated with the corresponding proportion.
[0149] In another implementation manner, when the acquisition unit 601 detects the common features from the multiple initial account features according to the feature values of the initial account features, it can be specifically used to:
[0150] Arranging the initial account features in descending order according to the feature values of the initial account features to obtain a feature sequence;
[0151] Each initial account feature located after the target arrangement position in the feature sequence is taken as a common feature.
[0152] In another implementation, when the processing unit 602 uses the multiple target account features to construct the account relationship network, it can be specifically used to:
[0153] Generate multiple nodes, and based on the principle that one node records the target account feature of one account, record the multiple target account features into the multiple nodes according to the accounts corresponding to the respective target account features;
[0154] For any two nodes, detecting the same target account features recorded by the any two nodes; and when the same target account features are detected, counting the number of features of the same target account features;
[0155] If the number of features is greater than the number threshold, an edge is used to connect any two nodes.
[0156] In another implementation manner, the number of features is represented by M, where M is a positive integer; when the processing unit 602 connects the two arbitrary nodes with an edge if the number of features is greater than the number threshold, it can be specifically used to:
[0157] If the number of features is greater than the number threshold, counting the number of different account attributes among the account attributes corresponding to the M target account features;
[0158] If the number of different account attributes counted is greater than the attribute quantity threshold, an edge is used to connect any two nodes.
[0159] In another implementation manner, when the processing unit 602 counts the number of different account attributes among the account attributes corresponding to the M target account features, it may be specifically configured to:
[0160] Initialize the account attribute set. The account attribute set after initialization is an empty set.
[0161] Traversing each target account feature in the M target account features, and determining whether the account attribute corresponding to the currently traversed target account feature is in the account attribute set;
[0162] If not, then add the account attribute corresponding to the currently traversed target account feature to the account attribute set; if it is, then continue to traverse the M target account features;
[0163] After each of the M target account features is traversed, the number of account attributes in the account attribute set is counted to obtain the number of different account attributes.
[0164] In another implementation, the processing unit 602 may also be used to:
[0165] Performing vector representation processing on the target account attributes of the accounts corresponding to the respective nodes in the account relationship network to obtain initial representation vectors of the respective nodes;
[0166] Alternatively, a vector is randomly generated for each node in the account relationship network according to the target dimension, and each randomly generated vector is used as an initial representation vector of the corresponding node;
[0167] Alternatively, for any node in the account relationship network, taking any node as a starting point, a random walk is performed in the account relationship network according to a preset walk length, and an initial characterization vector of the any node is generated according to the random walk result.
[0168] In another implementation, when the processing unit 602 trains and optimizes the vector representation model according to the initial representation vectors of each node in each triple, it can be specifically used to:
[0169] The vector representation model is called to perform vector representation on each node in each triple according to the initial representation vector of each node in each triple, so as to obtain an intermediate representation vector of each node in each triple;
[0170] Calculating a model loss value of the vector representation model according to the intermediate representation vectors of each node in each triple and the vector difference conditions that the intermediate representation vectors of each node in a single triple need to satisfy;
[0171] The vector representation model is trained and optimized in a direction of reducing the loss value of the model.
[0172] In another implementation, the clustering result includes multiple types of accounts, and one type of account includes at least one account; when the processing unit 602 identifies an abnormal account from multiple accounts participating in the clustering according to the clustering result, it can be specifically used to:
[0173] For any type of account in the clustering result, divide the any type of account into a plurality of account pairs;
[0174] According to the characteristic value of the target account characteristic of each account in each account pair of the multiple account pairs, the account association degree of each account pair is calculated respectively; wherein, if two accounts in any account pair have the same target account characteristic, the account association degree of any account pair is obtained by summing the characteristic values of the same target account characteristic;
[0175] performing a sum operation on the account association degree of each account pair in the plurality of account pairs to obtain a total account association degree of any one type of account;
[0176] If the total account association degree of any one type of account is greater than or equal to a preset association threshold, each account in any one type of account is identified as an abnormal account.
[0177] According to one embodiment of the present application, Figure 2 or Figure 4 Each step involved in the method shown can be Figure 6 The account identification device shown in FIG. 1 is executed by each unit in the account identification device shown in FIG. Figure 2 The step S201 shown in FIG. 1 can be performed by Figure 6 The acquisition unit 601 shown in FIG. 1 is executed, and steps S202-S205 can be performed by Figure 6 The processing unit 602 shown in is executed. For example, Figure 4 The steps S401-S405 shown in FIG. 4 can be performed by Figure 6 The acquisition unit 601 shown in FIG. 6 is executed, and steps S406-S409 can be performed by Figure 6 Processing unit 602 is shown executing, among other things.
[0178] According to another embodiment of the present application, Figure 6 The various units in the account identification device shown can be individually or completely combined into one or several other units to form, or one (some) of the units can be further divided into multiple functionally smaller units to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the account identification device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0179] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory medium (RAM), a read-only memory medium (ROM), and other processing elements and storage elements. Figure 2 or Figure 4A computer program (including program code) for each step involved in the corresponding method shown in Figure 6 The account identification device shown in and the account identification method of the embodiment of the present application are implemented. The computer program can be recorded on, for example, a computer storage medium, and loaded into the above-mentioned computing device through the computer storage medium and run therein.
[0180] The embodiment of the present application can obtain multiple account attributes of each account to be identified, that is, it can simultaneously refer to account attributes of more dimensions, and support the expansion of account attributes. Accordingly, each account attribute of each account can be feature extracted to obtain multiple target account features; and multiple target account features are used to construct a more accurate account relationship network, in which a node in the account relationship network records the target account feature of an account, and two nodes connected by an edge record the same target account feature. Then, multiple triples can be sampled from the account relationship network, so that according to the initial representation vector of each node in each triple, the vector representation model is trained and optimized, and the optimized vector representation model is used to determine a more accurate target representation vector of each node in the account relationship network, so as to improve the accuracy of the representation vector of each node, and then the target representation vector of each node is used to better represent the account corresponding to the corresponding node; accordingly, based on the target representation vector of each node, the account corresponding to each node can be clustered to obtain a more accurate clustering result, and the abnormal account can be identified in the multiple accounts participating in the clustering according to the clustering result, so as to improve the recall rate of the abnormal account without affecting the recognition accuracy.
[0181] Based on the description of the above method embodiment and device embodiment, the present application embodiment also provides a computer device. Figure 7 The computer device at least includes a processor 701, an input interface 702, an output interface 703, and a computer storage medium 704. The processor 701, the input interface 702, the output interface 703, and the computer storage medium 704 in the computer device may be connected via a bus or other means.
[0182] The computer storage medium 704 can be stored in the memory of the computer device, and the computer storage medium 704 is used to store a computer program, and the computer program includes program instructions, and the processor 701 is used to execute the program instructions stored in the computer storage medium 704. The processor 701 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; in one embodiment, the processor 701 described in the embodiment of the present application can be used to perform a series of account identification, specifically including: obtaining multiple account attributes of each account to be identified, and extracting features of each account attribute of each account to obtain multiple target account features; using the multiple target account features to construct an account relationship network, in which a node in the account relationship network records the target account feature of an account, and two nodes connected by an edge record the same target account Features; sampling multiple triplets from the account relationship network, a triplet includes a central node, a positive sample node and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node; according to the initial representation vector of each node in each triplet, the vector representation model is trained and optimized, and the optimized vector representation model is used to determine the target representation vector of each node in the account relationship network; based on the target representation vector of each node, the accounts corresponding to the nodes are clustered, and abnormal accounts are identified among the multiple accounts participating in the clustering according to the clustering results, and so on.
[0183] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer storage medium provides a storage space, which stores the operating system of the computer device. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor. In one embodiment, the processor can load and execute one or more instructions stored in the computer storage medium to implement the above-mentioned related Figure 2 or Figure 4 The various method steps in the embodiment of the account identification method are shown.
[0184] The embodiment of the present application can obtain multiple account attributes of each account to be identified, that is, it can simultaneously refer to account attributes of more dimensions, and support the expansion of account attributes. Accordingly, each account attribute of each account can be feature extracted to obtain multiple target account features; and multiple target account features are used to construct a more accurate account relationship network, in which a node in the account relationship network records the target account feature of an account, and two nodes connected by an edge record the same target account feature. Then, multiple triples can be sampled from the account relationship network, so that according to the initial representation vector of each node in each triple, the vector representation model is trained and optimized, and the optimized vector representation model is used to determine a more accurate target representation vector of each node in the account relationship network, so as to improve the accuracy of the representation vector of each node, and then the target representation vector of each node is used to better represent the account corresponding to the corresponding node; accordingly, based on the target representation vector of each node, the account corresponding to each node can be clustered to obtain a more accurate clustering result, and the abnormal account can be identified in the multiple accounts participating in the clustering according to the clustering result, so as to improve the recall rate of the abnormal account without affecting the recognition accuracy.
[0185] It should be noted that according to one aspect of the present application, a computer program product or a computer program is also provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer storage medium. The processor of the computer device reads the computer instructions from the computer storage medium, and the processor executes the computer instructions, so that the computer device performs the above Figure 2 or Figure 4 The method provided in various optional aspects of the account identification method embodiment shown.
[0186] Furthermore, it should be understood that what is disclosed above is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. An account identification method, characterized in that: include: Acquire multiple account attributes of each account to be identified, and perform feature extraction on each account attribute of each account to obtain multiple target account features; An account relationship network is constructed using the multiple target account features, wherein a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature; A plurality of triplets are sampled from the account relationship network, wherein a triplet includes a central node, a positive sample node, and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node; According to the initial representation vectors of each node in each triple, the vector representation model is trained and optimized, and the target representation vectors of each node in the account relationship network are determined by using the optimized vector representation model; Based on the target characterization vector of each node, the accounts corresponding to each node are clustered, and abnormal accounts are identified among the multiple accounts participating in the clustering according to the clustering result.
2. The method according to claim 1, characterized in that: The feature extraction is performed on each account attribute of each account to obtain multiple target account features, including: Extracting features from each account attribute of each account to obtain a plurality of initial account features, wherein one account attribute corresponds to one or more initial account features; Calculating the feature value of each initial account feature according to the feature type of each initial account feature; Detecting common features from the multiple initial account features according to the feature values of the initial account features, wherein the common features refer to: initial account features held by at least K accounts, where K is a positive integer; The remaining initial account features among the multiple initial account features except the common features are all used as target account features.
3. The method according to claim 2, characterized in that The feature extraction is performed on each account attribute of each account to obtain a plurality of initial account features, including: For any account attribute, determining a feature extraction method corresponding to the any account attribute according to the attribute type of the any account attribute; According to the feature extraction method, feature extraction is performed on any of the account attributes to obtain one or more initial account features corresponding to the any of the account attributes.
4. The method according to claim 3, characterized in that When the attribute type is a text type, the feature extraction method includes at least one of the following: an original string extraction method, a word segmentation extraction method, and a keyword extraction method; the original string extraction method is used to indicate: taking any of the account attributes as the corresponding initial account feature, the word segmentation extraction method is used to indicate: performing word segmentation processing on any of the account attributes, and taking the word segmentation processing result as the corresponding initial account feature, and the keyword extraction method is used to indicate: taking all the extracted keywords as the corresponding initial account feature; When the attribute type is an image type, the feature extraction method includes an image quantization extraction method, and the image quantization extraction method is used to indicate: performing image quantization processing on any of the account attributes, and using the image quantization processing result as the corresponding initial account feature; When the attribute type is a numeric type or a character type, the feature extraction method includes an original string extraction method.
5. The method according to any one of claims 2 to 4, characterized in that: The calculating the feature value of each initial account feature according to the feature type of each initial account feature includes: Grouping the multiple initial account features according to the feature type of each initial account feature to obtain multiple account feature groups, where one account feature group corresponds to one feature type; For any initial account feature, a feature value of the initial account feature is determined according to the proportion of the initial account feature in the corresponding account feature group; wherein the feature value of the initial account feature is negatively correlated with the corresponding proportion.
6. The method according to any one of claims 2 to 4, characterized in that: The detecting common features from the multiple initial account features according to the feature values of the initial account features includes: Arranging the initial account features in descending order according to the feature values of the initial account features to obtain a feature sequence; Each initial account feature located after the target arrangement position in the feature sequence is taken as a common feature.
7. The method according to any one of claims 1 to 4, wherein the step of constructing an account relationship network using the multiple target account characteristics comprises: Generate multiple nodes, and based on the principle that one node records the target account feature of one account, record the multiple target account features into the multiple nodes according to the accounts corresponding to the respective target account features; For any two nodes, detecting the same target account features recorded by the any two nodes; When the same target account features are detected, the number of features of the same target account features is counted; If the number of features is greater than the number threshold, an edge is used to connect the any two nodes.
8. The method according to claim 7, characterized in that The number of features is represented by M, where M is a positive integer; if the number of features is greater than a quantity threshold, an edge is used to connect any two nodes, including: If the number of features is greater than the number threshold, counting the number of different account attributes among the account attributes corresponding to the M target account features; If the number of different account attributes counted is greater than the attribute quantity threshold, an edge is used to connect any two nodes.
9. The method according to claim 8, characterized in that The counting of the number of different account attributes among the account attributes corresponding to the M target account features includes: Initialize the account attribute set. The account attribute set after initialization is an empty set. Traversing each target account feature in the M target account features, and determining whether the account attribute corresponding to the currently traversed target account feature is in the account attribute set; If not, then add the account attribute corresponding to the currently traversed target account feature to the account attribute set; if it is, then continue to traverse the M target account features; After each of the M target account features is traversed, the number of account attributes in the account attribute set is counted to obtain the number of different account attributes.
10. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Performing vector representation processing on the target account attributes of the accounts corresponding to the respective nodes in the account relationship network to obtain initial representation vectors of the respective nodes; Alternatively, a vector is randomly generated for each node in the account relationship network according to the target dimension, and each randomly generated vector is used as an initial representation vector of the corresponding node; Alternatively, for any node in the account relationship network, taking any node as a starting point, a random walk is performed in the account relationship network according to a preset walk length, and an initial characterization vector of the any node is generated according to the random walk result.
11. The method according to any one of claims 1 to 4, characterized in that: The training and optimization of the vector representation model according to the initial representation vectors of each node in each triplet includes: The vector representation model is called to perform vector representation on each node in each triple according to the initial representation vector of each node in each triple, so as to obtain an intermediate representation vector of each node in each triple; Calculating a model loss value of the vector representation model according to the intermediate representation vectors of each node in each triple and the vector difference conditions that the intermediate representation vectors of each node in a single triple need to satisfy; The vector representation model is trained and optimized in a direction of reducing the loss value of the model.
12. The method according to any one of claims 1 to 4, characterized in that: The clustering result includes multiple categories of accounts, and one category of accounts includes at least one account; and identifying abnormal accounts from the multiple accounts participating in the clustering according to the clustering result includes: For any type of account in the clustering result, divide the any type of account into a plurality of account pairs; According to the characteristic value of the target account characteristic of each account in each account pair of the multiple account pairs, the account association degree of each account pair is calculated respectively; wherein, if two accounts in any account pair have the same target account characteristic, the account association degree of any account pair is obtained by summing the characteristic values of the same target account characteristic; performing a sum operation on the account association degree of each account pair in the plurality of account pairs to obtain a total account association degree of any one type of account; If the total account association degree of any one type of account is greater than or equal to a preset association threshold, each account in any one type of account is identified as an abnormal account.
13. An account identification device, characterized in that: include: an acquisition unit, configured to acquire multiple account attributes of each account to be identified, and perform feature extraction on each account attribute of each account to obtain multiple target account features; A processing unit, configured to construct an account relationship network using the multiple target account features, wherein a node in the account relationship network records a target account feature of an account, and two nodes connected by an edge record the same target account feature; The processing unit is further used to sample a plurality of triplets from the account relationship network, wherein a triplet includes a central node, a positive sample node, and a negative sample node; the central node is a node selected from the account relationship network, the positive sample node is a node connected to the central node, and the negative sample node is a node not connected to the central node; The processing unit is further used to train and optimize the vector representation model according to the initial representation vector of each node in each triple, and use the optimized vector representation model to determine the target representation vector of each node in the account relationship network; The processing unit is further used to cluster the accounts corresponding to the nodes based on the target characterization vectors of the nodes, and identify abnormal accounts among the multiple accounts participating in the clustering according to the clustering results.
14. A computer device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 12 is implemented.
15. A computer storage medium, characterized in that: The computer storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Risk user management method and device, computer equipment and storage medium
CN111489095A
Abnormal account identification method, system and device and readable storage medium
CN113254672A