Data processing method, device and equipment
By constructing graph structure data between users and using graph comparison learning, the problem of user relationship graph characterization is solved, and efficient graph characterization learning is achieved, suitable for large-scale data.
Patent Information
- Application Number
- CN202510097500.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
In risk prevention and control scenarios, the interaction between users is very important, but the prior art is difficult to effectively characterize the relationship diagram between users, resulting in high algorithm complexity and inconsistent point feature distribution with the community structure in the graph.
By obtaining graph structure data based on financial transaction data between different users, using the encoder to determine the graph coded data, and processing the module degree data through the preset graph walk algorithm. Finally, through graph comparison learning, positive sample pairs and negative sample pairs are constructed to integrate topological structures and node attributes and reduce algorithm overhead.
The graph characterization learning of perceived community structure is realized, which can effectively integrate the topological structure and node attributes in the graph, reduce algorithm overhead, and is suitable for large-scale graph structure data.
Smart Images

Figure CN119988685A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and in particular to a data processing method, device and equipment. Background Art
[0002] Commonly used user group mining processing is feature-based and focuses more on the similarity of user portraits. However, in risk prevention and control scenarios, the interactive relationship between users is very important. For example, in the bottom-line risk scenario, the same user group with a specified risk will use the same device to engage in risky behaviors (such as fraud, stealing user privacy data, etc.). Starting from the relationship between shared devices between users, the user group of users with specified risks can be well mined, but the relationship between users is difficult to characterize, and is more like a relationship graph. However, some scenarios have high feature dimensions, resulting in excessive algorithm complexity, and the point feature distribution and the community structure distribution in the graph are likely to be inconsistent. To this end, it is necessary to provide a better graph feature learning method, so that the topological structure and node attributes in the graph can be integrated, and at the same time, the technical solution for dimensionality reduction can be achieved to reduce the algorithm overhead. Summary of the invention
[0003] The purpose of the embodiments of this specification is to provide a better graph feature learning method, so as to integrate the topological structure and node attributes in the graph and achieve a technical solution for dimensionality reduction to reduce algorithm overhead.
[0004] In order to implement the above technical solution, the embodiments of this specification are implemented as follows: A data processing method provided in an embodiment of the present specification includes: obtaining graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent associations between different nodes. Based on the graph structure data, graph encoding data corresponding to the graph structure data is determined by an encoder. The graph structure data is processed based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data. Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph comparison learning.
[0005] A data processing device provided in an embodiment of the present specification includes: a graph data acquisition module, which acquires graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent associations between different nodes. An encoding module, which determines graph encoding data corresponding to the graph structure data through an encoder based on the graph structure data. A modularity determination module, which processes the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data. A data pair construction module, which constructs positive sample pairs and / or negative sample pairs for the graph structure data through graph comparison learning based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data.
[0006] A data processing device provided in an embodiment of the present specification comprises: a processor; and a memory arranged to store computer executable instructions, wherein the executable instructions, when executed, cause the processor to: obtain graph structure data constructed based on financial transaction data between different users, wherein the graph structure data comprises nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent associations between different nodes. Based on the graph structure data, graph encoding data corresponding to the graph structure data is determined by an encoder. The graph structure data is processed based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data. Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed by graph comparison learning.
[0007] The embodiments of this specification also provide a storage medium, which is used to store computer executable instructions, and the executable instructions implement the following process when executed by a processor: obtain graph structure data constructed based on financial transaction data between different users, the graph structure data includes nodes and edges, the nodes are constructed based on relevant information of accounts of different users, and the edges represent the association relationship between different nodes. Based on the graph structure data, determine the graph encoding data corresponding to the graph structure data through an encoder. Process the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data. Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, construct positive sample pairs and / or negative sample pairs for the graph structure data through graph comparison learning.
[0008] The embodiments of this specification also provide a computer program product, including a computer program, which implements the following process when executed by a processor: obtaining graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent associations between different nodes. Based on the graph structure data, graph encoding data corresponding to the graph structure data is determined by an encoder. The graph structure data is processed based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data. Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph comparison learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative labor. Figure 1 This is an embodiment of a data processing method of this specification; Figure 2 Another data processing method embodiment of this specification; Figure 3 A schematic diagram of a multiple random walk processing process in this specification; Figure 4 A schematic diagram of a data processing process of this specification; Figure 5 This is an embodiment of a data processing device of the present specification; Figure 6 This is an embodiment of a data processing device in this specification. DETAILED DESCRIPTION
[0010] The embodiments of this specification provide a data processing method, device and equipment.
[0011] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0012] The embodiments of this specification provide a graph representation learning mechanism that can perceive community structure. Crowd expansion based on seeds is a common requirement in many scenarios. For example, in advertising scenarios, a user group similar to the seed user can be found in a large user database through a specified algorithm. In risk prevention and control scenarios, there are also similar requirements, that is, based on a given user with a specified risk, other users who are closely connected with the user can be mined, thereby realizing the mining of user groups with specified risks. Common user group mining processing is based on features and pays more attention to the similarity of user portraits, but in risk prevention and control scenarios, the interactive relationship between users is very important. For example, in the bottom-line risk scenario, the same user group with a specified risk will use the same device to engage in risky behavior. Starting from the relationship between users sharing devices, the user group where the user with the specified risk is located can be well mined, but the relationship between users is difficult to characterize, and it is more like a relationship graph. In view of this feature, a corresponding solution has been developed, that is, the fact edge relationship (such as capital transactions, device sharing, social relationships, etc.) can be used to construct a graph, and the constructed relationship graph can be cut into communities with dense internal edges, and then each community is purified based on the users with the specified risk and the weight of the edge to obtain the final user group. The above method has been verified to be effective in multiple scenarios. However, it is not enough to rely solely on user features or factual interactions between users to mine user groups. Based only on feature similarity, the mined user groups lack factual interactions, have poor interpretability and are difficult to use in business. Using only factual interactions will ignore the similarities of user groups in some risk features, resulting in inaccurate or insufficient coverage of mined user groups. Therefore, based on the above processing methods, a user group mining mechanism based on feature similarity and interaction relationships is provided. Specifically, user groups are mined based on point features and edge relationships. However, some scenarios have high feature dimensions, which leads to excessive algorithm complexity, and the point feature distribution and the community structure distribution in the graph are likely to be inconsistent.
[0013] Graph feature learning can be performed through graph contrastive learning. The initial graph contrastive learning uses instance recognition as a pre-text task, aiming to learn representations that are invariant to different enhanced views of the graph. In applications, a graph contrastive learning method that does not require enhancement is proposed. Specifically, contrast information can be extracted from the graph structure data itself and corresponding pseudo labels can be constructed. Algorithms such as clustering can be used to mine potential clusters within graph features, and the clustering results can be used as the basis for positive / negative sample pairs. In addition, a random walk algorithm can be used to mine positive samples in the neighborhood of each node in the graph. The above-mentioned graph contrastive learning methods have achieved remarkable success in graph clustering tasks. However, the above-mentioned enhancement-based graph contrastive learning methods usually rely on data enhancement to create multiple views and use multiple encoders to obtain corresponding representations. However, when applied to large-scale graph structure data, the above methods bring computational challenges and scalability issues. In addition, feature-based graph contrastive learning methods that do not require enhancement require the use of high-cost algorithms such as clustering algorithms to construct pseudo labels, which brings challenges when expanding to large-scale datasets. In addition, pre-text tasks play an important role in making graph contrast learning methods better adapt to downstream tasks. The typical approach is to define instance-level "enhancements", and then define the problem as closely aligning the learned "enhancements" with the original data while separating them from other data. However, this type of pre-text task ignores the inherent structure of graph-structured data, which may cause semantic drift in downstream clustering tasks. Graph contrast learning methods that do not require enhancement can alleviate semantic drift by mining the information inherently carried by graph-structured data. However, as mentioned above, the clustering-based approach lacks scalability, and using a simple random walk algorithm to mine positive samples within a node neighborhood may easily lead to community semantic drift, especially when a node is at the edge of the community, the semantic drift will be more obvious. To this end, it is necessary to provide a graph representation learning method that perceives community structure, so that the topological structure and node attributes in the graph can be integrated, and at the same time, a technical solution can be achieved to reduce dimensionality to reduce algorithm overhead. For specific processing, please refer to the specific content in the following embodiments.
[0014] like Figure 1As shown, an embodiment of this specification provides a data processing method, and the execution subject of the method can be a terminal device or a server, etc., wherein the terminal device can be a mobile terminal device such as a mobile phone, a tablet computer, or a computer device such as a laptop or a desktop computer, or an IoT device (specifically such as a smart watch, a car-mounted device, etc.), etc., wherein the server can be an independent server, or a server cluster composed of multiple servers, etc., and the server can be a background server for financial services or online shopping services, or a background server for an application, etc. In this embodiment, the execution subject is taken as an example for detailed description. For the case where the execution subject is a terminal device, please refer to the following server situation processing, which will not be repeated here. The method can specifically include the following steps: In step S102, graph structure data constructed based on financial transaction data between different users is obtained, and the graph structure data includes nodes and edges. The nodes are constructed based on relevant information of accounts of different users, and the edges represent association relationships between different nodes.
[0015] Among them, financial transaction data can include multiple types, for example, financial transaction data can include the account numbers of both parties to the transaction, the transaction location, the transaction time, the transaction amount, the information of the goods traded, the delivery method of the goods (such as express delivery or self-collection, etc.), etc., which can be set according to the actual situation. The nodes in the graph structure data can be constructed based on the relevant information of different users' accounts, and the relevant information of the accounts can include multiple types, such as the account number, name, code and other identification information. The association relationship between different nodes can include multiple types, such as friend relationship, relationship between buyer and seller, information interaction relationship, etc.
[0016] During implementation, when a user performs a financial transaction business, the financial transaction data generated in the process of the user performing the financial transaction business can be recorded. When the financial transaction data of the financial transaction business is needed, the above-recorded financial transaction data can be obtained. In actual applications, the financial transaction data generated in a specified time period can also be selected from the recorded financial transaction data, or the financial transaction data generated in a specified time period can be obtained from a database used to store the financial transaction data generated in the process of performing the financial transaction business. In addition to obtaining the financial transaction data generated in the financial transaction business in the above-mentioned manner, the financial transaction data can also be obtained in a variety of other different ways, which can be set according to actual conditions.
[0017] Each financial transaction data can be analyzed, and based on the analysis results, the key information contained in the financial transaction data can be determined, which can include building a node based on the relevant information of the user's account, building a node based on the relevant information of different users' accounts, and obtaining the association relationship between different nodes. The edges between the corresponding nodes are constructed based on the obtained association relationship, and then the corresponding graph structure data can be constructed based on the obtained key information.
[0018] In step S104, based on the graph structure data, graph encoding data corresponding to the graph structure data is determined by an encoder.
[0019] Among them, the encoder can be an encoder directly obtained from a specified database (such as an autoencoder AE), or an encoder constructed and trained using sample data and a specified algorithm, etc. The encoder can be constructed by a specified algorithm, or it can be constructed by a neural network (such as a neural network in the figure, etc.), etc., which can be set specifically according to actual conditions.
[0020] In implementation, the graph structure data can be input into an encoder, and the nodes and edges in the graph structure data can be encoded separately through the algorithm and / or network in the encoder to obtain the encoding information of the nodes and edges in the graph structure data, and then the graph encoding data corresponding to the graph structure data can be determined.
[0021] In step S106, the graph structure data is processed based on a preset graph walking algorithm to obtain modularity data corresponding to the graph structure data.
[0022] Among them, the graph walk algorithm can be to walk on the graph structure data to obtain a sequence of walk nodes, and use the relationship between nodes through graph representation learning to obtain a one-dimensional representation of the node, and then use these one-dimensional representations to perform downstream tasks. The goal of the graph walk algorithm is to learn a low-dimensional representation of each node in the graph structure data. After obtaining these low-dimensional representations, these low-dimensional representations can be used to perform the next downstream tasks (such as node classification, etc.). The graph walk algorithm first referred to the Word2vec model of NLP. One of the algorithms of the Word2vec model is the Skip Gram algorithm, which predicts the context based on the central word, and then optimizes it through negative sampling. Combining the idea of the Word2vec model with graph structure data will result in a graph walk algorithm. Modularity data can be an indicator of the quality of community division in graph structure data. Specifically, it can be a measure of the closeness of nodes in the community. It can be used as an optimization target to discover high-quality communities. In practical applications, modularity information can be determined by modularity matrices, etc. The modularity is as follows ,if ,but
[0023] in, Q represents modularity, B is the modularity matrix, each element in the modularity matrix represents the modularity coefficient, P It is a one-hot matrix, indicating the community to which each node belongs, N represents the number of communities, and C represents the community.
[0024] In implementation, the graph walk algorithm can be used to start from a node in the graph structure data. Each step of the walk randomly selects an edge from the edge connected to the current node or selects a specified edge, and moves to the next node along the selected edge. The above process is repeated until the obtained node sequence cannot continue to go down or reaches the specified maximum length. After walking many times, multiple walk node sequences can be obtained. The way to select the next node in the random walk sequence through the graph walk algorithm can be uniformly randomly distributed, so that the nodes connected to the current node have the same probability of being selected, or it can be pre-specified. In the graph walk algorithm, multiple sequences are sampled for each node in the graph structure data. After obtaining these node sequences, the Word2vec model can be used for processing. Finally, the modularity data corresponding to the graph structure data can be obtained.
[0025] It should be noted that the processing of the above-mentioned step S104 and step S106 can be as described above, first executing the processing of step S104 and then executing the processing of step S106, or first executing the processing of step S106 and then executing the processing of step S104, or simultaneously executing the processing of step S104 and step S106. The specific setting can be based on actual conditions, and the embodiments of this specification do not limit this.
[0026] In step S108, based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph contrast learning.
[0027] In implementation, each element in the modularity matrix in the modularity data corresponding to the graph structure data represents the modularity coefficient. At the same time, the graph coding data corresponding to the graph structure data can be determined in the above manner. In this way, the graph coding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data can be used to calculate the similarity between different nodes in the graph structure data. The calculated similarities can be compared, and the comparison results can be used to determine whether the two nodes belong to a positive sample pair or a negative sample pair, so that nodes in the same community are used as positive sample pairs, and nodes in different communities are used as negative sample pairs. In the above manner, positive sample pairs and / or negative sample pairs for graph structure data can be constructed.
[0028] The embodiment of the present specification provides a data processing method, which obtains graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on the relevant information of the accounts of different users, and the edges represent the association relationship between different nodes. Then, based on the graph structure data, the graph encoding data corresponding to the graph structure data can be determined through an encoder, and then the graph structure data can be processed based on a preset graph walking algorithm to obtain modularity data corresponding to the graph structure data. Finally, based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, a graph comparison learning method for the graph structure data can be constructed. In this way, the community structure in the graph structure data is perceived through the modularity data corresponding to the graph structure data to generate positive and negative sample pairs. The cost is very small and it can be easily expanded to graph structure data with a node scale of hundreds of millions or more. In addition, since the nodes in the same community are used as positive sample pairs, and the nodes in different communities are used as negative sample pairs, and graph representation learning is performed based on contrastive learning, the influence of community semantic shift can be effectively eliminated. In addition, graph representation learning based on contrastive learning can integrate the topological structure and node attributes in the graph structure data, and achieve dimensionality reduction to reduce algorithm overhead.
[0029] In practical applications, the specific processing methods of the above step S104 can be varied. An optional processing method is provided below, such as Figure 2 As shown, the steps may specifically include the following steps S1042 and S1044.
[0030] In step S1042, the graph structure data is input into an encoder constructed based on a graph neural network to obtain initial graph encoding data corresponding to the graph structure data.
[0031] Among them, the encoder constructed by the graph neural network GNN may include a variety of different network layers. Different network layers may include different parameters. Through different network layers, the relevant information of nodes and edges in the graph structure data can be encoded and processed to obtain corresponding encoded data.
[0032] In implementation, an initial encoder can be constructed based on the graph neural network GNN, and a certain number of training samples can be obtained. The initial encoder is trained using the training samples to obtain a trained encoder, which can encode the relevant information of the nodes and edges in the graph structure data to obtain corresponding encoded data. The graph structure data obtained above can be input into the trained encoder, and the relevant information of the nodes and edges in the graph structure data can be encoded by the trained encoder to obtain the initial graph encoding data corresponding to the graph structure data.
[0033] In step S1044, the initial graph coding data is normalized to obtain graph coding data corresponding to the graph structure data.
[0034] Among them, the normalization processing may include multiple types, for example, the normalization processing may include normalization processing, standardization processing, etc., which can be specifically set according to actual conditions.
[0035] In practical applications, the commonly used modularity function only focuses on the first-order proximity in the network, which goes against people's intuition, that is, two nodes belong to the same community often because there is a high-order proximity between them (for example, they have many common neighbors). Therefore, a sampling strategy based on a two-stage random walk can be set to effectively perceive the high-order proximity in the network. Specifically, the specific processing method of the above step S106 can be varied. An optional processing method is provided below, which may specifically include the following: performing multiple random walk processing in the graph structure data based on a preset deep random walk algorithm, and determining the modularity data corresponding to the graph structure data based on the number of mutual visits between nodes in the graph structure data obtained from the multiple random walk processing.
[0036] Among them, the deep random walk algorithm can be a random walk on the graph structure data to obtain a walk node sequence, and through graph representation learning, the association relationship between nodes is used to obtain a one-dimensional representation of the node, and then these one-dimensional representations are used for downstream tasks. The goal of the deep random walk algorithm is to learn a low-dimensional representation of each node in the graph structure data. After obtaining these low-dimensional representations, these low-dimensional representations can be used to perform the next downstream tasks.
[0037] In implementation, the deep random walk algorithm can be used to start from a node in the graph structure data. Each step of the random walk randomly selects an edge from the edge connected to the current node, moves along the selected edge to the next node, and repeats the above process until the obtained node sequence cannot continue to go down or reaches the specified maximum length. After multiple random walks, multiple random walk node sequences can be obtained. The way to select the next node in the random walk sequence through the deep random walk algorithm can be uniformly randomly distributed, so for nodes connected to the current node, there is an equal probability of being selected. In the deep random walk algorithm, multiple sequences are sampled for each node in the graph structure data. After obtaining these node sequences, the Word2vec model can be used for processing. Finally, the modularity data corresponding to the graph structure data can be obtained.
[0038] In practical applications, the above-mentioned deep random walk algorithm based on the preset performs multiple random walk processing in the graph structure data, and determines the modularity data corresponding to the graph structure data based on the number of mutual visits between nodes in the graph structure data obtained from the multiple random walk processing. The specific processing method can be varied, and an optional processing method is provided below, which can specifically include the following steps A2 to A8.
[0039] In step A2, each node in the graph structure data is used as a starting node, and multiple random walk processes are performed starting from the starting node, and the average number of visits is determined based on the number of mutual visits between nodes in the graph structure data obtained from the multiple random walk processes.
[0040] In implementation, Figure 3 As shown, for any given set of n nodes, multiple deep random walk processes can be performed starting from each node, and a candidate positive sample set of the node can be constructed based on the node whose access times are greater than the mean access times. Based on this, for multiple nodes in the graph structure data, a node can be selected from multiple nodes in the graph structure data at will. For this node, this node can be used as the starting node, and multiple random walk processes can be performed starting from the starting node. For other nodes, the above-mentioned processes can be performed. For specific random walk processes, please refer to the aforementioned related content, which will not be repeated here. Through the above-mentioned multiple random walk processes, the number of mutual visits between nodes in the graph structure data obtained in the multiple random walk processes can be obtained, and the average value (i.e., mean) of the number of node accesses can be calculated. Finally, the mean of the number of accesses can be obtained.
[0041] In step A4, a candidate positive sample set of the current starting node is constructed based on nodes whose access times are greater than the average access times.
[0042] In step A6, multiple random walk processes are performed on each node in the candidate positive sample set of each starting node, and the similarity between different nodes in the candidate positive sample set is determined based on the number of mutual visits between nodes in the candidate positive sample set obtained from the multiple random walk processes.
[0043] In implementation, Figure 3As shown, a certain number of nodes can be selected from the candidate positive sample set to construct a subset (or batch), and then, a second multiple random walk process is performed on each node in the subset. Through the above second multiple random walk process, the number of mutual visits between nodes in the subset obtained in the second multiple random walk process can be obtained. The number of mutual visits between nodes can be normalized to obtain the normalized number of mutual visits between nodes. The normalized number of mutual visits between nodes can be used as the similarity between nodes, and then a similarity matrix can be generated, where each element in the similarity matrix can be
[0044] in, Representation Node v and nodes u The similarity between represents a sub-collection, Representation Node v and nodes u The number of visits between each other, Representation Node v and nodes Through the above method, the similarity between different nodes in the candidate positive sample set can be obtained.
[0045] In step A8, modularity data corresponding to the graph structure data is determined based on the similarities between different nodes in the candidate positive sample set.
[0046] In implementation, Figure 3 As shown, based on the above processing, the following mini-Batch modularity matrix can be generated
[0047] Through the modularity matrix in the form of mini-batch, modularity data corresponding to the graph structure data can be further generated.
[0048] In practical applications, the specific processing methods of the above step S108 can be varied. An optional processing method is provided below, which may specifically include the following contents: based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, construct positive sample pairs and / or negative sample pairs for the graph structure data through graph contrast learning of a graph contrast learning model. The graph contrast learning model is a model based on a machine learning network and combined with a contrast learning algorithm.
[0049] Among them, the machine learning network may include multiple types, for example, a graph neural network or a convolutional neural network, or a network built based on a Transformer module, etc., which can be specifically set according to actual conditions.
[0050] In practical applications, the above graph contrast learning model can be trained by the following methods: Figure 4 As shown, please refer to the processing of steps B2 to B8 below for details.
[0051] In step B2, a graph structure data sample constructed based on historical financial transaction data between different users is obtained.
[0052] The specific processing method of the above step B2 can refer to the specific content of the above step S102, which will not be repeated here.
[0053] In step B4, based on the graph structure data sample, a graph encoding data sample corresponding to the graph structure data sample is determined by an encoder.
[0054] The specific processing method of the above step B4 can refer to the specific content of the above step S104, which will not be repeated here.
[0055] In addition, in practical applications, the processing of the above step B4 can be implemented in the following way: input the graph structure data sample into the encoder constructed based on the graph neural network to obtain the initial graph encoding data sample corresponding to the graph structure data sample; normalize the initial graph encoding data sample to obtain the graph encoding data sample corresponding to the graph structure data sample. The above specific processing method can be found in the aforementioned related content and will not be repeated here.
[0056] In step B6, the graph structure data sample is processed based on a preset graph walking algorithm to obtain a modularity data sample corresponding to the graph structure data sample.
[0057] The specific processing method of the above step B6 can refer to the specific content of the above step S106, which will not be repeated here.
[0058] In addition, in practical applications, the processing of the above step B6 can be implemented in the following way: based on the preset deep random walk algorithm, multiple random walk processing is performed in the graph structure data sample, and based on the number of mutual visits between nodes in the graph structure data sample obtained from the multiple random walk processing, the modularity data sample corresponding to the graph structure data sample is determined. The above specific processing method can be found in the aforementioned related content and will not be repeated here.
[0059] In addition, if Figure 3As shown, the above-mentioned deep random walk algorithm based on the preset performs multiple random walk processing in the graph structure data sample, and determines the specific processing of the modularity data sample corresponding to the graph structure data sample based on the number of mutual visits between the nodes in the graph structure data sample obtained in the multiple random walk processing. It can also be implemented in the following way: take each node in the graph structure data sample as the starting sample node, and perform multiple random walk processing starting from the starting sample node, and determine the mean number of visits sample based on the number of mutual visits between the nodes in the graph structure data sample obtained in the multiple random walk processing; take the node with a visit number greater than the mean number of visits sample as the first candidate positive sample set of the current starting sample node; perform multiple random walk processing on each node in the first candidate positive sample set of each starting sample node, and determine the similarity between different nodes in the first candidate positive sample set based on the number of mutual visits between the nodes in the first candidate positive sample set obtained in the multiple random walk processing; determine the modularity data sample corresponding to the graph structure data sample based on the similarity between different nodes in the first candidate positive sample set. The above-mentioned specific processing method can be referred to the aforementioned related content, which will not be repeated here.
[0060] In step B8, based on the graph coding data samples corresponding to the graph structure data samples and the modularity data samples corresponding to the graph structure data samples, the graph contrast learning model is trained by a preset loss function to obtain a trained graph contrast learning model.
[0061] In implementation, based on the graph coding data samples corresponding to the graph structure data samples and the modularity data samples corresponding to the graph structure data samples, the graph contrast learning of the graph contrast learning model is used to construct positive sample pairs and / or negative sample pairs for the graph structure data samples, and the graph contrast learning model is trained accordingly. At the same time, during the model training process, the graph contrast learning model is trained using a preset loss function until the loss function converges, thereby obtaining a trained graph contrast learning model. The loss function may include multiple types, such as the cross entropy loss function, the Info NCE (Noise Contrastive Estimation) loss function, etc. In practical applications, the above loss function can also be constructed in the following way: given a training subset And the corresponding modularity matrix Afterwards, for any node v , whose positive and negative sample pairs are and
[0062] The corresponding loss function can be as follows
[0063] in, Indicates that for the node v The loss function is Representation Node v The representation of Representation Node u The representation of Representation Node The representation of Representation Node v With Node u The similarity between Representation Node v With Node The specific setting can also be made according to the actual situation, and this specification embodiment does not limit this.
[0064] In practical applications, the above loss functions include the SimCLR contrast loss function based on the Softmax function.
[0065] In implementation, the SimCLR contrast loss function based on the Softmax function can be as follows
[0066] in, Indicates that for the node v SimCLR contrast loss function based on Softmax function.
[0067] In practical applications, based on the above positive sample pairs and / or negative sample pairs for graph structure data, risk prevention and control processing can be performed in the following manner, which can specifically include the processing of steps C2 to C6.
[0068] In step C2, based on the positive sample pairs and / or negative sample pairs for the graph structure data and the association relationship between different nodes in the graph structure data, one or more different communities included in the graph structure data are determined.
[0069] In implementation, based on the positive sample pairs and / or negative sample pairs for the graph structure data, and the association relationship between different nodes in the graph structure data, the situation of the edges and nodes in the graph structure data can be analyzed to determine the clustering of the nodes in the graph structure data. Based on the clustering of the nodes, the nodes clustered together in the graph structure data can be divided into a community. In addition to being implemented in the above-mentioned manner, the above-mentioned specific processing method for community division of graph structure data can also be implemented in a variety of different ways, such as pre-setting a community discovery algorithm, such as the Louvain algorithm or the Bi-Louvain algorithm, etc., and the community discovery algorithm can be used to perform community division processing on the graph structure data to obtain one or more different communities, which can be specifically processed based on the calculation rules corresponding to the corresponding algorithm, which will not be repeated here.
[0070] In addition, after obtaining the community in the above manner, the reliability of the obtained community can be evaluated to evaluate the stability or importance of each node in the community, or to evaluate the stability or importance of the nodes in the community. If a node has a higher reliability in the community, the node is more important or more stable in the community. Specifically, a corresponding evaluation algorithm can be pre-set, and the evaluation algorithm can include multiple types. For example, the evaluation algorithm can include an algorithm for the similarity between different nodes, an algorithm constructed based on the situation of the edges connected to the nodes (including the weight of the edges, the number of edges, etc.), etc. It can be set specifically according to the actual situation, and this specification embodiment does not limit this. The evaluation algorithm can be used to evaluate the reliability of each community obtained.
[0071] In step C4, based on the nodes and edges included in each community and the information of the nodes with preset risks included in each community, a node group with preset risks is determined from the graph structure data.
[0072] Among them, the preset risks may include multiple types, such as fraud risk, privacy leakage risk, illegal financial activity risk, etc., which can be set according to actual conditions. The node with preset risk can be a node that has been predetermined to have preset risk. For example, an account is reported or complained by multiple users: the account has fraud risk. The business platform can determine whether the account has fraud risk by conducting a risk assessment on the account. If the account has fraud risk, the corresponding mark (such as setting a fraud label, etc.) can be set for the account to indicate that the account is a node with preset risk.
[0073] In implementation, a preset group mining mechanism can be used to mine node groups with preset risks in graph structure data. Specifically, based on the nodes and edges contained in each community and the information of nodes with preset risks contained in each community, the group mining mechanism can be used to mine node groups with preset risks in graph structure data to obtain at least one node group with preset risks. The group mining mechanism can include multiple types, such as an algorithm based on the eigenvalue of the Laplace matrix, and can be set according to actual conditions.
[0074] In step C6, risk prevention and control processing is performed on the accounts corresponding to different nodes in the above node group.
[0075] In implementation, risk control can be performed on accounts corresponding to different nodes in the above-mentioned node group in a variety of different ways. For example, accounts corresponding to different nodes in the above-mentioned node group can be frozen or the collection or payment authority of the account can be limited, etc. The specific settings can be made according to actual conditions.
[0076] The embodiment of the present specification provides a data processing method, which obtains graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on the relevant information of the accounts of different users, and the edges represent the association relationship between different nodes. Then, based on the graph structure data, the graph encoding data corresponding to the graph structure data can be determined through an encoder, and then the graph structure data can be processed based on a preset graph walking algorithm to obtain modularity data corresponding to the graph structure data. Finally, based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, a graph comparison learning method for the graph structure data can be constructed. In this way, the community structure in the graph structure data is perceived through the modularity data corresponding to the graph structure data to generate positive and negative sample pairs. The cost is very small and it can be easily expanded to graph structure data with a node scale of hundreds of millions or more. In addition, since the nodes in the same community are used as positive sample pairs, and the nodes in different communities are used as negative sample pairs, and graph representation learning is performed based on contrastive learning, the influence of community semantic shift can be effectively eliminated. In addition, graph representation learning based on contrastive learning can integrate the topological structure and node attributes in the graph structure data, and achieve dimensionality reduction to reduce algorithm overhead.
[0077] In addition, the above processing provides a new pre-training task for graph contrastive learning. By perceiving the community structure in the graph structure data, nodes in the same community are used as positive sample pairs, and nodes in different communities are used as negative sample pairs. It can be integrated into any pre-training framework for graph contrastive learning, thereby improving the scalability of graph contrastive learning.
[0078] The above is a data processing method provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Figure 5 shown.
[0079] The data processing device comprises: a graph data acquisition module 501, an encoding module 502, a modularity determination module 503 and a data pair construction module 504, wherein: A graph data acquisition module 501 acquires graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent associations between different nodes; An encoding module 502, based on the graph structure data, determines the graph encoding data corresponding to the graph structure data through an encoder; A modularity determination module 503 processes the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data; The data pair construction module 504 constructs positive sample pairs and / or negative sample pairs for the graph structure data through graph comparison learning based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data.
[0080] In the embodiment of this specification, the encoding module 502 includes: An encoding unit, inputting the graph structure data into an encoder constructed based on a graph neural network to obtain initial graph encoding data corresponding to the graph structure data; The normalization unit performs normalization processing on the initial graph coding data to obtain graph coding data corresponding to the graph structure data.
[0081] In the embodiment of the present specification, the modularity determination module 503 performs multiple random walk processes in the graph structure data based on a preset deep random walk algorithm, and determines the modularity data corresponding to the graph structure data based on the number of mutual visits between nodes in the graph structure data obtained from the multiple random walk processes.
[0082] In the embodiment of this specification, the modularity determination module 503 includes: a mean value determining unit, taking each node in the graph structure data as a starting node, and performing multiple random walk processes from the starting node, and determining a mean value of the number of visits based on the number of mutual visits between the nodes in the graph structure data obtained from the multiple random walk processes; A set construction unit, constructing a candidate positive sample set of the current starting node based on nodes whose access times are greater than the average of the access times; a similarity determination unit, performing multiple random walk processes on each node in the candidate positive sample set of each starting node, and determining the similarity between different nodes in the candidate positive sample set based on the number of mutual visits between the nodes in the candidate positive sample set obtained in the multiple random walk processes; The modularity determination unit determines the modularity data corresponding to the graph structure data based on the similarities between different nodes in the candidate positive sample set.
[0083] In an embodiment of the present specification, the data pair construction module 504, based on the graph coding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, constructs positive sample pairs and / or negative sample pairs for the graph structure data through graph contrast learning of a graph contrast learning model, wherein the graph contrast learning model is a model constructed based on a machine learning network and in combination with a contrast learning algorithm.
[0084] In the embodiment of this specification, the device further includes: The sample acquisition module acquires graph structure data samples constructed based on historical financial transaction data between different users; A sample encoding module, based on the graph structure data sample, determines a graph encoding data sample corresponding to the graph structure data sample through the encoder; A sample modularity determination module processes the graph structure data sample based on a preset graph walk algorithm to obtain a modularity data sample corresponding to the graph structure data sample; The training module performs model training on the graph contrast learning model through a preset loss function based on the graph coding data samples corresponding to the graph structure data samples and the modularity data samples corresponding to the graph structure data samples to obtain a trained graph contrast learning model.
[0085] In the embodiment of this specification, the loss function includes a SimCLR contrast loss function based on a Softmax function.
[0086] In the embodiment of this specification, the device further includes: A community division module, which determines one or more different communities contained in the graph structure data based on positive sample pairs and / or negative sample pairs for the graph structure data and association relationships between different nodes in the graph structure data; A group determination module, based on the nodes and edges contained in each community and the information of the nodes with preset risks contained in each community, determines a node group with preset risks from the graph structure data; The risk prevention and control module performs risk prevention and control on the accounts corresponding to different nodes in the node group.
[0087] The embodiment of the present specification provides a data processing device, which obtains graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on the relevant information of the accounts of different users, and the edges represent the association relationship between different nodes. Then, based on the graph structure data, the graph encoding data corresponding to the graph structure data can be determined through an encoder, and then the graph structure data can be processed based on a preset graph walking algorithm to obtain modularity data corresponding to the graph structure data. Finally, based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, a graph comparison learning method can be constructed for the graph structure data. In this way, the community structure in the graph structure data is perceived through the modularity data corresponding to the graph structure data to generate positive and negative sample pairs. The cost is very small and it can be easily expanded to graph structure data with a node scale of hundreds of millions or more. In addition, since the nodes in the same community are used as positive sample pairs, and the nodes in different communities are used as negative sample pairs, and graph representation learning is performed based on contrastive learning, the influence of community semantic shift can be effectively eliminated. In addition, graph representation learning based on contrastive learning can integrate the topological structure and node attributes in the graph structure data, and achieve dimensionality reduction to reduce algorithm overhead.
[0088] In addition, the above processing provides a new pre-training task for graph contrastive learning. By perceiving the community structure in the graph structure data, nodes in the same community are used as positive sample pairs, and nodes in different communities are used as negative sample pairs. It can be integrated into any pre-training framework for graph contrastive learning, thereby improving the scalability of graph contrastive learning.
[0089] The above is a data processing device provided in the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, such as Figure 6 shown.
[0090] The data processing device may provide a terminal device or a server, etc. for the above-mentioned embodiments.
[0091] The data processing device may have relatively large differences due to different configurations or performances, and may include one or more processors 601 and memory 602, and the memory 602 may store one or more storage applications or data. Among them, the memory 602 may be a short-term storage or a persistent storage. The application stored in the memory 602 may include one or more modules (not shown in the figure), and each module may include a series of computer executable instructions in the data processing device. Furthermore, the processor 601 may be configured to communicate with the memory 602 and execute a series of computer executable instructions in the memory 602 on the data processing device. The data processing device may also include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input and output interfaces 605, and one or more keyboards 606.
[0092] Specifically in this embodiment, the data processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer executable instructions in the data processing device, and the one or more programs are configured to be executed by one or more processors, including computer executable instructions for performing the following: Acquire graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent association relationships between different nodes; Based on the graph structure data, determining, by an encoder, graph encoding data corresponding to the graph structure data; Processing the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data; Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph contrast learning.
[0093] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the data processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0094] The embodiment of the present specification provides a data processing device, which obtains graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on the relevant information of the accounts of different users, and the edges represent the association relationship between different nodes. Then, based on the graph structure data, the graph encoding data corresponding to the graph structure data can be determined through an encoder, and then the graph structure data can be processed based on a preset graph walking algorithm to obtain modularity data corresponding to the graph structure data. Finally, based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, a graph comparison learning method can be constructed for the graph structure data. In this way, the community structure in the graph structure data is perceived through the modularity data corresponding to the graph structure data to generate positive and negative sample pairs. The cost is very small and it can be easily expanded to graph structure data with a node scale of hundreds of millions or more. In addition, since the nodes in the same community are used as positive sample pairs, and the nodes in different communities are used as negative sample pairs, and graph representation learning is performed based on contrastive learning, the influence of community semantic shift can be effectively eliminated. In addition, graph representation learning based on contrastive learning can integrate the topological structure and node attributes in the graph structure data, and achieve dimensionality reduction to reduce algorithm overhead.
[0095] Furthermore, based on the above Figures 1 to 4 In one embodiment, the present specification further provides a storage medium for storing computer executable instruction information. In a specific embodiment, the storage medium may be a USB flash drive, an optical disk, a hard disk, etc. When the computer executable instruction information stored in the storage medium is executed by the processor, the following process can be implemented: Acquire graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent association relationships between different nodes; Based on the graph structure data, determining, by an encoder, graph encoding data corresponding to the graph structure data; Processing the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data; Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph contrast learning.
[0096] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0097] The embodiment of the present specification provides a storage medium, which obtains graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent the association relationship between different nodes. Then, based on the graph structure data, the graph encoding data corresponding to the graph structure data can be determined through an encoder, and then the graph structure data can be processed based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data. Finally, based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, a graph comparison learning method for the graph structure data can be constructed. Positive sample pairs and / or negative sample pairs. In this way, the community structure in the graph structure data is perceived through the modularity data corresponding to the graph structure data to generate positive and negative sample pairs. The cost is very low and it can be easily expanded to graph structure data with a node scale of hundreds of millions or more. In addition, since the nodes in the same community are used as positive sample pairs, and the nodes in different communities are used as negative sample pairs, and graph representation learning is performed based on contrastive learning, the influence of community semantic shift can be effectively eliminated. In addition, graph representation learning based on contrastive learning can integrate the topological structure and node attributes in the graph structure data, thereby achieving dimensionality reduction to reduce algorithm overhead.
[0098] Furthermore, based on the above Figures 1 to 4 In one or more embodiments of the present specification, a computer program product is provided, including a computer program. When the computer program in the computer program product is executed by a processor, the following process can be implemented: Acquire graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent association relationships between different nodes; Based on the graph structure data, determining, by an encoder, graph encoding data corresponding to the graph structure data; Processing the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data; Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph contrast learning.
[0099] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0100] The embodiments of the present specification provide a computer program product, which obtains graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent the association relationship between different nodes. Then, based on the graph structure data, graph encoding data corresponding to the graph structure data can be determined through an encoder, and then the graph structure data can be processed based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data. Finally, based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, a graph comparison learning method can be constructed for the graph structure data. In this way, the community structure in the graph structure data is perceived by the modularity data corresponding to the graph structure data to generate positive and negative sample pairs. The cost is very low and it can be easily expanded to graph structure data with a node scale of hundreds of millions or more. In addition, since the nodes in the same community are used as positive sample pairs, and the nodes in different communities are used as negative sample pairs, and graph representation learning is performed based on contrastive learning, the influence of community semantic shift can be effectively eliminated. In addition, graph representation learning based on contrastive learning can integrate the topological structure and node attributes in the graph structure data, thereby achieving dimensionality reduction to reduce algorithm overhead.
[0101] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0102] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0103] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0104] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0105] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0106] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, one or more embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0107] The embodiments of this specification are described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable fraud case serial and parallel device to produce a machine, so that the instructions executed by the processor of the computer or other programmable fraud case serial and parallel device generate instructions for implementing the processes in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0108] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable fraud case serial and parallel device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0109] These computer program instructions may also be loaded onto a computer or other programmable device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0110] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0111] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0112] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0113] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0114] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0115] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0116] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0117] The above description is only an embodiment of this specification and is not intended to limit this document. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this specification.
Claims
1. A data processing method, the method comprising: Acquire graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent association relationships between different nodes; Based on the graph structure data, determining, by an encoder, graph encoding data corresponding to the graph structure data; Processing the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data; Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph contrast learning.
2. The method according to claim 1, wherein determining the graph encoding data corresponding to the graph structure data by an encoder based on the graph structure data comprises: Inputting the graph structure data into an encoder constructed based on a graph neural network to obtain initial graph encoding data corresponding to the graph structure data; The initial graph coding data is normalized to obtain graph coding data corresponding to the graph structure data.
3. According to the method of claim 2, the processing of the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data comprises: Based on a preset deep random walk algorithm, multiple random walk processes are performed in the graph structure data, and based on the number of mutual visits between nodes in the graph structure data obtained from the multiple random walk processes, the modularity data corresponding to the graph structure data is determined.
4. The method according to claim 3, wherein the method performs multiple random walk processes in the graph structure data based on a preset deep random walk algorithm, and determines the modularity data corresponding to the graph structure data based on the number of mutual visits between nodes in the graph structure data obtained in the multiple random walk processes, including: Taking each node in the graph structure data as a starting node, performing multiple random walk processes from the starting node, and determining a mean number of visits based on the number of mutual visits between nodes in the graph structure data obtained from the multiple random walk processes; Constructing a candidate positive sample set of the current starting node based on nodes whose access times are greater than the average of the access times; Performing multiple random walk processes on each node in the candidate positive sample set of each starting node, and determining the similarity between different nodes in the candidate positive sample set based on the number of mutual visits between the nodes in the candidate positive sample set obtained from the multiple random walk processes; Based on the similarities between different nodes in the candidate positive sample set, modularity data corresponding to the graph structure data is determined.
5. According to the method of claim 1, constructing positive sample pairs and / or negative sample pairs for the graph structure data through graph contrast learning based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, comprising: Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph contrast learning of a graph contrast learning model, and the graph contrast learning model is a model constructed based on a machine learning network and in combination with a contrast learning algorithm.
6. The method according to claim 5, further comprising: Obtain graph structure data samples based on historical financial transaction data between different users; Based on the graph structure data sample, determining, by the encoder, a graph encoding data sample corresponding to the graph structure data sample; Processing the graph structure data sample based on a preset graph walk algorithm to obtain a modularity data sample corresponding to the graph structure data sample; Based on the graph encoding data samples corresponding to the graph structure data samples and the modularity data samples corresponding to the graph structure data samples, the graph contrast learning model is trained by a preset loss function to obtain a trained graph contrast learning model.
7. The method according to claim 6, wherein the loss function comprises a SimCLR contrast loss function based on a Softmax function.
8. The method according to claim 1, further comprising: Determine one or more different communities contained in the graph structure data based on positive sample pairs and / or negative sample pairs for the graph structure data and association relationships between different nodes in the graph structure data; Based on the nodes and edges contained in each community and the information of the nodes with preset risks contained in each community, determining a node group with preset risks from the graph structure data; Perform risk prevention and control on accounts corresponding to different nodes in the node group.
9. A data processing device, comprising: A graph data acquisition module, which acquires graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent associations between different nodes; An encoding module, based on the graph structure data, determines the graph encoding data corresponding to the graph structure data through an encoder; A modularity determination module processes the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data; A data pair construction module constructs positive sample pairs and / or negative sample pairs for the graph structure data through graph comparison learning based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data.
10. A data processing device, comprising: processor; as well as a memory arranged to store computer executable instructions which, when executed, cause the processor to: Acquire graph structure data constructed based on financial transaction data between different users, wherein the graph structure data includes nodes and edges, wherein the nodes are constructed based on relevant information of accounts of different users, and the edges represent association relationships between different nodes; Based on the graph structure data, determining, by an encoder, graph encoding data corresponding to the graph structure data; Processing the graph structure data based on a preset graph walk algorithm to obtain modularity data corresponding to the graph structure data; Based on the graph encoding data corresponding to the graph structure data and the modularity data corresponding to the graph structure data, positive sample pairs and / or negative sample pairs for the graph structure data are constructed through graph contrast learning.