Data processing method and device, equipment, storage medium and computer program product
In the distributed training of graph neural network model, the resource interaction graph is reasonably divided by load balancing coefficients and weight parameters, and the problems of load imbalance and communication overhead are solved, and the model training efficiency is improved.
Patent Information
- Application Number
- CN202510323415.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-29
AI Technical Summary
In the distributed training of existing graph neural network models, the division scheme of resource interaction graph ignores the internal connection between data, resulting in unbalanced load during training and large communication overhead, which affects the model training efficiency.
By obtaining the load balancing coefficient and weight parameters of the resource interaction graph, reasonably divide the resource interaction graph, allocate shared nodes to each classification result, reduce cross-computer communication, and determine the allocation result with load balancing coefficient and weight parameters, ensuring that closely connected nodes are in the same classification result and reduce cross-trainer access.
Load balancing is realized and communication access overhead is reduced, model training efficiency is improved, model training is ensured that the model mainly accesses local data during training, and cross-machine communication is reduced.
Smart Images

Figure CN120387498A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of artificial intelligence, and particularly relates to a data processing method, apparatus, device, storage medium, and computer program product. Background Art
[0002] Distributed training of a graph neural network model is a technique that distributes the training task of the graph neural network model to multiple computing account nodes for parallel computing. To complete the distributed training of the graph neural network model, it is often necessary to partition the resource interaction graph formed by the training data into multiple subgraphs based on a graph partitioning strategy, so that each computing account node processes the data in the corresponding subgraph to train the model.
[0003] Current resource interaction graph partitioning schemes often aim to minimize the cut edges for subgraph partitioning, or partition the account nodes and edges with similar timestamps in the resource interaction graph into the same subgraph, or partition the changing part of the resource interaction graph (i.e., the incremental graph of the resource interaction graph) into subgraphs. These partitioning schemes ignore the internal connections between the data in the resource interaction graph, which will lead to load imbalance of subgraphs during training, large differences in training time for each computing account node, and the sharing nodes being concentrated in a few subgraphs, resulting in frequent access by computing account nodes to the data in other computing account nodes and large communication overhead, seriously affecting the training efficiency of the model.
[0004] Therefore, how to achieve a reasonable partitioning of the resource interaction graph to balance the load, reduce the communication access overhead, and improve the model training efficiency has become a technical problem that urgently needs to be solved at present. Summary of the Invention
[0005] Embodiments of this application provide a data processing method, apparatus, device, storage medium, and computer program product, which can solve the problem of how to achieve a reasonable partitioning of the resource interaction graph to balance the load, reduce the communication access overhead, and improve the model training efficiency.
[0006] In a first aspect, an embodiment of this application provides a data processing method, the method including:
[0007] Obtain a resource interaction graph corresponding to the data to be trained, where the data to be trained is used to indicate the resource interaction situation between multiple accounts, the resource interaction graph includes multiple account nodes, where there is at least one resource interaction event between at least two account nodes, and the multiple account nodes include shared nodes, and the resource interaction graph corresponds to multiple classification results, and each classification result includes shared nodes;
[0008] Obtain the load balancing coefficients corresponding to multiple classification results respectively, and obtain the weight parameters corresponding to the account nodes in multiple classification results respectively. The load balancing coefficient represents the load situation corresponding to the classification result. Among them, the nth weight parameter corresponding to the ith account node is used to indicate that the ith account node is related to the neighbor nodes in the nth classification result, and there is a resource interaction event between the neighbor nodes and the ith account in the resource interaction graph. i and n are positive integers;
[0009] Based on multiple load balancing coefficients and multiple weight parameters, determine the allocation results corresponding to multiple classification results respectively. The allocation results include the account nodes allocated to the classification results and the resource interaction events related to the account nodes;
[0010] Train the sample model based on the allocation results corresponding to multiple classification results respectively to obtain a target model. The target model is used to identify abnormal resource interaction events or abnormal account nodes.
[0011] In some embodiments, obtaining the load balancing coefficients corresponding to multiple classification results respectively includes:
[0012] Obtain the first node quantity of the account nodes in the nth classification result and the first event quantity of the resource interaction events;
[0013] Obtain the timestamp parameters and event quantities of the resource interaction events corresponding to multiple classification results except the nth classification result respectively, and the target timestamp of the resource interaction events in the nth classification result. The timestamp parameters include a first timestamp and a second timestamp, and the event quantities include a second event quantity and a third event quantity;
[0014] Determine the first ratio result between the result quantity of multiple classification results and the total quantity of account nodes in the resource interaction graph, and based on the second ratio result between the first ratio result and the first node quantity, determine the node load coefficient of the nth classification result;
[0015] Determine the event load coefficient of the nth classification result according to the difference between the second event quantity, the third event quantity and the first event quantity;
[0016] Determine the time load coefficient of the nth classification result according to the difference between the first timestamp, the second timestamp and the target timestamp;
[0017] Determine the load balancing coefficient corresponding to the nth classification result according to the product result of the node load coefficient, the event load coefficient and the time load coefficient.
[0018] In some embodiments, obtaining the weight parameters corresponding to the account nodes in multiple classification results respectively includes:
[0019] Obtain the target neighbor nodes of the i-th account node in the n-th classification result, the third timestamp of the i-th account node, and the fourth timestamp corresponding to the resource interaction graph. There is a resource interaction event between the target neighbor nodes and the i-th account node in the resource interaction graph, and the existing resource interaction events are assigned to the n-th classification result. The fifth timestamp of the target neighbor nodes is included in the n-th classification result;
[0020] Determine the difference result between the third timestamp and the fifth timestamp, and determine the weight parameter corresponding to the i-th account node in the n-th classification result based on the product result of the difference result and the reciprocal of the fourth timestamp.
[0021] In some embodiments, the resource interaction graph includes an unassigned x-th resource interaction event. The x-th resource interaction event corresponds to a first account node and a second account node. The first account node corresponds to a first weight parameter in the n-th classification result, and the second account node corresponds to a second weight parameter in the n-th classification result. x is a positive integer;
[0022] Determine the allocation results corresponding to multiple classification results based on multiple load balancing coefficients and multiple weight parameters, including:
[0023] In the case where the x-th resource interaction event does not meet the preset conditions, determine the sum result between the first weight parameter and the second weight parameter. The preset conditions include that the first account node and the second account node are not shared nodes, and the first account node and the second account node are assigned to the same classification result, or one of the first account node and the second account node is a shared node and the other is assigned to a classification result;
[0024] Determine the n-th evaluation result corresponding to the n-th classification result according to the product result of the sum result and the load balancing coefficient corresponding to the n-th classification result;
[0025] In the case where the n-th evaluation result meets the first condition, assign the x-th resource interaction event to the n-th classification result to obtain the allocation result corresponding to the n-th classification result. The first condition is that the n-th evaluation result is greater than the maximum value among the evaluation results corresponding to multiple classification results.
[0026] In some embodiments, the method further includes:
[0027] In the case where there is an unassigned account node among the first account node and the second account node, assign the unassigned account node to the n-th classification result to obtain the allocation result corresponding to the n-th classification result.
[0028] In some embodiments, the method further includes:
[0029] When the first account node and the second account node do not belong to the shared nodes, and the first account node and the second account node are assigned to the nth classification result, the xth resource interaction event is assigned to the nth classification result to obtain the allocation result corresponding to the nth classification result;
[0030] When the first account node is assigned to the nth classification result, the second account node is a shared node, and the first account node does not belong to the shared node, the xth resource interaction event is assigned to the nth classification result;
[0031] When the second account node is assigned to the nth classification result, the first account node is a shared node, and the second account node does not belong to the shared node, the xth resource interaction event is assigned to the nth classification result.
[0032] In some embodiments, the method further includes:
[0033] Before obtaining the load balancing coefficients respectively corresponding to multiple classification results and obtaining the weight parameters corresponding to the account nodes in each classification result, the method further includes:
[0034] Determine the degree of each account node according to the number of neighbor nodes of each account node in the resource interaction graph;
[0035] Select the top K account nodes with the highest corresponding degrees from each account node included in the resource interaction graph as the shared nodes, where K is a positive integer.
[0036] In a second aspect, an embodiment of the present application provides a data processing device, including:
[0037] A first acquisition module, configured to acquire a resource interaction graph corresponding to the data to be trained, where the data to be trained is used to indicate the resource interaction situation between multiple accounts, the resource interaction graph includes multiple account nodes, where there is at least one resource interaction event between at least two account nodes, and multiple account nodes include shared nodes, and the resource interaction graph corresponds to multiple classification results, and each classification result includes shared nodes;
[0038] A second acquisition module, configured to acquire the load balancing coefficients respectively corresponding to multiple classification results, and acquire the weight parameters respectively corresponding to the account nodes in multiple classification results, where the load balancing coefficient represents the load situation corresponding to the classification result, where the nth weight parameter corresponding to the ith account node is used to indicate that the ith account node is related to the neighbor nodes in the nth classification result, and there is a resource interaction event between the neighbor nodes and the ith account in the resource interaction graph, and i and n are positive integers;
[0039] A determination module, configured to determine allocation results corresponding to multiple classification results based on multiple load balancing coefficients and multiple weight parameters, where the allocation results include account nodes allocated to the classification results and resource interaction events related to the account nodes;
[0040] A training module, configured to train a sample model based on the allocation results corresponding to multiple classification results to obtain a target model, where the target model is used to identify abnormal resource interaction events or abnormal account nodes.
[0041] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the data processing method described in any embodiment of the first aspect.
[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the data processing method described in any embodiment of the first aspect is implemented.
[0043] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is run, the data processing method described in any embodiment of the first aspect is executed.
[0044] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:
[0045] Each classification result corresponding to the resource interaction graph includes shared nodes among multiple account nodes, so that when training the sample model, the trainer can read the features of the shared nodes locally, without obtaining relevant data of the shared nodes by accessing other trainers, which can reduce the cross-machine communication access overhead, thereby improving the training efficiency of the model. When allocating account nodes other than the shared nodes in the resource interaction graph, considering the load situation (i.e., the load balancing coefficient) of each classification result and the resource interaction situation (i.e., the weight parameter) between each account node and the account nodes in the classification result, determine the allocation result corresponding to each classification result, and realize considering the load situation of the classification result when dividing the resource interaction graph to ensure load balancing among classification results. And considering the influence of the account nodes in the classification result on the currently allocated account nodes, considering the internal connection between data, ensuring that account nodes with close internal connections are allocated to the same classification result as much as possible, so as to reduce the cross-machine communication access overhead during training, realize a reasonable division of the resource interaction graph, and improve the training efficiency of the model at the same time. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 is a flowchart of a current distributed training MTGNN model;
[0048] Figure 2 is a schematic diagram of an application scenario of the data processing method applicable to the present application;
[0049] Figure 3 is a schematic flowchart of a data processing method provided by an embodiment of the present application;
[0050] Figure 4 is a schematic diagram of a resource interaction diagram provided by an embodiment of the present application;
[0051] Figure 5 is a schematic flowchart of a process for obtaining load balancing coefficients corresponding to multiple classification results in an embodiment of the present application;
[0052] Figure 6 is a schematic flowchart of a process for obtaining weight parameters corresponding to an account node in each classification result in an embodiment of the present application;
[0053] Figure 7 is a schematic flowchart of another data processing method provided by an embodiment of the present application;
[0054] Figure 8 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;
[0055] Figure 9 is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application. Detailed implementation manners
[0056] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.
[0057] It should be understood that, as used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups.
[0058] It should also be understood that the term "and / or" as used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0059] As used in the specification of this application and the appended claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0060] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0061] The reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0062] In the financial transaction scenario, financial transaction records are often stored in the form of graph data. Graph data is a data type represented and stored in a graph structure, consisting of nodes (Vertices) and edges (Edges), and is used to describe entities and the relationships between them. In the graph data, each account in the financial transaction scenario is regarded as a node in the resource interaction graph, and each transaction is regarded as an edge on the graph. That is to say, the edge represents the existence of an interaction event between two nodes.
[0063] In practical applications, by introducing a time dimension based on graph data, a resource interaction graph with timestamps can be constructed, and the graph neural network model can be trained using the resource interaction graph, so that the trained graph neural network model can predict whether a transaction will occur between two accounts at a future moment to identify potential abnormal transaction behaviors or abnormal accounts in financial transaction scenarios. Generally speaking, the graph neural network models adopted often have excellent capabilities in node feature representation and link relationship prediction. For example, the Memory based Temporal Graph Neural Network (MTGNN) or the Conditional Temporal Dynamic Graphs (CTDGs) model, etc. However, the training speed of such graph neural network models is relatively slow, and generally users will accelerate the training speed of the graph neural network model through distributed training.
[0064] Generally speaking, the distributed training of graph neural network models usually includes two methods: model parallelism and data parallelism. In the model parallelism method, different layers or different modules of the graph neural network model are assigned to different trainers for parallel computing to accelerate the training speed. The data parallelism method is to install the same graph neural network model in multiple trainers, but each trainer uses different training data for model training. Generally speaking, the data parallelism method will divide a large-scale training data set into multiple subsets, and each trainer is responsible for processing one subset. For example, when processing large-scale traffic flow time series data, the data of different road segments or different time periods are assigned to different trainers. The graph neural network model on each trainer performs forward propagation and calculates gradients and other operations based on local data, and then aggregates and updates the gradients or model parameters of each trainer through a communication mechanism, so that each trainer can obtain global training information, thereby realizing the overall optimization of the model.
[0065] The distributed training of the graph neural network model involved in the embodiments of this application takes the data parallelism method as an example. The following combines Figure 1 , taking the graph neural network model as the MTGNN model as an example to introduce the process of distributed training of the graph neural network. Figure 1 is a flowchart of distributed training of the MTGNN model. As Figure 1 shown, the MTGNN model includes a memory update layer and a message passing layer. When the MTGNN model is distributedly trained by multiple trainers, the MTGNN model is installed on multiple trainers in advance, so that each trainer integrates the memory update layer and the message passing layer of the MTGNN model.
[0066] Before starting to train the MTGNN model, the resource interaction graph required by the MTGNN model (i.e., the training dataset) needs to be divided into multiple subgraphs (also called classification results, corresponding to subsets of the training dataset), and the graph classification step in Figure 1 is completed. And each trainer is set to manage a subgraph, that is, the local storage in Figure 1 . Each subgraph contains the nodes and edges in the resource interaction graph divided into itself, the features of the nodes and edges (the features include the static point-edge features in Figure 1 ), the memory units corresponding to the nodes and edges, and the last transaction information stored in the memory units.
[0067] During distributed training, as shown in Figure 1 , the data in different subgraphs is trained in batches in parallel according to the time order. That is, each trainer processes a batch of edges in the subgraph it manages simultaneously according to the time order (the edges at this time include the two nodes it connects, as well as the features of the edges and nodes). Each trainer samples a subgraph from the locally stored subgraph (that is, samples the data required for batch training from the subgraph) for subsequent training. During the sampling process, the events being trained can be processed by negative sample sampling and the temporal neighbors of the nodes in the sampled subgraph. Subsequently, the neighbor nodes of the node stored in the remote server, as well as the edges, static features, and the state of the memory unit related to them (the neighbor nodes), are pulled from the remote machine.
[0068] The trainer inputs the relevant data of the pulled and sampled nodes and edges into the memory update layer of the corresponding MTGNN model. The memory update layer updates at least one memory unit and the last transaction information (the transaction information can be understood as the relevant information of newly occurred resource interaction events) in the subgraph managed by the trainer according to the input data. Figure 1 In , the update process is marked with different colors respectively. Then, through the indexes of the nodes and edges, and the timestamps of the nodes and edges, the relevant data of the nodes and edges in the remote server are aggregated and the node embeddings are updated. And during the model training process, a remote memory unit aggregation operation is performed to integrate the memory unit information from different subgraphs or remotely to obtain more comprehensive graph information. After the model calculations on different trainers are completed, model gradient synchronization is performed to ensure that the model parameter updates on each node are consistent, so as to achieve the coordination of distributed training and promote the continuous optimization of the model.
[0069] From the training process of the above model, it can be seen that the main communication stages involved in the process of distributed training of the graph neural network model are remote feature pulling, obtaining the state of the memory unit, and aggregating the state of the remote memory unit. The communication during the training process occupies a large amount of overhead, which affects the efficiency of distributed training.
[0070] Therefore, the partitioning of the resource interaction graph involved in the graph neural network model becomes crucial.
[0071] Currently, the partitioning schemes for resource interaction graphs generally include the following three types: The first is to minimize the cut edges, partition the resource interaction graph into several subgraphs, and ensure the minimization of the number of edges in each subgraph; the second is the graph partitioning scheme based on Distributed Temporal Graph Learning (DisTGL), which considers the time information ignored in the graph partitioning scheme that minimizes the cut edges, and tries to partition the nodes and edges with similar timestamps in the resource interaction graph into the same subgraph for additional load balancing; the third is that during the continuous evolution of the resource interaction graph, an incremental graph (i.e., the part where the graph structure and data change) will be generated. For each incremental graph, a graph partitioning operation is performed again. When there are new node, edge, or attribute changes, this part of the incremental graph data is extracted, and then according to a certain graph partitioning strategy, the incremental graph is partitioned into each subgraph (the third type is also called the Sven graph partitioning scheme).
[0072] In the distributed training process of the graph neural network, each trainer processes a batch of data in its assigned subgraph in a sequential and synchronous manner. Each trainer must wait for all other trainers to complete processing the same batch of data before synchronizing the memory. However, the above graph partitioning schemes all perform graph partitioning from the perspective of the overall structure of the resource interaction graph, ignoring the internal connections of the data in the resource interaction graph, resulting in the problem of unbalanced training task load during the model training process, leading to high synchronous computing overhead and affecting the model training efficiency.
[0073] Moreover, the above graph partitioning schemes often consider minimizing the number of edges within the subgraph while ignoring the connection relationships between nodes in the resource interaction graph, resulting in more cut edges. The high-degree nodes (also called shared nodes, which are frequently connected to other nodes in the resource interaction graph) in the resource interaction graph will be partitioned into different subgraphs. When training the model, the degree nodes become the nodes that are frequently accessed across trainers, resulting in high communication access overhead and further affecting the model training efficiency. In addition, the third method above requires re-partitioning for each incremental graph, with high computational complexity.
[0074] In view of the above problems, the embodiments of the present application provide a data processing method, apparatus, device, storage medium, and computer program product. When partitioning a resource interaction graph, each shared node is assigned to each classification result, so that the shared nodes are distributed locally on the trainer, reducing the number of cross-classification result samplings during training, reducing the overhead of cross-trainer communication, and improving the model training efficiency. When partitioning the resource interaction graph, the load balance between classification results is ensured by a load balancing coefficient, and the influence of the existing account nodes in the classification result on the currently assigned account node is considered through the weight parameters corresponding to the account nodes, taking into account the importance of the neighbors of the account nodes in the resource interaction graph, achieving a reasonable partitioning of the resource interaction graph, and making the closely related account nodes be assigned to the same classification result as much as possible, so that during model training, the trainer can access local data most of the time, reducing the communication overhead between trainers and further improving the training efficiency.
[0075] Figure 2 FIG. is a schematic diagram of an application scenario suitable for the data processing method of the present application. As Figure 2 shown, this application scenario includes a server 110 and multiple terminal devices 120, and the terminal devices 120 are communicatively connected to the server 110.
[0076] The terminal device 120 can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, an e-book reader, an intelligent voice interaction device, an intelligent home appliance, a vehicle-mounted terminal, etc.; a client related to interactive data processing can be installed on the terminal device, and the client can be software (such as a browser, a live broadcast software, a shopping software, an instant messaging software, etc.), or a web page, a small program, etc., and the server 110 is the background server corresponding to the software or the web page, the small program, etc., or a server specifically used for interactive data processing, and the present application does not make specific limitations. The server 110 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0077] It should be noted that the data processing method in each embodiment of the present application can be executed by an electronic device, and the electronic device can be the terminal device 120 or the server 110, that is, the method can be executed independently by the terminal device 120 or the server 110, or can be executed jointly by the terminal device 120 and the server 110.
[0078] Taking the independent execution of server 110 as an example, server 110 may include multiple trainers, and sample models may be deployed on each trainer. Server 110 may obtain a resource interaction graph corresponding to the data to be trained, and divide the account nodes and resource interaction events in the resource interaction graph into different classification results. In this way, the resource interaction graph is divided into multiple sub-graphs, and different sub-graphs are assigned to each trainer of its own, so as to perform distributed training on the sample model through its own trainers. The specific division process of the resource interaction graph will be introduced in the subsequent embodiments. The trained sample model can be used to predict the abnormal transaction behavior of users in the terminal device 120 or classify the transaction behavior of users, etc.
[0079] The data processing method of the present application will be introduced in detail below.
[0080] Figure 3 It is a schematic flowchart of a data processing method provided by an embodiment of the present application. As Figure 3 shown, the method includes the following steps:
[0081] S101, obtain a resource interaction graph corresponding to the data to be trained.
[0082] Among them, the data to be trained is used to indicate the resource interaction situation among multiple accounts. The resource interaction graph includes multiple account nodes, among which there are resource interaction events between at least two account nodes. The multiple account nodes include shared nodes, and the resource interaction graph corresponds to multiple classification results, and each classification result includes shared nodes.
[0083] The data to be trained may be a graph data representing the resource interaction relationship between accounts. In the financial transaction scenario, the data to be trained may include different accounts and transaction records (an example of resource interaction events) of different accounts at different times. For example, the data to be trained may include record 1 of account A transferring Y1 yuan to account B at time T1, record 2 of account B withdrawing Y2 yuan from account C at time T2, and record 3 of account C selling item C to account A and account B respectively at time T3. At this time, accounts A to C are respectively account nodes A to C in the resource interaction graph, and so on. Record 1 is a resource interaction event (i.e., an edge) between account node A and account node B.
[0084] The resource interaction graph is a continuous-time dynamic graph, and the account nodes and resource interaction events therein will change with time, and it can describe the process of the evolution of the data to be trained (graph structure) over time. The resource interaction graph includes a set of time-series edges E with timestamps and a set of account nodes V. Each edge in the set of time-series edges E is a resource interaction event existing between two account nodes. Specifically, the resource interaction graph can be defined as a series of time-ordered resource interaction events:
[0085] G = {e(u0, v0, t0, f0), e(u1, v1, t1, f1), …, e(u i , v i ,, t i , f i ), …, e(u n , v n , t_n, f n )}
[0086] Among them, G represents the resource interaction graph, and each resource interaction event e(u i , v i ,, t i , f i ) represents a directed temporal edge from the account node u i to the account node v i occurring at the timestamp t i . f i represents the feature vector of the resource interaction event e(u i , v i ,, t i , f i ) at the timestamp t i . It can be understood that the resource interaction event is the edge in the resource interaction graph. The set of account nodes V includes all the account nodes in G. For example, u n and v i , etc. The set of temporal edges E includes all the resource interaction events e in G. The resource interaction graph can also be represented as G = (V, E, X N , X E ), where X N is the feature matrix storing the features of the account nodes, and X E is the feature matrix storing the features of the edges.
[0087] The shared node is an account node that can be replicated in the resource interaction graph. It is connected to multiple account nodes in the resource interaction graph. That is, the shared node is an account node with a relatively high degree (number of degrees) in the resource interaction graph.
[0088] In one implementation, the electronic device may receive the data to be trained, the preset proportionality coefficient, and the number n (n>1) of classification results to be divided input by the user or other terminal devices, construct a corresponding resource interaction graph based on the data to be trained, and generate n empty classification results based on the number n. At least one shared node is selected from all the account nodes of the resource interaction graph based on the preset proportionality coefficient, and after obtaining the shared nodes, each shared node is divided into each classification result, that is, the features of each shared node are additionally stored in each trainer, so that the trainer can subsequently train the sample model based on its own allocation results. The electronic device also initializes the account nodes and resource interaction events other than the shared nodes in the resource interaction graph to an unallocated state, so as to subsequently allocate the account nodes and resource interaction events in the resource interaction graph to the n classification results based on S102 to S103.
[0089] The number n of classification results can be set by the user or the electronic device based on factors such as the number of trainers participating in the training of the sample model and the memory size, and the specific setting method is not limited in the embodiments of the present application. The preset proportionality coefficient (or the number of shared nodes) can be set by the user or the electronic device based on the connection situation between each account node in the resource interaction graph and the total number of account nodes in the resource interaction graph, and the specific setting method is not limited in the embodiments of the present application. For example, the number of shared nodes can be set to 5% or 1% of the total number of account nodes in the resource interaction graph, and the preset proportionality coefficient at this time is 5% or 1%.
[0090] It should be noted that since the shared nodes are frequently connected to multiple account nodes in the resource interaction graph, the shared nodes are usually frequently accessed by the trainer during model training. If the shared nodes are stored separately in a remote device or some of the multiple trainers, when training the sample model, the trainer not only needs to read local data, but also needs to frequently access the relevant data in the trainer storing the shared nodes, resulting in a large cross-machine communication overhead. After allocating all the shared nodes to each classification result, it can be ensured that each trainer additionally stores the features of the shared nodes, so that the trainer can obtain the relevant data through local communication without cross-trainer communication, reducing the communication overhead, thereby improving the efficiency of subsequent model training.
[0091] In one implementation, before obtaining the load balancing coefficients corresponding to multiple classification results and the weight parameters corresponding to the account nodes in each classification result, the method further includes:
[0092] Determine the degree of each account node according to the number of neighbor nodes of each account node in the resource interaction graph;
[0093] From each account node included in the resource interaction graph, select the top K account nodes with the highest corresponding degrees as shared nodes, where K is a positive integer.
[0094] The electronic device can calculate the product result between the preset proportional coefficient and the total number of account nodes, and round up the product to obtain K. For each account node, the electronic device can count the number of other account nodes connected to this account node through edges in the resource interaction graph, that is, count the number of its neighbor nodes, and use this number as the degree D(v) of this account node, where v represents the number of neighbor nodes of the account node, to determine the importance and connectivity of this account node. Furthermore, taking the top K account nodes with the highest corresponding degrees as shared nodes can ensure the reliability of the shared nodes and provide a data basis for subsequent model training.
[0095] Combined with Figure 4 the schematic diagram of a resource interaction graph in an embodiment of the present application shown in Figure 4 the resource interaction graph includes account nodes A and B and account nodes 1 to 15, a total of 17 account nodes, and edges e1 to e21.
[0096] If there are two resource interaction event transactions between account A and account 1 in the training data, then account nodes A and 1 can be added in the resource interaction graph, and edges e1 and e2 are formed between the two account nodes. If there are two resource interaction event transactions between account A and account 3, then account node 3 can be added in the resource interaction graph, and edges e1 and e4 are formed between the two account nodes, and so on. If there are three resource interaction event transactions B between account 7 and account B, then account nodes B and 3 can be added in the resource interaction graph, and edges e12, e13, and e14 are formed between the two account nodes, and finally the Figure 4 resource interaction graph shown in
[0097] Taking the preset proportional coefficient as 10% as an example, Figure 4 it includes 17 account nodes, 17 * 10% and after rounding up K = 2, Figure 4 there are 4 account nodes connected to account node A in , then the degree of account node A is 4. Similarly, the degree of account node B is 5, the degrees of account nodes 2, 4, and 5 are all 3, and the degrees of the remaining account nodes are 1 or 2. Then the electronic device will take the top 2 account nodes with the highest degrees, that is, account nodes A and B as shared nodes.
[0098] The edges (that is, resource interaction events) in the resource interaction graph can be directed or undirected. Figure 4 The directed edge is taken as an example in .
[0099] S102. Obtain the load balancing coefficients corresponding to multiple classification results respectively, and obtain the weight parameters corresponding to the account nodes in multiple classification results respectively.
[0100] Among them, the load balancing coefficient represents the load situation corresponding to the classification result. Among them, the nth weight parameter corresponding to the ith account node is used to indicate that the ith account node is related to the neighbor nodes in the nth classification result, and there is a resource interaction event between the neighbor nodes and the ith account in the resource interaction graph. i and n are positive integers.
[0101] The load balancing coefficient includes a node load coefficient, an event load coefficient, and a time load coefficient. Among them, the node load coefficient represents the situation of the account node in the classification result from the node load upper limit. The event load coefficient represents the difference between the number of resource interaction events in the classification result and the number of resource interaction events in the other classification results. The time load coefficient represents the difference between the timestamps of the resource interaction events in the classification result and the timestamps of the resource interaction events in the other classification results.
[0102] The electronic device can use the load balancing coefficient to ensure that the number of account nodes and resource interaction events allocated to each classification result is balanced. At the same time, introducing the time load coefficient into the load balancing coefficient can restrict the difference in the timestamps of the resource interaction events in each classification result, so as to avoid the resource interaction events being concentrated and allocated to one classification result within the same time range. In this way, it can be ensured that within the time range corresponding to each training batch in the subsequent process of training the sample model, the number of training tasks allocated to each trainer is balanced, and thus the load balancing during the subsequent training of the sample model can be ensured. The sample model in the embodiments of the present application is the graph neural network model to be trained.
[0103] The current graph partitioning scheme usually gives priority to reducing the cut edges and often ignores the connection between an account node and other account nodes it is connected to. However, in the training process of the sample model, considering the resource interaction and then determining the weight parameter corresponding to the account node in the classification result based on the timestamp of the neighbor nodes can effectively aggregate the information of the neighbor nodes, enhance the node representation ability, and thus improve the accuracy of the trained model in tasks such as prediction or classification.
[0104] Therefore, the embodiments of the present application propose weight parameters to represent the degree of dependence between an account node and neighbor nodes in the classification result. In this way, when the electronic device allocates account nodes in the resource interaction graph, the influence of the account nodes in the classification result on the currently participating account node can be considered through the weight parameters, ensuring that as much neighborhood aggregation information as possible is retained in the classification result, that is, nodes with close relationships are allocated to the same classification result as much as possible. Thus, when training the sample model subsequently, not only can the sample model more effectively aggregate the information of neighbor nodes, but also the trainer of the sample model can read relevant data locally to reduce cross-machine communication.
[0105] In the context of training a graph neural network model, the weights borne by neighbor nodes of an account node at different timestamps in the resource interaction graph follow an exponential distribution, and the attention weight of the resource interaction event between the account node and the neighbor node to the neighbor node decreases with time accumulation. Therefore, the weight parameter is also a parameter that decreases with time accumulation.
[0106] In one implementation, the electronic device traverses each resource interaction event in the resource interaction graph in chronological order, and for each traversed resource interaction event, obtains the current node load factor, event load factor, and time load factor corresponding to each classification result, and takes the product result of the obtained node load factor, event load factor, and time load factor as the load balancing factor corresponding to the classification result. For each account node connected to the resource interaction event, the neighbor nodes of the account node in each classification result are respectively obtained, and then the weight parameter corresponding to the account node in the classification result is determined based on the timestamps of the neighbor nodes. The specific determination processes of the load balancing factor and the weight parameter are introduced in subsequent embodiments.
[0107] It should be noted that there may be isolated account nodes in the resource interaction graph. When the electronic device traverses the resource interaction event, it will also traverse the two account nodes connected to it. Therefore, after the electronic device traverses the resource interaction event, the un-traversed account nodes can be determined as isolated nodes. The electronic device will continue to traverse the isolated account nodes in chronological order, and for each isolated account node, obtain the current load balancing factor corresponding to each classification result, and the weight parameter corresponding to each isolated account node in each classification result.
[0108] S103. Determine the allocation results corresponding to the multiple classification results based on the multiple load balancing factors and the multiple weight parameters.
[0109] Among them, the allocation result includes the account nodes allocated to the classification result and the resource interaction events related to the account nodes.
[0110] In one implementation, for each traversed resource interaction event, the electronic device determines the evaluation result corresponding to each classification result based on the load balancing coefficient corresponding to each classification result and the weight parameter corresponding to each account node connected to the resource interaction event in each classification result. The evaluation result is used to represent the possibility that the resource interaction event is assigned to the classification result. The resource interaction event is assigned to the classification result with the largest corresponding evaluation result. If there are unassigned account nodes among the two account nodes connected to the resource interaction event, the unassigned account nodes are continuously assigned to the classification result with the largest corresponding evaluation result. In this way, the electronic device can obtain the assignment results corresponding to multiple classification results respectively.
[0111] It can be understood that for isolated account nodes, the electronic device can also determine to assign them to the corresponding classification results based on a similar process. Only at this time, the electronic device determines the evaluation result corresponding to each classification result based on the load balancing coefficient corresponding to each classification result and the weight parameter corresponding to the account node in each classification result. The evaluation result at this time is used to represent the possibility that the account node is assigned to the classification result.
[0112] It should be noted that at the initial assignment, that is, when just starting to assign resource interaction events, each classification result only includes shared nodes. For the currently participating resource interaction event, the evaluation results corresponding to each classification result should be the same, that is, all classification results belong to the classification result with the largest evaluation result. At this time, the resource interaction event and the unassigned nodes among the two account nodes connected to it can be randomly assigned to a classification result.
[0113] Combined with Figure 4 , the resource interaction graph corresponds to two classification results (Classification Result 1 to 2). After the electronic device assigns 17 account nodes and 21 resource interaction events in the resource interaction graph based on multiple load balancing coefficients and multiple weight parameters, the obtained assignment results can be as shown in Classification Result 1 and 2 in Figure 4 . It can be seen that in the assignment results determined based on the load balancing coefficient and the weight parameter, the nodes that are closely related to each other in the resource interaction graph can be divided into the same classification result. Then, during the subsequent sample model training, the required data can be read from the local classification results, reducing the communication overhead and improving the model training efficiency.
[0114] S104. Train the sample model based on the assignment results corresponding to multiple classification results respectively to obtain a target model.
[0115] Among them, the target model is used to identify abnormal resource interaction events or abnormal account nodes. The sample model can be a graph neural network model such as an MTGNN model or a CTDGs model, which can be set by itself specifically, and the implementation of this application does not make specific limitations.
[0116] In the embodiments of the present application, distributed training is performed on the sample model. Taking the data parallelism method described above as an example, the electronic device can pre-control multiple trainers to load the sample model respectively, so that the same sample model is installed in each trainer. The trainer can be a hardware device such as a Graphics Processing Unit (GPU), a Central Processing Unit (CPU), or an Application-Specific Integrated Circuit (ASIC), or it can be a computing node in the electronic device. The present application does not limit the specific type of the trainer. The account nodes and resource interaction events in the allocation result corresponding to each classification result form a subgraph of the resource interaction graph.
[0117] After the electronic device obtains the allocation results corresponding to each classification result respectively, that is, after obtaining the subgraphs, it will read the features of the account nodes and the features of the resource interaction events in each subgraph from the resource interaction graph based on the indexes of the account nodes and the indexes of the resource interaction events in each subgraph, and store them in the subgraphs. Furthermore, the electronic device will allocate at least one subgraph to each trainer. In this way, each trainer can train its own sample model based on the allocated subgraph, and synchronize parameters such as the gradients of the sample models in each trainer through the communication mechanism, and finally obtain the target model. In the embodiments of the present application, it is taken as an example that one trainer is allocated one subgraph.
[0118] The training process of the sample model is generally divided into multiple time batches (which can be understood as dividing the subgraph into multiple small graph structures, or sampling the subgraph multiple times), and the training is carried out in chronological order. In each time batch, the resource interaction events should be trained in the order of strictly increasing timestamps. In the distributed training process, each trainer is only responsible for training the resource interaction events allocated locally (in the subgraph), and the resource interaction events belonging to the same batch on different trainers are trained simultaneously.
[0119] Before training the sample model, the electronic device will create a memory unit for the account nodes in each classification result to store the transaction information generated during the training process through the memory unit.
[0120] Taking the CTDGs model as an example for the sample model in the embodiments of the present application, this model also includes a memory update layer and a message passing layer. Its training process is similar to Figure 1 the training process of the MTGNN model shown in Figure 1 The training process of the sample model is described by replacing the MTGNN model in with the CTDGs model (that is, replacing the TGNN data parallel model with the CTDGs model).
[0121] After the electronic device assigns subgraphs to each trainer, the electronic device can control each trainer to perform the following steps:
[0122] Step (1): Obtain the training events of the current time batch in chronological order from early to late.
[0123] The subgraphs assigned to the trainer are pre-divided by time batch. For the current time batch, the trainer samples its assigned subgraph. First, randomly sample a batch of account nodes in the subgraph for the current time batch, and use the sampled account nodes as negative samples. Among them, the randomly sampled account nodes are called target nodes. Then, sample the source nodes in the resource interaction graph for the current time batch. The account nodes that have resource interaction events with the target nodes in the resource interaction graph are source nodes. Then the trainer samples the source nodes and target nodes to construct a sampled subgraph. Specifically, the trainer extracts the source nodes, target nodes, and the resource interaction events between them from its own subgraph and the resource interaction graph as a small graph structure (that is, the sampled subgraph).
[0124] When the trainer extracts the source nodes from the resource interaction graph and the resource interaction events between the source nodes and target nodes, it determines the remote trainer (that is, other trainers except itself among multiple trainers) where the subgraph storing the source node is located based on the index of the source node, and determines the remote trainer where the subgraph storing the resource interaction event is located based on the index of the resource interaction event between the source node and target node. And the trainer pulls the memory unit of the source node, the features of the source node, and the features of the resource interaction event between the source node and target node in the sampled subgraph from the subgraph of the determined remote trainer to its own storage space. The trainer also extracts the memory unit of the target node and the features of the resource interaction event between the source node and target node from its own subgraph, and finally obtains the training event. The training event includes the memory unit of the source node, the memory unit of the target node, the features of the resource interaction event between the two, the features of the source node, the features of the target node, and the subgraph structure of the sampled subgraph.
[0125] Step (2): Input the collected training events into the sample model, so that the sample model updates the memory units corresponding to at least one node in the training events based on the training events, and performs feature splicing based on the updated memory units to obtain the latest transaction information corresponding to the trainer.
[0126] The trainer also obtains the last transaction information obtained before the current moment, and this transaction information is the information obtained after aggregating the transaction information in the memory unit for the last time before the current moment. The trainer inputs the sampled subgraph structure (i.e., the structure of the sampled subgraph constructed), the features of the target nodes, the features of the source nodes, the features of the resource interaction events between the target nodes, the memory units of each node, and the last obtained transaction information into a sample model including a memory update layer and a message passing layer. The memory update layer in the sample model outputs the latest memory unit based on the input data (training events as transaction information), and the memory update layer is mainly composed of a recurrent neural network (RNN) or a gated recurrent unit (GRU) network layer.
[0127] The memory update layer determines the newly occurred resource interaction events from the input data, concatenates the memory units (transaction information in them) of the source nodes, the memory units (transaction information in them) of the target nodes, and the features of the resource interaction events between the two in the newly occurred resource interaction events, and sends the concatenated features to the message passing layer. The message passing layer aggregates the historical features and the concatenated features to obtain the representation of the last interaction corresponding to the trainer at the current moment (i.e., the latest transaction message).
[0128] Step (3): Send the latest transaction information corresponding to the trainer to the remote memory unit aggregation module, so that the remote memory unit module aggregates the latest transaction information corresponding to each trainer to obtain the aggregated transaction information (i.e., the last transaction information at the current moment).
[0129] When the electronic device performs distributed training of the graph neural network model, in addition to including multiple trainers, it also includes a remote memory unit aggregation module. The trainer sends the latest transaction information and the updated memory unit obtained from the sample model to the remote memory unit aggregation module, so that the remote memory unit aggregation module performs transaction information aggregation operations according to the time stamps of the latest transaction information sent by each trainer and the indexes of the nodes corresponding to the updated memory units, and obtains the aggregated transaction according to the strategy of the largest time stamp.
[0130] The remote memory unit aggregation module also sends the aggregated transaction information to each trainer, and the trainer can update the mailbox storage and node memory storage of the corresponding nodes in its own subgraph. The mailbox storage of the node is used to receive messages from other trainers and module communications, and the node memory storage is used to save information such as the updated memory unit in the trainer's own subgraph.
[0131] Step (4): Based on the aggregated transaction information, update and synchronize the parameters of the sample model to obtain the updated sample model, and the parameters include the gradients of the sample model.
[0132] The trainer controls the sample model to calculate its latest parameters based on the aggregated transaction information and sends these latest parameters to the parameter aggregation module in the electronic device. The parameter aggregation module aggregates the latest parameters of the sample model under each trainer and sends the aggregated parameters to each trainer. Each trainer can then update its own sample model based on the aggregated parameters. Specifically, the aggregation operation can be a weighted average of each parameter.
[0133] In step (5), if the current time batch is not the last time batch, the next time batch is used as the current time batch, and the process returns to step (1) until the training events under all time batches are trained in sequence to obtain the target model.
[0134] In an embodiment of the present application, each classification result corresponding to the resource interaction graph includes a shared node among multiple account nodes, so that when training the sample model, the trainer can read the characteristics of the shared node locally, without having to access other trainers to obtain the relevant data of the shared node, which can reduce the communication access overhead across machines, thereby improving the training efficiency of the model. When allocating account nodes other than shared nodes in the resource interaction graph, the load situation of each classification result (i.e., load balancing coefficient) and the resource interaction situation between each account node and the account node in the classification result (i.e., weight parameter) are combined to determine the allocation result corresponding to each classification result, so that the load situation of the classification result is considered when dividing the resource interaction graph, ensuring load balancing between each classification result. The impact of the account node in the classification result on the currently allocated account node is also considered, and the internal connection between the data is considered to ensure that the account nodes with close internal connections are allocated to the same classification result as much as possible, thereby reducing the communication access overhead across machines during training, achieving a reasonable division of the resource interaction graph, and improving the training efficiency of the model.
[0135] Figure 5 FIG. 1 is a flow chart of obtaining load balancing coefficients corresponding to multiple classification results in an embodiment of the present application. Figure 5 As shown, obtaining the load balancing coefficients corresponding to the multiple classification results may include the following S201-S206:
[0136] S201: Obtain the first node quantity of the account nodes and the first event quantity of the resource interaction events in the nth classification result.
[0137] Any one of the multiple classification results in the nth classification result. The electronic device can respectively count the total number of account nodes and resource interaction events in the nth classification result, and the counted total number of nodes is used as the first node number, and the counted total number of resource interaction events is used as the first event number.
[0138] S202. Obtain the timestamp parameters and the number of events of the resource interaction events corresponding to multiple classification results except the nth classification result, and the target timestamp of the resource interaction events in the nth classification result.
[0139] Among them, the timestamp parameters include a first timestamp and a second timestamp, and the number of events includes a second number of events and a third number of events.
[0140] The first timestamp represents the maximum value among the timestamps of the resource interaction events in each classification result, the second timestamp represents the minimum value among the timestamps of the resource interaction events in each classification result, the second number of events represents the largest number of resource interaction events in each classification result, and the third number of events represents the smallest number of resource interaction events in each classification result. The target timestamp represents the maximum value among the timestamps between the resource interaction events in the nth classification result.
[0141] For each classification result except the nth classification result among the multiple classification results, the electronic device will count the number of events of the resource interaction events in this classification result and the maximum value of the timestamps of the resource interaction events in this classification result. The electronic device will select the largest number among the numbers of events corresponding to all classification results as the second number of events, the smallest number as the third number of events, select the largest timestamp among the timestamps of the resource interaction events of all classification results as the first timestamp, and select the smallest timestamp among the timestamps of the resource interaction events of all classification results as the second timestamp.
[0142] In an example, taking n as 2 and the multiple classification results being the 1st to 4th classification results respectively, except for the 2nd classification result, the number of events of the 1st classification result (resource interaction events) is 5, the number of events of the 3rd classification result is 10, and the number of events of the 4th classification result is 8, then the second number of events is 5 and the third number of events is 10.
[0143] Except for the 2nd classification result, the maximum value of the time indicated by the timestamp of the 1st classification result (resource interaction events) is 18:40:05, the maximum value of the time indicated by the timestamp of the 3rd classification result is 18:40:30, and the maximum value of the time indicated by the timestamp of the 3rd classification result is 18:39:27, then the time indicated by the first timestamp is 18:40:30 and the time indicated by the second timestamp is 18:39:27.
[0144] S203. Determine the first ratio result between the number of results of the multiple classification results and the total number of account nodes in the resource interaction graph, and determine the node load factor of the nth classification result based on the second ratio result between the first ratio result and the first number of nodes.
[0145] In one implementation, the electronic device may substitute the number of results of multiple classification results, the total number of account nodes in the resource interaction graph, and the number of the first nodes into Formula 1 to determine the node load factor of the nth classification result. Formula 1 is where BN(p n ) represents the node load factor of the nth classification result, p n represents the nth classification result, m is the number of results of multiple classification results, represents the number of the first nodes of account nodes in the nth classification result, N represents the total number of account nodes in the resource interaction graph, ∈ represents a hyperparameter of the sample model, (1 + ∈)N / m represents the upper limit of the number of account nodes that can be accommodated in each classification result, N / m represents the first ratio result, represents the second ratio result.
[0146] S204. Determine the event load factor of the nth classification result according to the difference between the third event quantity, the second event quantity, and the first event quantity.
[0147] In one implementation, the electronic device may substitute the third event quantity, the second event quantity, and the first event quantity into Formula 2 to determine the event load factor of the nth classification result. Formula 2 is where BE(p n ) represents the event load factor of the nth classification result, p j represents the jth classification, j is a positive integer and j is not equal to n, represents the number of the first events of resource interaction events in the nth classification result, represents the third event quantity, represents the second event quantity, p j <m represents that the jth classification result belongs to one of multiple classification results, ∈ E represents another hyperparameter of the sample model.
[0148] S205. Determine the time load factor of the nth classification result according to the difference between the first timestamp, the second timestamp, and the target timestamp, and determine the time load factor of the nth classification result.
[0149] In one implementation, the electronic device may substitute the first timestamp, the second timestamp, and the target timestamp into Formula 3 to determine the time load factor of the nth classification result. Formula 3 is where BT(p n ) represents the time load factor of the nth classification result, represents the second timestamp, represents the first timestamp, T(n) represents the target timestamp, ∈T Represents another hyperparameter of the sample model.
[0150] S206. Determine the load balancing coefficient corresponding to the nth classification result according to the product result among the node load coefficient, the event load coefficient, and the time load coefficient.
[0151] In one implementation, the electronic device may substitute the node load coefficient, the event load coefficient, and the time load coefficient of the nth classification result into Formula 4 to obtain the load balancing coefficient corresponding to the nth classification result. Formula 4 is F BAL (p n ) = BN(p n )·BE(p n )·BT(p n ), where F BAL (p n ) represents the load balancing coefficient corresponding to the nth classification result.
[0152] In the embodiments of the present application, the node load coefficient can reflect the difference between the allocated account nodes in the classification result and the node load upper limit in the classification result. The event load coefficient can reflect the difference between the allocated resource interaction events in the classification result and the event load upper limit in the classification result. The time load coefficient can reflect the time difference between the allocated resource interaction events in the classification result. Furthermore, the load balancing coefficient determined based on the product result among the node load coefficient, the event load coefficient, and the time load coefficient can comprehensively and accurately reflect the current load situation of the classification result from three aspects: the number of nodes, the number of resource interaction events, and the time of resource interaction events, providing a reliable data basis for subsequent allocation of nodes and events in the resource interaction graph, ensuring the load balance among each classification result, and also introducing timestamp information when subsequently allocating nodes and events in the resource interaction graph to ensure the time load balance of the classification result and avoid allocating resource interaction events within the same time range to the same classification result.
[0153] Figure 6 It is a schematic flowchart of a process for obtaining the weight parameters corresponding to an account node in multiple classification results in the embodiments of the present application. As Figure 6 shown, obtaining the weight parameters corresponding to the account node in each classification result includes the following S301 to S302:
[0154] S301. Obtain the target neighbor node of the ith account node and the third timestamp of the ith account node in the nth classification result, as well as the fourth timestamp corresponding to the resource interaction graph.
[0155] Among them, there is a resource interaction event between the target neighbor node and the i-th account node in the resource interaction graph, and the existing resource interaction event is assigned to the n-th classification result, and the fifth timestamp of the target neighbor node is included in the n-th classification result. The fourth timestamp is the maximum value of the timestamps of each resource interaction event in the resource interaction graph, and the i-th account node is any account node in the resource interaction graph.
[0156] The electronic device can count from the resource interaction graph each account node that has a resource interaction event with the i-th account node, as well as the resource interaction event between the two. For each counted account node, search for this account node in the n-th classification result, and after finding this account node, determine that this account node is a neighbor node of the i-th account node in the n-th classification result. For each counted resource interaction event, continue to search for this resource interaction event in the n-th classification result, and after finding this resource interaction event, determine that the found resource interaction event is a neighbor event of the i-th account node in the n-th classification result. The electronic device will also select the timestamp with the largest value from the timestamps of each resource interaction event in the resource interaction graph as the fourth timestamp.
[0157] S302. Determine the difference result between the third timestamp and the fifth timestamp, and based on the product result of the difference result and the reciprocal of the fourth timestamp, determine the weight parameter corresponding to the i-th account node in the n-th classification result.
[0158] The electronic device can substitute the third timestamp and the fifth timestamp into formula 5 to obtain the weight parameter corresponding to the i-th account node in the n-th classification result.
[0159] Formula 5 is Among them, Att(i,t,p m ) represents the weight parameter corresponding to the i-th account node at timestamp t in the n-th classification result, N u (t) represents the set of neighbor nodes of the i-th account node that have interacted before timestamp t in the n-th classification result, represents the set of resource interaction events related to the i-th account node before timestamp t in the n-th classification result, 9v,τ)∈N u (t) and jointly represent that the neighbor node v of the i-th account node is not only in the n-th classification result, but also the resource interaction event 9i,v,τ) between the i-th account and this neighbor node v is also in the n-th classification result, that is, 9v,τ) represents the target neighbor node, (i,v,τ) represents the resource interaction event between the i-th neighbor node and the target neighbor node, and τ represents the fifth timestamp. Among them, max e(u,v,t)∈EDenote the target timestamp as \(t\), and \(E\) represents the set composed of each resource interaction event in the resource interaction graph. The timestamp \(t\) can be the timestamp of the resource interaction event currently participating in the allocation corresponding to the \(i\)-th account node.
[0160] In the embodiment of the present application, the target neighbor node of the \(i\)-th account node in the classification result is the node that has a resource interaction event in the resource interaction graph and this resource interaction event is also in this classification result. Through the timestamp of the target neighbor node in this classification result and the target timestamp corresponding to the resource interaction graph, the influence of the neighbor nodes and neighbor events that have been allocated to the classification result on the \(i\)-th account node in the time dimension can be determined, so as to consider the importance of the neighbors of the account node in the classification result when allocating nodes and events in the resource interaction graph later, so as to retain as much neighborhood aggregation information as possible in the classification result.
[0161] Figure 7 This is another data processing method provided by the embodiment of the present application. As Figure 7 shown, this method includes the following S401 to S406:
[0162] S401, obtain the resource interaction graph corresponding to the data to be trained. For S401, please refer to Figure 3 S101 in the shown embodiment for details, which will not be elaborated here.
[0163] S402, obtain the load balancing coefficients corresponding to multiple classification results respectively, and obtain the weight parameters corresponding to the account node in multiple classification results respectively.
[0164] For details of S402, please refer to Figure 5 and Figure 6 the shown embodiment, which will not be elaborated here.
[0165] In one implementation, the resource interaction graph includes the \(x\)-th unallocated resource interaction event, the timestamp of the \(x\)-th resource interaction event, the first account node and the second account node, \(x\) is a positive integer, the first account node corresponds to the first weight parameter in the \(n\)-th classification result, and the second account node corresponds to the second weight parameter in the \(n\)-th classification result;
[0166] Based on multiple load balancing coefficients and multiple weight parameters, determine the allocation results corresponding to multiple classification results respectively, including S403 to S405:
[0167] S403, when the \(x\)-th resource interaction event does not meet the preset condition, determine the sum result between the first weight parameter and the second weight parameter.
[0168] Among them, the preset conditions include that the first account node and the second account node are not shared nodes, and the first account node and the second account node are assigned to the same classification result, or one of the first account node and the second account node is a shared node and the other is assigned to a classification result. The x-th resource interaction event is any resource interaction event in the resource interaction graph.
[0169] The electronic device traverses each resource interaction event in the resource interaction graph in chronological order, and executes the processing steps described in S403 to S405 for each traversed resource interaction event, so as to determine the classification result to which each resource interaction event belongs according to its account node. In the embodiments of the present application, the case where the electronic device traverses the x-th resource interaction event is taken as an example for description.
[0170] S404. Determine the n-th evaluation result corresponding to the n-th classification result according to the product result between the total result and the load balancing coefficient corresponding to the n-th classification result.
[0171] The electronic device may substitute the first weight parameter, the second weight parameter, and the load balancing coefficient corresponding to the n-th classification result into Formula 6 to obtain the n-th evaluation result corresponding to the n-th classification result. Formula 6 is S(e,p n )=[Att(q,t,p n )+Att(k,t,p n )+1]·F BAL (p n ) where S(e,p n ) represents the n-th evaluation result corresponding to the n-th classification result, e represents the x-th resource interaction event, q represents the first account node, k represents the second account node, Att(q,t,p n ) represents the first weight parameter, Att(k,t,p n ) represents the second weight parameter, and F BAL (p n ) represents the load balancing coefficient corresponding to the n-th classification result.
[0172] S405. When the n-th evaluation result meets the first condition, assign the x-th resource interaction event to the n-th classification result to obtain the assignment result corresponding to the n-th classification result.
[0173] Among them, the first condition is that the n-th evaluation result is greater than the maximum value among the evaluation results corresponding to multiple classification results.
[0174] For the x-th resource interaction event, the electronic device determines the evaluation result corresponding to each classification result based on Formula 6, and finds the maximum evaluation result from all the evaluation results, so as to assign the x-th resource interaction event to the classification result with the largest corresponding evaluation result, thereby obtaining the assignment result of this classification result. Specifically, the electronic device can be based on Formula 7: Determine the classification result with the largest corresponding evaluation result, where p represents the finally determined evaluation result.
[0175] For example, combined with Figure 4 , taking the resource interaction graph corresponding to two classification results (Classification Result 1 to 2), Figure 4 and the time order of the resource interaction events in being resource interaction events e1 to e21 (hereinafter referred to as edges for resource interaction events), and the x-th resource interaction event being the 6th resource interaction event (i.e., e6) as an example, for e6, the evaluation results corresponding to Classification Results 1 to 2 are F1 to F2 respectively, where F2 > F1, then e6 is assigned to Classification Result 2, and if F1 > F2, then e6 is assigned to Classification Result 1.
[0176] In one implementation, the method further includes, in the case that there are unassigned account nodes in the first account node and the second account node, assigning the unassigned account nodes to the n-th classification result to obtain the assignment result corresponding to the n-th classification result. When there are unassigned account nodes in the two account nodes of the resource interaction event, the unassigned account nodes are assigned to the classification result where the resource interaction event is located, which can distinguish the mutually related account nodes and resource interaction events and assign them to the same classification result as much as possible, ensuring that the account nodes and resource interactions within the classification result form a relatively complete local subgraph, maintaining the connectivity and integrity of the graph structure within the partition, enabling better information dissemination and interaction between the nodes within the partition, facilitating model training, and reducing the resource interaction events across partitions and the communication overhead during the training process.
[0177] For example, both account nodes j1 and j2 connected to e6 are unassigned. After determining that e6 is assigned to Classification Result 2, the account nodes j1 and j2 are also assigned to Classification Result 2.
[0178] There will be various situations regarding the number of shared nodes among two account nodes on different resource interaction events, and there will also be various situations regarding whether there are unassigned nodes among two account nodes on different resource interaction events. To accelerate the partitioning rate of the resource interaction graph, different allocation schemes can also be adopted based on different situations. In addition to the schemes based on S403 to S405 and the above-mentioned scheme of assigning unassigned account nodes to the classification result where the resource interaction event is located, in one implementation, the method further includes:
[0179] When the first account node and the second account node do not belong to the shared nodes, and the first account node and the second account node are assigned to the nth classification result, the xth resource interaction event is assigned to the nth classification result to obtain the assignment result corresponding to the nth classification result;
[0180] When the first account node is assigned to the nth classification result, the second account node is a shared node, and the first account node does not belong to the shared node, the xth resource interaction event is assigned to the nth classification result;
[0181] When the second account node is assigned to the nth classification result, the first account node is a shared node, and the second account node does not belong to the shared node, the xth resource interaction event is assigned to the nth classification result.
[0182] For example, if the two account nodes j1 and j2 connected by e6 are not shared nodes, and the account nodes j1 and j2 have both been assigned to classification result 2, then e6 is also assigned to classification result 2.
[0183] For another example, among the two account nodes j1 and j2 connected by e6, j1 is a shared node, j2 is not a shared node, and j2 has been assigned to classification result 1, then e6 is also assigned to classification result 1.
[0184] In the above technical solution, when the two account nodes of the resource interaction event are both assigned to the same classification result, the resource interaction event is also assigned to the classification result where the two account nodes are located, which can reduce the resource interaction events across partitions and reduce the communication overhead during the training process. When any one of the two account nodes of the resource interaction event is a shared node and the other node is a non-shared node and is assigned, the resource interaction event is divided into the classification result where the non-shared node is located, which can ensure the connectivity between the account nodes and the resource interaction events in the classification result where the non-shared node is located, reduce the number of cross-region resource interaction events, and avoid excessive resource interaction events in the classification result where the shared node is located, which helps to balance the load between different classification results.
[0185] It should be noted that for the isolated account nodes in the resource interaction graph, the electronic device will traverse the isolated account nodes in chronological order, and for each isolated account node, execute S403 to S405. Only at this time, the weight parameter corresponding to the isolated account node in the nth classification result obtained by the electronic device is used, and the sum result between the weight parameter and the preset value is determined. Then, according to the product result between the sum result and the load balancing coefficient corresponding to the nth classification result, when it comes to the isolated account node, the nth evaluation result corresponding to the nth classification result is determined. That is, transform formula 6 into S(e,p n ) = [Att(q,t,pn ) + 1]·F BAL (p n ), where q at this time represents the currently traversed isolated account node.
[0186] In an application scenario, combined with Figures 3 to 7 , the partitioning process of the resource interaction graph can be represented by the following innovative neighborhood-aware load balancing graph partitioning algorithm, where the neighborhood-aware load balancing graph partitioning algorithm represents the part of the data processing method in the embodiments of the present application for partitioning the resource interaction graph.
[0187] Innovative neighborhood-aware load balancing graph partitioning algorithm:
[0188] Input: edge set E, node set V, preset proportionality coefficient k, number of partitions m; where the number of partitions m is the total number of multiple classification results, the edges in the edge set E represent resource interaction events in the resource interaction graph, and the node set V includes account nodes in the resource interaction graph;
[0189] Output: node partition P N , edge partition P E and shared node set H; the node partition P N includes the account nodes in the resource interaction graph assigned to each partition, the edge partition P E includes the edges in the resource interaction graph assigned to each partition, and the shared node set H includes the determined shared nodes;
[0190] Calculate the degree D(v) of each (account) node;
[0191] For each node v ∈ V, H ← topk(D(v), k·|V|), where |V| represents the total number of account nodes in the resource interaction graph;
[0192] Initialize partition P N and P E , determine that the account nodes other than the shared nodes are unassigned. Then, for each resource model, perform the following loop operation:
[0193] For each edge e(u, v, t) ∈ E, if P(u) ≠ None and P(v) ≠ None and P(u) ≠ P(v), that is, the node u connected by the edge (u, v, t) has been assigned, the node v has also been assigned, and the partitions of the nodes u and v are different, enter the next judgment:
[0194] If u ∈ H and then P E (e) ← P N (v);
[0195] If v ∈ H and then PE (e) ← P N (u), that is, when there is a shared node between nodes u and v, the edge e(u, v, t) is assigned to the partition where the non-shared node is located.
[0196] Or, if P E (e) = None, then that is, if the edge e remains unassigned after the above judgment, the edge e is assigned to the partition with the largest corresponding evaluation result;
[0197] If P N (u) = None, then P N (u) ← P E (e);
[0198] If P N (v) ≠ None, then P N (v) ← P e (e), that is, the unassigned nodes in the edge e are assigned to the partition where the edge e is located;
[0199] Repeat the above process until the loop ends, and the algorithm returns P N 、P E and H, and the electronic device obtains the nodes, edges, and shared nodes in each partition.
[0200] After determining the partition to which each edge or node belongs, the electronic device updates the number of nodes, the number of edges, and the maximum timestamp stored under the corresponding partition, and updates the number of neighbor nodes and the timestamp of neighbor events of the nodes under the corresponding partition, so as to determine the evaluation result corresponding to the partition subsequently.
[0201] After the electronic device determines the nodes and edges assigned to each partition, based on the indexes of the nodes and the indexes of the edges, it reads the features of the nodes and the features of the edges corresponding to the indexes from the resource interaction graph (that is, the input edge set and node set), and creates corresponding memory units for the nodes in each partition, so as to perform training of the sample model subsequently.
[0202] S406. Train the sample model based on the allocation results corresponding to multiple classification results to obtain the target model. For details of S406, please refer to Figure 3 S104 in the illustrated embodiment, which will not be elaborated here.
[0203] In the embodiment of the present application, when the x-th resource interaction event does not meet the preset condition, the evaluation results determined by the weight parameter and the load balancing coefficient are used to comprehensively reflect the load conditions in each classification result and the influence of the existing nodes in each classification result on the two account nodes of the x-th resource interaction event. In this way, after the x-th resource interaction event is assigned to the classification result with the largest evaluation result among multiple classification results, it can be ensured that there are nodes closely related to the account nodes of the x-th resource interaction event in the classification result after the x-th resource interaction event is assigned, so that the closely related events and nodes can be assigned to the same classification result. Furthermore, the cross-machine communication overhead can be reduced during training. And the partitioning of the resource interaction graph with load balancing can be realized, avoiding the situation of load imbalance, thereby improving the training efficiency. When the x-th resource interaction event meets the preset condition, the classification result to which the x-th resource interaction event should be assigned is determined based on the classification results already assigned to the two account nodes in the x-th resource interaction event, realizing the rapid partitioning of the resource interaction graph.
[0204] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device 6 of this embodiment includes: at least one processor 60 ( Figure 8 only one is shown in the figure), a processor, a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the steps in any of the above-mentioned data processing method embodiments are implemented.
[0205] The electronic device 6 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 8 this is only an example of the electronic device 6 and does not constitute a limitation on the electronic device 6. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0206] The so-called processor 60 may be a Central Processing Unit (CPU), and the processor 60 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0207] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as the hard disk or memory of the electronic device 6. In other embodiments, the memory 61 may also be an external storage device of the electronic device 6, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 6. Further, the memory 61 may also include both the internal storage unit and the external storage device of the electronic device 6. The memory 61 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 61 may also be used to temporarily store data that has been output or will be output.
[0208] Corresponding to the data processing method described in the above embodiments, Figure 9 The structural block diagram of the data processing device provided by the embodiments of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0209] Referring to Figure 9 , the device includes:
[0210] A first acquisition module 100, configured to acquire a resource interaction graph corresponding to the data to be trained. The data to be trained is used to indicate the resource interaction situation between multiple accounts. The resource interaction graph includes multiple account nodes, where there are resource interaction events between at least two account nodes, and the multiple account nodes include shared nodes. The resource interaction graph corresponds to multiple classification results, and each classification result includes a shared node;
[0211] A second acquisition module 200, configured to acquire load balancing coefficients respectively corresponding to multiple classification results, and acquire weight parameters respectively corresponding to account nodes in the multiple classification results, where the load balancing coefficient represents the load condition corresponding to the classification result. Among them, the nth weight parameter corresponding to the ith account node is used to indicate that the ith account node is related to neighbor nodes in the nth classification result, and there is a resource interaction event between the neighbor nodes and the ith account in the resource interaction graph. i and n are positive integers;
[0212] A determination module 300, configured to determine distribution results respectively corresponding to the multiple classification results based on the multiple load balancing coefficients and the multiple weight parameters. The distribution results include the account nodes allocated to the classification results and the resource interaction events related to the account nodes;
[0213] A training module 400, configured to train a sample model based on the distribution results respectively corresponding to the multiple classification results to obtain a target model, where the target model is used to identify abnormal resource interaction events or abnormal account nodes.
[0214] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought can be specifically referred to in the method embodiment part, and will not be elaborated here.
[0215] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0216] The embodiment of the present application also provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.
[0217] The embodiment of the present application provides a computer program product, and when the computer program product runs on an electronic device, it enables the electronic device to implement the steps in the foregoing method embodiments when executed.
[0218] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0219] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0220] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0221] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0222] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0223] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining a resource interaction graph corresponding to the data to be trained, where the data to be trained is used to indicate the resource interaction situation among multiple accounts. The resource interaction graph includes multiple account nodes, and there is at least one resource interaction event between at least two account nodes. The multiple account nodes include shared nodes, and the resource interaction graph corresponds to multiple classification results, and each classification result includes the shared nodes; Obtaining the load balancing coefficients corresponding to the multiple classification results respectively, and obtaining the weight parameters corresponding to the account nodes in the multiple classification results respectively. The load balancing coefficient represents the load situation corresponding to the classification result. Among them, the nth weight parameter corresponding to the ith account node is used to indicate that the ith account node is related to the neighbor nodes in the nth classification result, and there is a resource interaction event between the neighbor nodes and the ith account in the resource interaction graph. i and n are positive integers; Based on the multiple load balancing coefficients and multiple weight parameters, determining the allocation results corresponding to the multiple classification results respectively. The allocation results include the account nodes allocated to the classification results and the resource interaction events related to the account nodes; Training a sample model based on the allocation results corresponding to the multiple classification results respectively to obtain a target model, where the target model is used to identify abnormal resource interaction events or abnormal account nodes.
2. The method according to claim 1, characterized in that, The obtaining the load balancing coefficients corresponding to the multiple classification results respectively includes: Obtaining the first node number of the account nodes in the nth classification result and the first event number of the resource interaction events; Obtaining the timestamp parameters and event numbers of the resource interaction events corresponding to the multiple classification results except the nth classification result respectively, and the target timestamp of the resource interaction events in the nth classification result. The timestamp parameters include a first timestamp and a second timestamp, and the event numbers include a second event number and a third event number; Determining a first ratio result between the result number of the multiple classification results and the total number of the account nodes in the resource interaction graph, and based on a second ratio result between the first ratio result and the first node number, determining the node load coefficient of the nth classification result; Determining the event load coefficient of the nth classification result according to the difference among the second event number, the third event number and the first event number; Determining the time load coefficient of the nth classification result according to the difference among the first timestamp, the second timestamp and the target timestamp; Determining the load balancing coefficient corresponding to the nth classification result according to the product result among the node load coefficient, the event load coefficient and the time load coefficient.
3. The method according to claim 1 or 2, characterized in that, The obtaining the weight parameters corresponding to the account nodes in the multiple classification results respectively includes: Obtain the target neighbor nodes of the \(i\)-th account node and the third timestamp of the \(i\)-th account node in the \(n\)-th classification result, as well as the fourth timestamp corresponding to the resource interaction graph. There is a resource interaction event between the target neighbor nodes and the \(i\)-th account node in the resource interaction graph, and the existing resource interaction events are assigned to the \(n\)-th classification result. The fifth timestamp of the target neighbor nodes is included in the \(n\)-th classification result; Determine the difference result between the third timestamp and the fifth timestamp, and determine the weight parameter corresponding to the \(i\)-th account node in the \(n\)-th classification result based on the product result of the difference result and the reciprocal of the fourth timestamp.
4. The method according to claim 3, wherein The resource interaction graph includes an unassigned \(x\)-th resource interaction event. The \(x\)-th resource interaction event corresponds to a first account node and a second account node. The first account node corresponds to a first weight parameter in the \(n\)-th classification result, and the second account node corresponds to a second weight parameter in the \(n\)-th classification result, where \(x\) is a positive integer; Determining the allocation results corresponding to the multiple classification results based on multiple load balancing coefficients and multiple weight parameters includes: When the \(x\)-th resource interaction event does not meet the preset conditions, determine the sum result between the first weight parameter and the second weight parameter. The preset conditions include that the first account node and the second account node are not the shared nodes, and the first account node and the second account node are assigned to the same classification result, or one of the first account node and the second account node is the shared node and the other is assigned to a classification result; Determine the \(n\)-th evaluation result corresponding to the \(n\)-th classification result according to the product result of the sum result and the load balancing coefficient corresponding to the \(n\)-th classification result; When the \(n\)-th evaluation result meets the first condition, assign the \(x\)-th resource interaction event to the \(n\)-th classification result to obtain the allocation result corresponding to the \(n\)-th classification result. The first condition is that the \(n\)-th evaluation result is greater than the maximum value among the evaluation results corresponding to the multiple classification results.
5. The method according to claim 4, characterized in that, The method further includes: When there is an unassigned account node among the first account node and the second account node, assign the unassigned account node to the \(n\)-th classification result to obtain the allocation result corresponding to the \(n\)-th classification result.
6. The method according to claim 4, characterized in that, The method further includes: When the first account node and the second account node do not belong to the shared nodes and the first account node and the second account node are assigned to the \(n\)-th classification result, assign the \(x\)-th resource interaction event to the \(n\)-th classification result to obtain the allocation result corresponding to the \(n\)-th classification result; When the first account node is assigned to the nth classification result and the second account node is the shared node and the first account node does not belong to the shared node, assign the xth resource interaction event to the nth classification result; When the second account node is assigned to the nth classification result and the first account node is the shared node and the second account node does not belong to the shared node, assign the xth resource interaction event to the nth classification result.
7. A data processing device, characterized in that, Comprising: A first acquisition module, configured to acquire a resource interaction graph corresponding to the data to be trained, the data to be trained being used to indicate resource interaction situations among multiple accounts, the resource interaction graph including multiple account nodes, wherein there are resource interaction events between at least two account nodes, the multiple account nodes including a shared node, the resource interaction graph corresponding to multiple classification results, and each classification result including the shared node; A second acquisition module, configured to acquire the load balancing coefficients respectively corresponding to the multiple classification results, and acquire the weight parameters respectively corresponding to the account nodes in the multiple classification results, the load balancing coefficient indicating the load situation corresponding to the classification result, wherein the nth weight parameter corresponding to the ith account node is used to indicate that the ith account node is related to neighbor nodes in the nth classification result, and there are resource interaction events between the neighbor nodes and the ith account in the resource interaction graph, and i and n are positive integers; A determination module, configured to determine the assignment results respectively corresponding to the multiple classification results based on the multiple load balancing coefficients and the multiple weight parameters, the assignment results including the account nodes assigned to the classification results and the resource interaction events related to the account nodes; A training module, configured to train a sample model based on the assignment results respectively corresponding to the multiple classification results to obtain a target model, the target model being used to identify abnormal resource interaction events or abnormal account nodes.
8. An electronic device, characterized in that, Comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the data processing method according to any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the data processing method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, Including a computer program, when the computer program is run, the data processing method according to any one of claims 1-6 is executed.