A method, device, equipment and storage medium for determining estimated values ​​of neighborhood edges

By introducing parameter servers, distributed storage and parallel data processing are implemented, the problem of driving single-point network bottleneck in the Spark platform is solved, and computing performance is improved and computing resources is saved.

CN114329081BActive Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111075473.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-14
Publication Date
2025-05-06
Estimated Expiration
2041-09-14

AI Technical Summary

Technical Problem

The neighborhood edge estimation method based on the Spark platform is prone to encounter driver single-point network bottlenecks, resulting in limited performance under large-scale data volumes.

Method used

The parameter server is introduced to realize distributed storage of large-scale data and supports parallel acquisition and update of data. Each executor sends update information to the parameter server. The parameter server determines the node's read counter based on the update information.

Benefits of technology

Eliminate driver single-point network bottlenecks, improve computing performance, save computing resources, and avoid data shuffling operations and increase network overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114329081B_ABST
    Figure CN114329081B_ABST
Patent Text Reader

Abstract

The present application discloses a method for determining the estimated value of the number of neighborhood edges in the field of big data and graph computing, which can be applied to the field of message processing, and specifically includes: obtaining the data of the adjacency table of a directed graph, the data of the adjacency table of the directed graph being stored on at least two executors; generating a node aggregation table according to the data of the adjacency table of the directed graph; in the first round of iterations, for each executor, counting the number of first-order neighborhood edges of each tail node according to the node aggregation table, so that each executor can respectively count the first round of update information of each tail node; sending the update information of each tail node to the parameter server through each executor, so that the parameter server can determine the first-order read counter of the corresponding node according to the number and update information of each tail node, and the first-order read counter stores the estimated value of the first-order neighborhood edges. The present application also provides an apparatus, a device and a medium. The present application can eliminate the problem of the single-point network bottleneck of the driver to improve computing performance and save computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data processing technology, and in particular to a method, device, equipment and storage medium for determining an estimated value of a neighborhood edge number. Background Art

[0002] In recent years, with the rapid development of Internet technology, more and more objects have joined various community networks. A community network can be regarded as a graph, which is a data structure that represents the relationship between a series of objects. The nodes (vertex) in the graph represent objects, and the edges (edge) in the graph represent the relationship between objects. Therefore, estimating the number of neighborhood edges can often provide important reference information for some specific scenarios.

[0003] At present, the HyperLoglog (HLL) algorithm can be used to estimate the number of neighborhood edges based on the Spark platform. The Spark platform is a fast and general computing engine designed for large-scale data processing, which can support interactive computing and more complex algorithms.

[0004] However, the solution implemented based on the Spark platform is prone to encountering the single-point network bottleneck of the driver. This is because data storage and updates are all performed on the driver. When encountering large-scale data volumes, the storage space and computing power of the single-point driver are difficult to handle, resulting in performance limitations. Summary of the invention

[0005] The embodiment of the present application provides a method, device, equipment and storage medium for determining the estimated value of the number of neighborhood edges. The parameter server introduced in the present application can not only realize the distributed storage of large-scale data, but also support operations such as parallel acquisition and update of data, thereby eliminating the problem of driver single-point network bottleneck, improving computing performance and saving computing resources.

[0006] In view of this, the present application provides, on one hand, a method for determining an estimated value of the number of neighborhood edges, comprising:

[0007] Obtaining directed graph adjacency table data, wherein the directed graph adjacency table data represents adjacency table data corresponding to a directed graph, the directed graph includes K nodes, each piece of data in the directed graph adjacency table data includes a node number of a source node and a node number of a tail node corresponding to a directed edge, the directed graph adjacency table data is stored on at least two executors, and K is an integer greater than 1;

[0008] Generate a node aggregation table according to the directed graph adjacency table data, wherein the node aggregation table includes a node number of a source node, a node number set of a tail node, and an edge number set, the node number set includes at least one node number, and the edge number set includes at least one edge number;

[0009] For each executor, the number of neighboring edges of each tail node is counted according to the node aggregation table, so that each executor can separately count and obtain the first round of update information of each tail node;

[0010] Each executor sends the first-round update information of each tail node to the parameter server, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first-round update information of each tail node, wherein the first-order read counter stores the estimated value of the first-order neighborhood edge number.

[0011] Another aspect of the present application provides a device for determining an estimated value of the number of neighborhood edges, comprising:

[0012] An acquisition module is used to acquire directed graph adjacency table data, wherein the directed graph adjacency table data represents adjacency table data corresponding to a directed graph, the directed graph includes K nodes, each piece of data in the directed graph adjacency table data includes a node number of a source node and a node number of a tail node corresponding to a directed edge, the directed graph adjacency table data is stored on at least two executors, and K is an integer greater than 1;

[0013] A generating module, configured to generate a node aggregation table according to the directed graph adjacency table data, wherein the node aggregation table includes a node number of a source node, a node number set of a tail node, and an edge number set, wherein the node number set includes at least one node number, and the edge number set includes at least one edge number;

[0014] A statistics module is used to count the number of neighboring edges of each tail node according to the node aggregation table for each executor, so that each executor can respectively count and obtain the first round of update information of each tail node;

[0015] A sending module is used to send the first round update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round update information of each tail node, wherein the first-order read counter stores the estimated value of the first-order neighborhood edge number.

[0016] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0017] An acquisition module is specifically used to acquire undirected graph adjacency table data, wherein the undirected graph adjacency table data represents the adjacency table data corresponding to the undirected graph, the undirected graph includes K nodes, and each forward data in the undirected graph adjacency table data includes a node number of a source node and a node number of a tail node;

[0018] For each forward data in the undirected graph adjacency table data, generate reverse data corresponding to each forward data, wherein the source node in the reverse data is the tail node in the corresponding forward data, and the tail node in the reverse data is the source node in the corresponding forward data;

[0019] Directed graph adjacency table data is generated according to each piece of forward data and the reverse data corresponding to each piece of forward data, wherein both the forward data and the reverse data belong to the data in the directed graph adjacency table data.

[0020] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0021] A generation module is specifically used to encode the edge corresponding to each data according to the directed graph adjacency table data to obtain the edge number corresponding to each edge;

[0022] Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, and the edge number;

[0023] According to the node and edge relationship table, the node numbers of the same source node are aggregated to obtain a node aggregation table.

[0024] In one possible design, in another implementation of another aspect of the embodiment of the present application, each piece of data in the directed graph adjacency table data further includes an edge label corresponding to the edge, and the node aggregation table further includes an edge label set, and the edge label set includes at least one edge label;

[0025] A generation module is specifically used to encode the edge corresponding to each data according to the directed graph adjacency table data to obtain the edge number corresponding to each edge;

[0026] Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, the edge number and the edge label;

[0027] According to the node and edge relationship table, the node numbers of the same source nodes are aggregated to obtain a node aggregation table.

[0028] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0029] A statistical module is specifically used to determine, for each executor, a counter corresponding to each neighbor edge in the neighbor edge set according to the node aggregation table, wherein each neighbor edge corresponds to a counter, and the counter stores the edge number and the cardinality estimate, and the cardinality estimate is 1;

[0030] For each executor, the counter corresponding to each neighborhood edge is fused with the counter of the corresponding tail node to obtain the first round of update information of each tail node.

[0031] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0032] A statistics module is specifically used to determine, for each executor, a counter corresponding to each neighbor edge in the neighbor edge set according to the node aggregation table, wherein each neighbor edge corresponds to a counter, and the counter stores the edge label, the edge number, and the cardinality estimate. When the edge label is a target edge label, the cardinality estimate is 1, and when the edge label is a non-target edge label, the cardinality estimate is 0;

[0033] For each executor, the counter corresponding to each neighborhood edge is fused with the counter of the corresponding tail node to obtain the first round of update information of each tail node.

[0034] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0035] The sending module is specifically used to send the first round of update information of the tail node to the parameter server through each executor after the parameter server creates the counter matrix, so that the parameter server can fuse the zero-order write counter of the corresponding node with the update information of the first round according to the node number of each tail node and the update information of the first round, until the first round of iteration is completed, and the parameter server determines the corresponding first-order read counter according to the first-order write counter of each node in the K nodes stored in the counter matrix, wherein the initial cardinality estimate of the zero-order write counter is 0.

[0036] In one possible design, in another implementation of another aspect of the embodiment of the present application, the apparatus for determining a neighborhood edge number estimate further includes a determination module;

[0037] The acquisition module is further used to send the first-round update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first-round update information of each tail node, and then obtains the first-order read counter corresponding to the source node from the parameter server for each executor;

[0038] The determination module is further used to determine, for each executor, a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node;

[0039] The determination module is further used to determine the second round update information corresponding to each tail node by fusing the first-order read counters of all source nodes corresponding to each tail node in the tail node set for each executor;

[0040] The sending module is also used to send the second-round update information of each tail node to the parameter server through each executor, so that the parameter server can fuse the first-order write counter of the corresponding node with the second-round update information according to the node number of each tail node and the second-round update information until the second round of iteration is completed. The parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, wherein the second-order read counter stores the estimated value of the second-order neighborhood edge number.

[0041] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0042] A determination module is specifically used to obtain, for each executor, a first-order read counter of a source node corresponding to each tail node in the tail node set;

[0043] For each executor, the first-order read counters of the source nodes corresponding to each tail node are fused to obtain the second round update information corresponding to each tail node.

[0044] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0045] The sending module is specifically used to send the second-round update information of each tail node to the parameter server through each executor, so that the parameter server fuses the first-order write counter of the same tail node with the second-round update information according to the second-round update information of each tail node and the node number of each tail node to obtain the second-order write counter, until the second round of iteration is completed, and the second-order write counter of each node in the K nodes of the parameter server determines the corresponding second-order read counter.

[0046] In one possible design, in another implementation of another aspect of the embodiment of the present application,

[0047] The acquisition module is further used to send the second-round update information of each tail node to the parameter server through each executor, so that the parameter server performs a fusion process on the first-order write counter of the corresponding node and the second-round update information according to the number of each tail node and the second-round update information, until the second round of iteration is completed. After the parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, for each executor, the second-order read counter corresponding to the source node is obtained from the parameter server;

[0048] The determination module is further used to determine, for each executor, a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node;

[0049] The determination module is further used to determine the third round of update information corresponding to each tail node by fusing the second-order read counters of all source nodes corresponding to each tail node in the tail node set for each executor;

[0050] The sending module is also used to send the third-round update information of each tail node to the parameter server through each executor, so that the parameter server can fuse the second-order write counter of the corresponding node with the third-round update information according to the number of each tail node and the third-round update information until the third round of iteration is completed. The parameter server determines the corresponding third-order read counter according to the third-order write counter of each node in the K nodes, wherein the third-order read counter stores the estimated value of the third-order neighborhood edge number.

[0051] In one possible design, in another implementation of another aspect of the embodiment of the present application, the apparatus for determining a neighborhood edge number estimate further includes a receiving module and a sending module;

[0052] The acquisition module is specifically used to receive a data upload instruction for a data upload control sent by a terminal device, wherein the data upload control is displayed on an interactive interface provided by the terminal device;

[0053] Obtain directed graph adjacency table data according to the data upload instruction;

[0054] A receiving module, configured to receive a label selection instruction for at least one edge label sent by a terminal device, wherein at least one selectable edge label is displayed on an interactive interface provided by the terminal device;

[0055] A determination module, used for determining a target edge label according to a label selection instruction;

[0056] The receiving module is further used to receive an order setting instruction for an order adjustment control sent by a terminal device, wherein the order adjustment control is displayed on an interactive interface provided by the terminal device;

[0057] The determination module is further used to determine the target order according to the order setting instruction;

[0058] The receiving module is further used to receive a node viewing instruction for any node sent by the terminal device, wherein at least one node is displayed on the interactive interface provided by the terminal device;

[0059] A sending module, used for sending the estimated value of the number of first-order neighboring edges corresponding to any node to the terminal device if the node viewing instruction carries the first-order identifier, so that the terminal device displays the estimated value of the number of first-order neighboring edges corresponding to any node;

[0060] The sending module is also used to send the estimated value of the second-order neighboring edges corresponding to any node to the terminal device if the node viewing instruction carries a second-order identifier, so that the terminal device displays the estimated value of the second-order neighboring edges corresponding to any node.

[0061] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer is enabled to execute the above-mentioned methods.

[0062] Another aspect of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided by the above aspects.

[0063] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0064] In an embodiment of the present application, a method for determining the estimated value of the number of neighborhood edges is provided. First, the directed graph adjacency table data is obtained, and the directed graph adjacency table data is stored on at least two executors. Then, a node aggregation table can be generated according to the directed graph adjacency table data. Based on this, for each executor, the number of neighborhood edges of each tail node can be counted according to the node aggregation table, so that each executor respectively obtains the first round of update information of each tail node. Finally, each executor sends the first round of update information of each tail node to the parameter server respectively, and the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round of update information of each tail node. Here, the first-order read counter stores the first-order neighborhood edge number estimate. Through the above method, the introduction of the parameter server can not only realize the distributed storage of large-scale data, but also support operations such as parallel acquisition of data and update of data. Therefore, on the one hand, the problem of driver single-point network bottleneck can be eliminated and the computing performance can be improved. On the other hand, there is no need to create and store elastic distributed data sets (RDD) for the update of node count information, thereby saving computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 A schematic diagram of an environment of a system for determining a neighborhood edge number estimation value in an embodiment of the present application;

[0066] Figure 2 A schematic diagram of the architecture of a system for determining a neighborhood edge number estimation value in an embodiment of the present application;

[0067] Figure 3 A schematic diagram of a flow chart of a method for determining an estimated value of the number of neighborhood edges in an embodiment of the present application;

[0068] Figure 4 A schematic diagram of the conversion relationship between a directed graph and an adjacency list in an embodiment of the present application;

[0069] Figure 5 A schematic diagram of converting an undirected graph into a directed graph in an embodiment of the present application;

[0070] Figure 6 A schematic diagram of a directed graph in an embodiment of the present application;

[0071] Figure 7 Another schematic diagram of a directed graph in an embodiment of the present application;

[0072] Figure 8 A schematic diagram of an interactive interface in an embodiment of the present application;

[0073] Fig. 9 This is another schematic diagram of the interactive interface in the embodiment of the present application;

[0074] Fig.10 This is another schematic diagram of the interactive interface in the embodiment of the present application;

[0075] Fig.11 This is another schematic diagram of the interactive interface in the embodiment of the present application;

[0076] Fig.12 This is another schematic diagram of the interactive interface in the embodiment of the present application;

[0077] Fig.13 A schematic diagram of a device for determining an estimated value of the number of neighborhood edges according to an embodiment of the present application;

[0078] Fig.14 A schematic diagram of the structure of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0079] The embodiment of the present application provides a method, device, equipment and storage medium for determining the estimated value of the number of neighborhood edges. The parameter server introduced in the present application can not only realize the distributed storage of large-scale data, but also support operations such as parallel acquisition and update of data, thereby eliminating the problem of driver single-point network bottleneck, improving computing performance and saving computing resources.

[0080] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0081] Graph theory is a discipline that studies the relationship between elements of a set. People often use points on a plane (or space) to represent elements, and use line segments (or arcs) connecting points to represent certain relationships between elements. In short, a graph is composed of nodes (vertices) and edges (edges).

[0082] Graphs can be divided into directed graphs and undirected graphs. Directed graphs have an indicated direction, while undirected graphs do not. Exemplarily, taking a social network as an example, assuming that the social network is an undirected graph, the nodes in the graph represent users, and the undirected edges in the graph represent friend relationships. For example, node A points to node B, that is, the user corresponding to node A and the user corresponding to node B are friends with each other. Exemplarily, taking a transaction network as an example, assuming that the transaction network is a directed graph, the nodes in the graph represent user accounts, and the directed edges in the graph represent transfer relationships. For example, node A points to node B, that is, the user account corresponding to node A initiates a transfer to the user account corresponding to node B.

[0083] Graphs can also be divided into directed graphs, homogeneous graphs and heterogeneous graphs, where homogeneous graphs indicate that there is only one type of node and edge in the data, while heterogeneous graphs may have more than one type of node and edge. Exemplarily, taking a graph as a transaction network as an example, assuming that the transaction network is a homogeneous graph, the nodes in the graph represent user accounts, and the directed edges in the graph represent transfer relationships. Exemplarily, taking a graph as a transaction network as an example, assuming that the transaction network is a heterogeneous graph, the nodes in the graph represent user accounts, and the directed edges in the graph represent legal transfer relationships or illegal transfer relationships.

[0084] In the financial risk control scenario of electronic payment, the number of neighborhood edges of each node in the graph can be estimated to understand the number of illegal transactions. The edges in the financial risk control scenario represent the transfers between the two parties, and the estimated value of the number of neighborhood edges is the estimated number of transfers in the neighborhood. If the edge is labeled, for example, the label is "illegal transfer", then the number of illegal transfers is estimated, which can provide important information in the financial risk control scenario. It can be understood that the above scenario is only an illustration. In actual applications, other scenarios may also be involved, such as social scenarios.

[0085] For example, take the Internet of Vehicles as an example, assuming that the Internet of Vehicles is a homogeneous graph, the nodes in the graph represent vehicles, and the directed edges in the graph represent the trip sharing relationship between vehicles. In the Internet of Vehicles scenario, the number of neighboring edges of each node in the graph can be estimated to understand the frequency of sharing. As a result, important information can be provided in the Internet of Vehicles scenario, for example, the same or similar music, radio stations, songs, videos, audios, advertisements, audiobooks or travel can be pushed to vehicles that share frequently.

[0086] In order to efficiently and accurately estimate the number of neighboring edges of a node in the above scenario, this application proposes a method for determining the estimated value of the number of neighboring edges. The method is applied to Figure 1The neighborhood edge number estimation value determination system shown in the figure, as shown in the figure, the neighborhood edge number estimation value determination system may include a server and a terminal device, and the client is deployed on the terminal device, wherein the client can be run on the terminal device in the form of a browser, or can be run on the terminal device in the form of an independent application (application, APP), etc., and the specific presentation form of the client is not limited here. Users can transfer money, trade, and add friends with other users through the client. The server involved in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal device includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc., but is not limited to this. The terminal device and the server can be directly or indirectly connected by wired or wireless communication, and this application is not limited here. The number of servers and terminal devices is also not limited. The solution provided in this application can be completed independently by a terminal device, can be completed independently by a server, or can be completed by a terminal device and a server in cooperation. This application does not make any specific limitations on this.

[0087] based on Figure 1The neighborhood edge number estimation value determination system shown in FIG. takes the financial risk control scenario as an example. User A uses the client to transfer money to user B, thereby generating a transfer record, which includes the account of user A and the account of user B. The transfer record can be represented as two nodes and an edge in the figure. Since there are a large number of transfer records in the actual financial risk control scenario, the processing of these data can be realized through cloud computing. Among them, cloud computing refers to the delivery and use mode of information technology (IT) infrastructure, which refers to obtaining the required resources through the network in an on-demand and easily scalable manner; generalized cloud computing refers to the delivery and use mode of services, which refers to obtaining the required services through the network in an on-demand and easily scalable manner. This service can be IT and software, Internet-related, or other services. Cloud computing is the product of the integration of the development of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing.

[0088] With the development of the Internet, real-time data streams, and the diversification of connected devices, as well as the demand for search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from the previous parallel distributed computing, the emergence of cloud computing will promote revolutionary changes in the entire Internet model and enterprise management model from a conceptual perspective.

[0089] In addition, the present application may also involve the processing of big data, where big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time frame. It is a massive, high-growth, and diversified information asset that requires new processing models to have stronger decision-making power, insight discovery, and process optimization capabilities. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process large amounts of data within a tolerable elapsed time. Technologies applicable to big data include large-scale parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.

[0090] For further understanding, please refer to Figure 2 , Figure 2This is a schematic diagram of the architecture of a system for determining the estimated value of the number of neighborhood edges in an embodiment of the present application. As shown in the figure, the present application provides a method for determining the estimated value of the number of neighborhood edges based on a parameter server (PS) that can be implemented on a distributed computing engine. Specifically, different executors process different parts of the graph, that is, process the partitioned graph. In step S1, if an undirected graph is obtained, the undirected graph needs to be converted into a directed graph. If a directed graph is obtained, subsequent processing is performed, that is, encoding each edge in the directed graph. In step S2, the parameter server creates and initializes the counter matrix of the node. In step S3, each executor pulls the data of the read counter (Read-Counters) in the parameter server. In step S4, each executor calculates the update information (message-counter) corresponding to the node. In step S5, after completing the message-counter calculation of a batch, each executor pushes the message-counter corresponding to the node of the batch to the partition corresponding to the parameter server, and the parameter server can update the write counter (Write-Counters).

[0091] Before introducing the embodiments of the present application, for ease of understanding, the relevant terms involved in the present application will be introduced below.

[0092] (1) r-order neighborhood vertex set: In the graph G(V,E), V represents the vertex set and E represents the edge set. For any two vertices x, y∈V, the shortest path length from x to y is denoted as d(x,y). Then, the r-order (r≥1) neighborhood vertex set of node x is defined as N r (x) = {y|y∈V; d(y,x)≤r}, that is, the set of vertices in graph G whose shortest path to node x is no greater than r, including node x itself.

[0093] (2) r-order neighborhood edge set: Based on the definition of the r-order neighborhood vertex set, the r-order (r ≥ 1) neighborhood edge set of node x is defined as ε r (x)={(m,n)|m,n∈N r (x); (m, n) ∈ E; d (m, x) ≠ d (n, x) = r}.

[0094] (3) r-order: The order is used to define the size of a node’s neighborhood. The r-order is the maximum value of the shortest distance from any node in the r-order neighborhood of node x to node x.

[0095] (4) Cardinality: It refers to the number of non-repeating elements in a set. Estimating the number of order-neighborhood edges of a vertex is to estimate the cardinality of its order-neighborhood edge set. Mathematically speaking, the detailed description of the cardinality estimation problem is that for a data stream {x1, x2, ..., xs}, it may have repeated elements. Let n represent the number of different elements of this data stream, and this set can be represented as {e1, ..., en}.

[0096] (5) HyperLogLog: It is a probabilistic set cardinality estimation algorithm. Cardinality estimation is to estimate the number of non-repeating elements in a batch of data.

[0097] (6) Parameter Server (PS): is a programming framework for solving distributed machine learning problems. The framework mainly includes a server, a client, and a scheduler. The main function of the server is to store the parameters of the computing task, receive update information from the client, and update the local parameters. The main functions of the client are twofold: one is to obtain the latest parameters from the server, and the other is to use the data of the local or remote node and the parameters obtained from the server to calculate the update information and finally send the update information to the server. The main function of the scheduler is to manage the server and client nodes, complete data synchronization between nodes, and add and delete nodes.

[0098] (7) Spark: is a fast and general computing engine designed for large-scale data processing.

[0099] (8) Distributed computing engine: It is a high-performance distributed computing platform that combines the parameter server function with Spark's large-scale data processing capabilities, supporting traditional machine learning, deep learning, and various graph algorithms.

[0100] (9) Source node: indicates the node to which the arrow points.

[0101] (10) Tail node: represents the node pointed by the arrow.

[0102] (11) Incoming edge: refers to the edge pointing to a vertex.

[0103] (12) Out-edge: refers to the edge pointing out from a vertex.

[0104] (13) Parameter r: represents the neighborhood order, for example, the rth order.

[0105] (14) Parameter i: represents the round, for example, the i-th round, where parameter i is less than or equal to parameter r.

[0106] (15) Parameter D v : The set of tail nodes of node v, that is, the set of tail nodes pointed to by directed edges with node v as the source node.

[0107] (16) Parameter l e : represents the edge label of edge e.

[0108] (17) Parameter L: represents the target edge label set.

[0109] (18) The read counter of node v in round i.

[0110] (19) Write counter of node v in round i.

[0111] (20)m v (i): Update information of node v in the i-th round.

[0112] Combined with the above introduction, the following will introduce the method for determining the estimated value of the number of neighborhood edges in this application. Please refer to Figure 3 , an embodiment of the method for determining the estimated value of the number of neighborhood edges in the embodiment of the present application includes:

[0113] 110. Obtain directed graph adjacency table data, wherein the directed graph adjacency table data represents adjacency table data corresponding to a directed graph, the directed graph includes K nodes, each piece of data in the directed graph adjacency table data includes a node number of a source node and a node number of a tail node corresponding to a directed edge, the directed graph adjacency table data is stored on at least two executors, and K is an integer greater than 1;

[0114] In one or more embodiments, directed graph adjacency table data is obtained, and directed graph adjacency table data refers to adjacency table data corresponding to the directed graph, wherein each piece of data includes at least a source node and a tail node. If the directed graph adjacency table data also carries an edge label, then each piece of data may also include an edge label. In practical applications, the directed graph adjacency table data is distributedly stored in at least two executors, and each executor has the functions of data storage, data calculation, and data communication.

[0115] For ease of explanation, this application uses the order as the basis for distinguishing node types. For example, "source node" refers to the source node defined when counting the estimated value of the first-order neighborhood edge count, and "tail node" refers to the tail node defined when counting the estimated value of the first-order neighborhood edge count. For another example, "source node" may refer to the source node defined when counting the estimated value of the second-order neighborhood edge count, and "tail node" may refer to the tail node defined when counting the estimated value of the second-order neighborhood edge count.

[0116] Specifically, for ease of understanding, see Figure 4 , Figure 4 Schematic diagram of the conversion relationship between the directed graph and the adjacency list in the embodiment of the present application. Figure 4 As shown in Figure (A), take the 7 nodes (i.e., K=7) shown in the figure as an example, and encode each node. This application adopts the default encoding method of continuous non-negative integer encoding starting from 0, that is, node A is node 0, node B is node 1, and so on. Based on this, construct Figure 4 The adjacency table shown in Figure (B) in the figure. For example, when node 0 is the "source node", node 1 is the "tail node". For example, when node 1 is the "source node", nodes 2, 4 and 5 are all "tail nodes". By analogy, the adjacency table data of the directed graph shown in Table 1 below can be obtained.

[0117] Table 1

[0118] (0,1) (1,2) (1,4) (1,5) (2,4) (3,2) (4,1) (4,3) (5,6)

[0119] Wherein, each data in the directed graph adjacency table data is represented as (source node ID, tail node ID), the source node ID represents the node number of the source node (i.e., node ID), and the tail node ID represents the node number of the tail node (i.e., node ID). It should be noted that Table 1 is only a schematic and should not be construed as a limitation on the present application.

[0120] 120. Generate a node aggregation table according to the directed graph adjacency table data, wherein the node aggregation table includes a node number of a source node, a node number set of a tail node, and an edge number set, the node number set includes at least one node number, and the edge number set includes at least one edge number;

[0121] In one or more embodiments, it is also necessary to use a default encoding method to encode the edge, and the default encoding method is a continuous non-negative integer encoding starting from 0. Among them, the edges between the same two nodes have the same edge ID (i.e., edge ID), with Figure 4 Taking Figure (A) as an example, the edge between node B and node E has the same edge ID as the edge between node E and node B.

[0122] Specifically, according to the node ID of the source node, the node IDs and edge IDs of all tail nodes are aggregated to generate a node aggregation table. Each data in the node aggregation table is represented as (source node ID, Iterator (tail node ID, edge ID)), wherein the source node ID represents the node ID of the source node, the tail node ID represents the node ID of the tail node, and the same source node ID corresponds to at least one tail node ID and at least one edge ID. Optionally, if the directed graph adjacency table data also includes an edge label, each data in the node aggregation table is represented as (source node ID, Iterator (tail node ID, edge ID, edge label)). The node aggregation table can aggregate the relevant information of the source node together, which is convenient for subsequent processing and makes the code implementation easier.

[0123] 130. For each executor, the number of neighboring edges of each tail node is counted according to the node aggregation table, so that each executor can respectively count and obtain the first round of update information of each tail node;

[0124] In one or more embodiments, the first round update information (message-counter) of each tail node may be obtained according to the node aggregation table.

[0125] Specifically, the directed graph adjacency table data is stored on at least two executors, and each executor has the ability to process data independently. That is, the corresponding task is assigned to each executor in advance. For example, executor No. 1 processes the 1st to 100th data in the directed graph adjacency table data, and executor No. 2 processes the 101st to 200th data in the directed graph adjacency table data. Based on this, each executor can count the number of neighborhood edges of each tail node based on the node aggregation table and the assigned data, so as to obtain the first round of update information of each tail node.

[0126] 140. Send the first-round update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first-round update information of each tail node, wherein the first-order read counter stores the estimated value of the first-order neighborhood edge number.

[0127] In one or more embodiments, each executor sends the first round of update information corresponding to each tail node in a batch to the parameter server. Since the parameter server has multiple partitions, the executor needs to push the first round of update information to the specified partition. When the amount of processing is small, each executor can complete the upload of the first round of update information in one batch in the same round. When the amount of processing is large, each executor often needs to upload multiple batches of the first round of update information in the same round. Among them, the executors upload the first round of update information in parallel. For the parameter server, the first round of update information reported by the executor is processed in parallel and in real time.

[0128] Specifically, the parameter server merges the first-round update information of the same node with the first-order write counter based on the first-round update information reported by each executor. For example, when the parameter server receives the first-round update information uploaded by executor No. 1 in the first batch, the first-round update information of the same node is merged with the first-order write counter corresponding to the node. At the same time, when the parameter server receives the first-round update information uploaded by executor No. 2 in the first batch, the first-round update information of the same node is still merged with the first-order write counter corresponding to the node. After a round ends, the parameter server obtains the updated first-order write counter, and then updates the first-order read counter corresponding to the node to the first-order write counter corresponding to the node.

[0129] Since some nodes may not be tail nodes in the first round, but these nodes still have corresponding first-order read counters, the estimated value of the number of first-order neighbor edges corresponding to each node can be determined based on the first-order read counter corresponding to each node in the K nodes.

[0130] The "fusion" involved in this application means fusing two counters. Each counter stores an edge set and an estimated value of the number of neighboring edges, takes the union of the elements in the two counters, and then fuses the two counters into an updated counter, which stores the updated edge set and the updated estimated value of the number of neighboring edges.

[0131] It should be noted that the method provided in the present application can be executed by a computer device, wherein at least two executors are deployed on the computer device. The computer device can be a server or a terminal device, or a system composed of a server and a terminal device, which is not limited here.

[0132] In an embodiment of the present application, a method for determining an estimated value of the number of neighborhood edges is provided. Through the above method, the introduction of a parameter server can not only realize distributed storage of large-scale data, but also support operations such as parallel acquisition of data and update of data. Therefore, on the one hand, it can eliminate the problem of driver single-point network bottleneck and improve computing performance. On the other hand, there is no need to additionally create and store RDD for updating node count information, thereby saving computing resources.

[0133] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, obtaining the adjacency table data of the directed graph may specifically include:

[0134] Obtaining undirected graph adjacency table data, wherein the undirected graph adjacency table data represents adjacency table data corresponding to the undirected graph, the undirected graph includes K nodes, and each forward data in the undirected graph adjacency table data includes a node number of a source node and a node number of a tail node;

[0135] For each forward data in the undirected graph adjacency table data, generate reverse data corresponding to each forward data, wherein the source node in the reverse data is the tail node in the corresponding forward data, and the tail node in the reverse data is the source node in the corresponding forward data;

[0136] Directed graph adjacency table data is generated according to each piece of forward data and the reverse data corresponding to each piece of forward data, wherein both the forward data and the reverse data belong to the data in the directed graph adjacency table data.

[0137] In one or more embodiments, a method for converting an undirected graph to a directed graph is described. For an undirected graph adjacency list, the undirected graph is converted to an equivalent directed graph. Typically, only one directed edge is used to replace a bidirectional edge in the undirected graph, and an additional reverse edge is generated for each edge in the undirected graph. If there are edge labels, the consistency of the edge labels should be maintained.

[0138] Specifically, assume that there is a forward data (0,1) in the undirected graph adjacency table data, where "0" in the forward data represents the node ID of the source node, and "1" represents the node ID of the tail node. In response to this, a corresponding reverse data is generated, and the reverse data is (1,0), where "1" in the reverse data represents the node ID of the source node, and "0" represents the node ID of the tail node. Similarly, after obtaining the reverse data corresponding to each forward data, combined with each forward data in the undirected graph adjacency table data, the directed graph adjacency table data is obtained.

[0139] For easier understanding, see Figure 5 , Figure 5This is a schematic diagram of converting an undirected graph into a directed graph in an embodiment of the present application. As shown in the figure, each edge in the undirected graph becomes two edges after conversion. Assume that node A is node 0 and node B is node 1. Therefore, the undirected edge between node A and node B can be converted into two directed edges, and the two directed edges are represented as (0,1) and (1,0) respectively.

[0140] Secondly, in the embodiment of the present application, a method for converting an undirected graph into a directed graph is provided. Through the above method, considering that in practical applications, many data are presented as undirected graphs. Therefore, combining the characteristics of the graph being undirected and directed, a conversion function of an undirected graph and a directed graph is designed. Thus, the application scenarios in which the input is compatible with both undirected and directed graphs are achieved.

[0141] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, generating a node aggregation table according to the directed graph adjacency table data may specifically include:

[0142] According to the data of the directed graph adjacency table, the edge corresponding to each data is encoded to obtain the edge number corresponding to each edge;

[0143] Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, and the edge number;

[0144] According to the node and edge relationship table, the node numbers of the same source node are aggregated to obtain a node aggregation table.

[0145] In one or more embodiments, a method for generating a node aggregation table without edge labels is introduced. First, each edge can be encoded according to the default encoding method to obtain the edge ID corresponding to each edge. The directionality of the edge is not considered when encoding the edge. As long as the two connected nodes are the same, the same edge ID is used. Next, a node and edge relationship table is generated based on the directed graph adjacency table data and the edge code corresponding to each edge. Finally, combined with the node and edge relationship table, the node numbers of the same source node are aggregated to obtain a node aggregation table. This will be introduced in detail with examples below.

[0146] Specifically, for ease of understanding, see Figure 6 , Figure 6 This is a schematic diagram of a directed graph in an embodiment of the present application. As shown in the figure, each node and each edge are first encoded, where node A is node 0, node B is node 1, node C is node 2, node D is node 3, and node E is node 4. Thus, the directed graph adjacency table data shown in Table 2 below is obtained.

[0147] Table 2

[0148] (0,3) (1,3) (2,1) (2,3) (3,4) (4,0) (4,1)

[0149] Wherein, each data in the directed graph adjacency table data is represented as (source node ID, tail node ID), the source node ID represents the node ID of the source node, and the tail node ID represents the node ID of the tail node. It should be noted that Table 2 is only a schematic and should not be understood as a limitation of the present application. Based on this, in combination with the edge ID corresponding to each edge, a node and edge relationship table is further generated. Thus, a node and edge relationship table as shown in Table 3 below is obtained.

[0150] Table 3

[0151]

[0152]

[0153] Among them, each data in the node and edge relationship table is represented as (source node ID, tail node ID, edge ID), the source node ID represents the node ID of the source node, and the tail node ID represents the node ID of the tail node. It should be noted that Table 3 is only a schematic and should not be understood as a limitation of this application. Based on this, the node numbers of the same source node are aggregated, thereby obtaining a node aggregation table as shown in Table 4 below.

[0154] Table 4

[0155] (0,Iterator(3,0)) (1,Iterator(3,3)) (2,Iterator((1,5),(3,4))) (3,Iterator(4,1)) (4,Iterator((0,2),(1,6)))

[0156] Among them, each data in the node aggregation table is represented as (source node ID, Iterator ((tail node ID, edge ID))), the source node ID represents the node ID of the source node, and the tail node ID represents the node ID of the tail node. It should be noted that Table 4 is only a schematic and should not be understood as a limitation on the present application.

[0157] Secondly, in the embodiment of the present application, a method for generating a node aggregation table without edge labels is provided. Through the above method, the source node is used as the clustering basis, and the node numbers of the same source node are aggregated to obtain a node aggregation table. Based on the node aggregation table, subsequent statistics and fusion can be facilitated, thereby improving the feasibility and operability of the solution.

[0158] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, each piece of data in the directed graph adjacency table data further includes an edge label corresponding to the edge, and the node aggregation table further includes an edge label set, and the edge label set includes at least one edge label;

[0159] Generate a node aggregation table based on the directed graph adjacency table data, which may specifically include:

[0160] According to the data of the directed graph adjacency table, the edge corresponding to each data is encoded to obtain the edge number corresponding to each edge;

[0161] Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, the edge number and the edge label;

[0162] According to the node and edge relationship table, the node numbers of the same source nodes are aggregated to obtain a node aggregation table.

[0163] In one or more embodiments, a method for generating a node aggregation table in the presence of edge labels is introduced. It can be seen from the aforementioned embodiments that each piece of data in the directed graph adjacency table data can also include the edge label corresponding to each edge. Therefore, the node aggregation table also includes an edge label set, and the edge label set includes at least one edge label. First, each edge can be encoded according to the default encoding method to obtain the edge ID corresponding to each edge. Next, a node and edge relationship table is generated based on the directed graph adjacency table data and the edge ID corresponding to each edge. Finally, in combination with the node and edge relationship table, the node numbers of the same source node are aggregated to obtain a node aggregation table. This will be introduced in detail with examples below.

[0164] Specifically, for ease of understanding, see Figure 7 , Figure 7 This is another schematic diagram of a directed graph in an embodiment of the present application. As shown in the figure, each node and each edge are first encoded, where node A is node 0, node B is node 1, node C is node 2, node D is node 3, and node E is node 4. At the same time, each edge has an edge label, and there are two types of edge labels, namely, the edge label of "legal transfer" and the edge label of "illegal transfer". Thus, the directed graph adjacency table data shown in Table 5 below is obtained.

[0165] Table 5

[0166] (0,3,Illegal transfer) (1,3,Legal transfer) (2,1,Legal transfer) (2,3, legal transfer) (3,4, legal transfer) (4,0,Illegal transfer) (4,1, illegal transfer)

[0167] Wherein, each data in the directed graph adjacency table data is represented as (source node ID, tail node ID, edge label), the source node ID represents the node ID of the source node, and the tail node ID represents the node ID of the tail node. It should be noted that Table 5 is only a schematic and should not be understood as a limitation of the present application. Based on this, in combination with the edge ID corresponding to each edge, a node and edge relationship table is further generated. Thus, a node and edge relationship table as shown in Table 6 below is obtained.

[0168] Table 6

[0169] (0,3,0,Illegal transfer) (1,3,3, legal transfer) (2,1,5,Legal transfer) (2,3,4, legal transfer) (3,4,1,Legal transfer) (4,0,2,Illegal transfer) (4,1.6,Illegal transfer)

[0170] Among them, each data in the node and edge relationship table is represented as (source node ID, tail node ID, edge ID, edge label), the source node ID represents the node ID of the source node, and the tail node ID represents the node ID of the tail node. It should be noted that Table 6 is only a schematic and should not be understood as a limitation of this application. Based on this, the node numbers of the same source node are aggregated, thereby obtaining a node aggregation table as shown in Table 7 below.

[0171] Table 7

[0172] (0,Iterator(3,0,illegal transfer)) (1,Iterator(3,3,Legal Transfer)) (2,Iterator(1,5,Legal Transfer)(3,4,Legal Transfer)) (3,Iterator(4,1,Legal Transfer)) (4,Iterator(0,2,illegal transfer)(1,6,illegal transfer))

[0173] Among them, each data in the node aggregation table is represented as (source node ID, Iterator (tail node ID, edge ID, edge label)), the source node ID represents the node ID of the source node, and the tail node ID represents the node ID of the tail node. It should be noted that Table 7 is only a schematic and should not be understood as a limitation on the present application.

[0174] Secondly, in the embodiment of the present application, a method for generating a node aggregation table with edge labels is provided. The above method only supports the case where the edges of the input graph are homogeneous and unlabeled, and is difficult to apply to heterogeneous graphs or input graphs with edge labels. However, in scenarios such as financial payment transaction detection, graphs are often heterogeneous or edges are labeled. Therefore, adding edge label processing can expand the scope of application and adapt to more scenarios. In addition, the node aggregation table can facilitate subsequent statistics and fusion, thereby improving the feasibility and operability of the solution.

[0175] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, for each executor, the number of neighborhood edges of each tail node is counted according to the node aggregation table, so that each executor respectively counts the first round of update information of each tail node, which may specifically include:

[0176] For each executor, according to the node aggregation table, determine the counter corresponding to each neighbor edge in the neighbor edge set, where each neighbor edge corresponds to a counter, and the counter stores the edge number and the cardinality estimate, and the cardinality estimate is 1;

[0177] For each executor, the counter corresponding to each neighborhood edge is fused with the counter of the corresponding tail node to obtain the first round of update information of each tail node.

[0178] In one or more embodiments, a method for determining the update information of the first round in the first round is introduced. It can be seen from the above embodiments that, without involving multiple edge labels, each executor can search for the tail node corresponding to the source node according to the aggregation result of each source node in the node aggregation table, thereby facilitating the counting of the number of neighboring edges of each tail node to obtain the update information of the first round of each tail node.

[0179] Specifically, before entering the statistics of the first round (i.e., round 0), the Spark driver applies to the parameter server to create a counter matrix of HyperLoglog, which is used to store two sets of counters corresponding to each node, namely, read counters (read-counter) and write counters (write-counter). Based on this, each data in the counter matrix of HyperLoglog can be expressed as (node ​​ID, read-counter, write-counter). Both sets of counters are used to store the estimated number of neighborhood edges of the node neighborhood edge set. The difference is that the read counter remains unchanged before the end of each round, while the write counter may be updated multiple times in each round. This is because multiple batches may need to be updated in the same round, and setting two counters can prevent the update changes from causing calculation errors.

[0180] According to the definition of the r-order neighborhood edge set, the 0-order edge set of any node is empty, that is, the estimated value of its neighborhood edge number is 0. Therefore, the initial read counter of each node is The stored estimates of the number of neighborhood edges are set to 0, and the initial write counters for each node are The stored estimates of the number of neighborhood edges are also set to 0.

[0181] When entering the first round (i.e., round 1), each executor processes its stored data in batches. In a processed batch, each executor generates a corresponding counter for each neighbor edge. Without distinguishing the edge label, the cardinality estimate in the corresponding counter of each neighbor edge is set to 1. Each executor performs statistics in the "incoming edge" direction, and merges the counter of each neighbor edge with the counter of its corresponding tail node (w) in turn, and obtains the fused counter as the update information of the first round of the tail node (w), i.e., m w (1). Similarly, other tail nodes also use similar methods to obtain their corresponding first round update information.

[0182] It should be noted that the present application performs statistics on neighboring edges in the "incoming edge" direction. In practical applications, the neighboring edges can also be counted in the "outgoing edge" direction. This is only an illustration and should not be understood as a limitation on the present application.

[0183] Secondly, in an embodiment of the present application, a method for determining the update information of the first round in the first round is provided. Through the above method, each executor can respectively count the update information of the first round, realize parallel processing and real-time processing, thereby improving data processing efficiency.

[0184] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, for each executor, the number of neighborhood edges of each tail node is counted according to the node aggregation table, so that each executor respectively counts the first round of update information of each tail node, which may specifically include:

[0185] For each executor, according to the node aggregation table, determine the counter corresponding to each neighbor edge in the neighbor edge set, where each neighbor edge corresponds to a counter, and the counter stores the edge label, edge number, and cardinality estimate. When the edge label is the target edge label, the cardinality estimate is 1, and when the edge label is the non-target edge label, the cardinality estimate is 0;

[0186] For each executor, the counter corresponding to each neighborhood edge is fused with the counter of the corresponding tail node to obtain the first round of update information of each tail node.

[0187] In one or more embodiments, another method of determining the update information of the first round in the first round is introduced. In the case of multiple edge labels, each executor can respectively search for the tail node and the edge label corresponding to the source node according to the aggregation result of each source node in the node aggregation table, thereby facilitating the counting of the number of neighboring edges of each tail node to obtain the update information of the first round of each tail node.

[0188] Before entering the first round (i.e., round 0) of statistics, the Spark driver requests the parameter server to create a HyperLoglog counter matrix to store two sets of counters corresponding to each node, namely the read counter and the write counter. And the initial read counter of each node The stored estimates of the number of neighborhood edges are set to 0, and the initial write counters for each node are The stored estimated value of the number of neighborhood edges is also set to 0. The specific process has been introduced in the above embodiment, so it will not be repeated here.

[0189] When entering the first round (i.e., round 1), each executor processes its stored data in batches. In a processed batch, each executor generates a corresponding counter for each neighbor edge. In the case of distinguishing edge labels, if the edge label corresponding to the neighbor edge (e) is the target edge label, i.e., l e ∈L, the cardinality estimator in the counter corresponding to the neighbor edge is set to 1. If the edge label corresponding to the neighbor edge (e) is not the target edge label, that is, Then the cardinality estimate in the counter corresponding to the neighborhood edge is set to 0. Based on this, if each executor performs statistics in the "incoming edge" direction, the counter of each neighborhood edge is merged with the counter of its corresponding tail node (w) in turn, and the fused counter is used as the first round of update information of the tail node (w), that is, m w (1). Similarly, other tail nodes also use similar methods to obtain their corresponding first round update information.

[0190] It should be noted that this application uses the "incoming edge" direction to count the neighboring edges. In actual applications, the neighboring edges can also be counted in the "outgoing edge" direction. This is only an illustration and should not be understood as a limitation on this application. In addition, in actual situations, multiple target edge labels can also be set (i.e., a target edge label set is set). This application takes setting a target edge label as an example for introduction, however, this should not be understood as a limitation on this application.

[0191] Secondly, in an embodiment of the present application, another method for determining the update information of the first round in the first round is provided. Through the above method, on the one hand, each executor can respectively count the update information of the first round, realize parallel processing and real-time processing, thereby improving data processing efficiency. On the other hand, it only supports the situation where the edges of the input graph are homogeneous and unlabeled, which is difficult to apply to heterogeneous graphs or input graphs with labeled edges. However, in scenarios such as financial payment transaction detection, graphs are often heterogeneous or edges are labeled. Therefore, adding processing for edge labels can expand the scope of application and adapt to more scenarios. Furthermore, it is possible to support scenarios with edge labels, so that the estimated value of the number of neighborhood edges with specific labels in the node neighborhood edge set can be estimated.

[0192] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, each executor sends the first round of update information of each tail node to the parameter server, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round of update information of each tail node, which may specifically include:

[0193] After the parameter server creates the counter matrix, each executor sends the first round of update information of the tail node to the parameter server, so that the parameter server fuses the zero-order write counter of the corresponding node with the update information of the first round according to the node number of each tail node and the update information of the first round, until the first round of iteration is completed, the parameter server determines the corresponding first-order read counter according to the first-order write counter of each node in the K nodes stored in the counter matrix, where the initial cardinality estimate of the zero-order write counter is 0.

[0194] In one or more embodiments, a method for implementing distributed fusion information based on a parameter server is introduced. As can be seen from the above embodiments, the parameter server has multiple partitions, each partition fuses counters of a portion of nodes, and usually, different partitions will not repeatedly fuse the same nodes.

[0195] Specifically, each executor pushes the first round of update information of each tail node of a batch to the corresponding partition of the parameter server. For example, partition A is responsible for fusing the counter of the tail node 1, then the executor needs to push the first round of update information of the tail node 1 to partition A. Based on this, the parameter server pushes the first round of update information (m w (1)) and the zero-order write counter The zero-order write counter is updated by fusion, thereby obtaining the first-order write counter, that is,

[0196] It can be understood that for the first round of update information of the first batch received by the PS, the parameter server is the first round of update information (m w (1)) and the zero-order write counter Fusion is performed to complete the zero-order write counter The update of the first-order write counter Right now It can be seen that the zero-order write counter The estimated value of the number of neighboring edges is 0. The first-order write counter is continuously updated starting from the next batch.

[0197] After all batches are processed, one round of iteration is completed. Therefore, in the counter matrix on the parameter server, for each node (v), let At this time, the first-order read counter and the first-order write counter corresponding to each node in the counter matrix both store the estimated value of the first-order neighborhood edge number of the node.

[0198] Secondly, in the embodiment of the present application, a method for implementing distributed fusion information based on a parameter server is provided. Through the above method, the parameter server supports operations such as obtaining and updating model parameters in parallel, and has the characteristics of high performance, scalability and fault tolerance. By introducing PS, the single-point network bottleneck of the Spark driver can be eliminated, and a large amount of data shuffling operations and huge network overhead generated when calculating update information and updating counters can be avoided, thereby improving performance. The actual application scenarios of this application are wider, and the performance on ultra-large-scale graph data is superior.

[0199] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, another optional embodiment provided by the embodiment of the present application sends the first round of update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round of update information of each tail node, and can also include:

[0200] For each executor, get the first-order read counter corresponding to the source node from the parameter server;

[0201] For each executor, determine a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node;

[0202] For each executor, according to the first-order read counters of all source nodes corresponding to each tail node in the tail node set, the second round update information corresponding to each tail node is determined by fusing the first-order read counters of all corresponding source nodes;

[0203] The second-round update information of each tail node is sent to the parameter server through each executor, so that the parameter server fuses the first-order write counter of the corresponding node with the second-round update information according to the number of each tail node and the second-round update information until the second round of iteration is completed. The parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, where the second-order read counter stores the estimated value of the second-order neighborhood edge number.

[0204] In one or more embodiments, a method for determining an estimated value of the number of second-order neighbor edges is introduced. As can be seen from the above embodiments, the first-order read counter and the first-order write counter also store edge number information, thereby achieving deduplication and fusion processing of edges. The following will introduce a method for determining the estimated value of the number of second-order neighbor edges corresponding to each node in the second round.

[0205] Specifically, in the second round, each executor processes its stored data in batches. In each processing batch, each executor pulls the corresponding first-order read counter from the parameter server according to the source node number. Then, each source node pulls the first-order read counter Send to each tail node in the corresponding tail node set. At this time, each tail node (w) in the tail node set receives the first-order read counter of all source nodes To fuse (i.e. Where N1(w) represents the first-order neighborhood vertex set of the tail node w), and the second round update information of each tail node is obtained (m w (2)).

[0206] Each executor sends the second round update information corresponding to each tail node in a batch to the parameter server. Since the parameter server has multiple partitions, the executor needs to push the second round update information to the specified partition. For the parameter server, the second round update information reported by the executor is processed in parallel and in real time.

[0207] Based on the second round update information of each tail node reported by each executor, the parameter server merges the first-order write counter of the corresponding numbered node with the second round update information to obtain the second-order write counter of the node. After a round ends, the parameter server obtains the second-order write counters of all nodes. Therefore, the second-order read counter corresponding to the node Updated to the corresponding second-order write counter At this point, the second-order read counter corresponding to each node and a second-order write counter The estimated number of second-order neighbor edges of the node is stored.

[0208] Secondly, in an embodiment of the present application, a method for determining the estimated value of the second-order neighborhood edge number is provided. Through the above method, it can be seen that the method for determining the estimated value of the second-order neighborhood edge number is different from the method for determining the estimated value of the first-order neighborhood edge number. There is no need to count the edges again, but directly merge the counters corresponding to the nodes, thereby improving processing efficiency.

[0209] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, for each executor, according to the first-order read counter of the source node corresponding to each tail node in the tail node set, determining the second round update information corresponding to each tail node, may specifically include:

[0210] For each executor, obtain the first-order read counter of the source node corresponding to each tail node in the tail node set;

[0211] For each executor, the first-order read counters of the source nodes corresponding to each tail node are fused to obtain the second round update information corresponding to each tail node.

[0212] In one or more embodiments, a method for determining the update information of the second round in the second round is introduced. Starting from the second round, the first-order read counters of all source nodes corresponding to each tail node are merged to obtain the update information of the second round corresponding to the tail node.

[0213] Specifically, when entering the second round, each executor processes its stored data in batches. In a processed batch, the same source node has a corresponding tail node set. After obtaining multiple tail node sets, for each tail node, the first-order read counters of all its source nodes can be fused. Exemplarily, assume that node 2 is the tail node, and its corresponding source node is node 7. Node 3 is the tail node, and its corresponding source node is also node 7. Based on this, the source nodes can be fused. Taking the fusion of node 7 as the source node as an example, that is, the first-order read counters of node 7 are fused, thereby obtaining the update information of the second round corresponding to node 7, that is, m w (2). Similarly, other source nodes also use similar methods to obtain their corresponding second-round update information.

[0214] Again, in an embodiment of the present application, a method for determining the update information of the second round in the second round is provided. Through the above method, each executor can respectively count the update information of the second round, realize parallel processing and real-time processing, thereby improving data processing efficiency.

[0215] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, each executor sends the second round update information of each tail node to the parameter server, so that the parameter server performs a fusion process on the first-order write counter of the corresponding node and the second round update information according to the node number of each tail node and the second round update information until the second round of iteration is completed. The parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, which may specifically include:

[0216] Each executor sends the second-round update information of each tail node to the parameter server, so that the parameter server fuses the first-order write counter of the same tail node with the second-round update information according to the second-round update information of each tail node and the node number of each tail node to obtain the second-order write counter until the second round of iterations is completed. The second-order write counter of each node in the K nodes of the parameter server determines the corresponding second-order read counter.

[0217] In one or more embodiments, a method for implementing distributed fusion information based on a parameter server is introduced. As can be seen from the above embodiments, the parameter server has multiple partitions, each partition updates the counters of a portion of nodes respectively, and usually, different partitions will not repeat the update of the same node.

[0218] Specifically, each executor pushes the second round update information of each tail node of a batch to the corresponding partition of the parameter server. For example, partition B is responsible for updating the counter of the tail node 1, then the executor needs to push the second round update information of the tail node 1 to partition B. Based on this, the parameter server pushes the second round update information (m w (2)) and the first-order write counter The fusion is performed to complete the update of the first-order write counter, that is,

[0219] It can be understood that for the second round of update information of the first batch received by the PS, the parameter server is the second round of update information (m w (2)) and the first-order write counter The first-order write counter is updated by fusion to obtain the second-order write counter, that is, The second-order write counter is continuously updated starting from the next batch.

[0220] After all batches are processed, one round of iteration is completed. Therefore, in the counter matrix on the parameter server, for each node (v), let At this time, the second-order read counter and the second-order write counter corresponding to each node in the counter matrix both store the estimated value of the second-order neighborhood edge number of the node.

[0221] Again, in the embodiment of the present application, a method for implementing distributed fusion information based on a parameter server is provided. Through the above method, the parameter server supports operations such as obtaining and updating model parameters in parallel, and has the characteristics of high performance, scalability and fault tolerance. By introducing PS, the single-point network bottleneck of the Spark driver can be eliminated, and a large number of data shuffling operations and huge network overhead generated when calculating update information and updating counters can be avoided, thereby improving performance. The actual application scenarios of this application are wider, and the performance on ultra-large-scale graph data is superior.

[0222] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, each executor sends the second round update information of each tail node to the parameter server, so that the parameter server performs a fusion process on the first-order write counter of the corresponding node and the second round update information according to the number of each tail node and the second round update information until the second round of iteration is completed. After the parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, it may also include:

[0223] For each executor, obtain the second-order read counter corresponding to the source node from the parameter server;

[0224] For each executor, determine a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node;

[0225] For each executor, according to the second-order read counters of all source nodes corresponding to each tail node in the tail node set, the third round update information corresponding to each tail node is determined by fusing the second-order read counters of all corresponding source nodes;

[0226] The third-round update information of each tail node is sent to the parameter server through each executor, so that the parameter server fuses the second-order write counter of the corresponding node with the third-round update information according to the number of each tail node and the third-round update information until the third round of iteration is completed. The parameter server determines the corresponding third-order read counter according to the third-order write counter of each node in the K nodes, where the third-order read counter stores the estimated value of the third-order neighborhood edge number.

[0227] In one or more embodiments, a method for determining an estimated value of the number of third-order neighbor edges is introduced. The second-order read counter and the second-order write counter also include edge IDs, thereby enabling edge fusion processing. The following describes a method for determining an estimated value of the number of third-order neighbor edges corresponding to each node in the third round.

[0228] Specifically, in the third round, each executor processes its stored data in batches. In each processing batch, each executor pulls the corresponding second-order read counter from the parameter server according to the node ID of the source node. Then, each source node sends the pulled second-order read counter to each tail node in the corresponding tail node set. At this time, each tail node (w) in the tail node set fuses the second-order read counters received from all source nodes (i.e., Where N2(w) represents the second-order neighborhood vertex set), and the third round of update information of each tail node is obtained (m w (3)).

[0229] Each executor sends the third round update information corresponding to each tail node in a batch to the parameter server. Since the parameter server has multiple partitions, the executor needs to push the third round update information to the specified partition. For the parameter server, the third round update information reported by the executor is processed in parallel and in real time.

[0230] Based on the third-round update information reported by each executor, the parameter server fuses the third-round update information of the node with the same number with the second-order write counter. That is, the parameter server obtains the third-order write counter of each node by fusing the second-order write counter of each node in the K nodes with the corresponding third-round update information. After a round ends, the parameter server obtains the updated third-order write counter, and then updates the third-order read counter of the node to the third-order write counter. At this point, the estimated value of the number of third-order neighborhood edges of each node can be stored based on the third-order read counter and the third-order write counter of each node in the K nodes.

[0231] It should be noted that, when the number is greater than or equal to the fourth round, a similar method is used to determine the estimated value of the number of neighborhood edges corresponding to the corresponding order, which will not be elaborated here.

[0232] Again, in an embodiment of the present application, a method for determining the estimated value of the third-order neighborhood edge number is provided. Through the above method, since the method for determining the estimated value of the third-order neighborhood edge number is similar to the method for determining the estimated value of the second-order neighborhood edge number, only one more round of calculation is required to obtain the desired result, thereby improving the feasibility and operability of the solution.

[0233] Optionally, in the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiment of the present application, obtaining the directed graph adjacency table data may specifically include:

[0234] Receiving a data upload instruction for a data upload control sent by a terminal device, wherein the data upload control is displayed on an interactive interface provided by the terminal device;

[0235] Obtain directed graph adjacency table data according to the data upload instruction;

[0236] It may also include:

[0237] Receiving a label selection instruction for at least one edge label sent by a terminal device, wherein at least one selectable edge label is displayed on an interactive interface provided by the terminal device;

[0238] Determine the target edge label according to the label selection instruction;

[0239] It may also include:

[0240] Receiving an order setting instruction for an order adjustment control sent by a terminal device, wherein the order adjustment control is displayed on an interactive interface provided by the terminal device;

[0241] Determine the target order according to the order setting instruction;

[0242] It may also include:

[0243] Receiving a node viewing instruction for any node sent by a terminal device, wherein at least one node is displayed on an interactive interface provided by the terminal device;

[0244] If the node viewing instruction carries a first-order identifier, the estimated value of the number of first-order neighboring edges corresponding to any node is sent to the terminal device, so that the terminal device displays the estimated value of the number of first-order neighboring edges corresponding to any node;

[0245] If the node viewing instruction carries a second-order identifier, the estimated value of the second-order neighboring edges corresponding to any node is sent to the terminal device, so that the terminal device displays the estimated value of the second-order neighboring edges corresponding to any node.

[0246] In one or more embodiments, a method for implementing related operations based on an interactive interface is introduced. The user can also upload directed graph adjacency table data on his own, or set the corresponding target order, or select the required target edge label, or choose to view the estimated value of the number of r-order neighboring edges corresponding to a node. The functions provided by the interactive interface will be introduced below in conjunction with the diagrams.

[0247] 1. Autonomous directed graph adjacency table data;

[0248] Specifically, for ease of understanding, see Figure 8 , Figure 8 This is a schematic diagram of an interactive interface in an embodiment of the present application. As shown in the figure, A1 is used to indicate a data upload control, wherein: Figure 8 Six data upload controls are shown. Assuming that the user clicks the data upload control corresponding to "transfer data B", the data upload instruction for the data upload control is triggered, and then the neighborhood edge number estimation value determination device obtains the directed graph adjacency table data corresponding to "transfer data B".

[0249] 2. Customize the target edge label;

[0250] Specifically, for ease of understanding, see Fig. 9 , Fig. 9 This is another schematic diagram of the interactive interface in the embodiment of the present application, where B1 is used to indicate an optional edge label, wherein: Fig. 9 Five optional edge labels are shown. Assuming that the user clicks the edge label corresponding to "illegal transfer", the label selection instruction for the edge label is triggered, and then the neighborhood edge number estimation value determination device takes "illegal label" as a target edge label.

[0251] 3. Customize the target order;

[0252] Specifically, for ease of understanding, see Fig.10 , Fig.10 This is another schematic diagram of the interactive interface in the embodiment of the present application, C1 is used to indicate the order adjustment control, C2 is used to indicate the confirmation control, wherein, Fig.10 The order adjustment control shown has the functions of increasing and decreasing. Assuming the default order is 1, when the user clicks the increase button in the order adjustment control, the order increases. When the user clicks the decrease button in the order adjustment control, the order decreases. After selecting the order, click "Confirm Control" to confirm the target order set by the user.

[0253] It should be noted that the target order is an integer greater than or equal to 1.

[0254] 4. Check the estimated value of the number of r-order neighboring edges corresponding to a certain node;

[0255] Specifically, for ease of understanding, see Fig.11 , Fig.11 It is another schematic diagram of the interactive interface in the embodiment of the present application, and D1 is used to indicate the confirmation control. When the user clicks on any node in the interface (for example, node C), the node view instruction for the node (i.e., node C) is triggered. In the node setting stage, the user has set the target order. Based on this, assuming that the target order is "first order", the node view instruction carries a first-order identifier, and what the user needs to view is the first-order neighborhood edge estimate. Similarly, assuming that the target order is "second order", the node view instruction carries a second-order identifier, and what the user needs to view is the second-order neighborhood edge estimate. And so on, it is not exhaustive here.

[0256] See also Fig.12 , Fig.12 This is another schematic diagram of the interactive interface in the embodiment of the present application. As shown in the figure, after the user selects the corresponding node, the estimated value of the number of neighboring edges corresponding to the node can be displayed on the interface. Taking the estimated value of the second-order neighboring edges of node C as an example, the interface displays "The estimated value of the second-order neighboring edges of node C (user 775895) is 125".

[0257] Secondly, in an embodiment of the present application, a method for implementing relevant operations based on an interactive interface is provided. Through the above method, a user can trigger relevant instructions through a terminal device, and display corresponding content in combination with the instructions triggered by the user. Thereby, on the one hand, the flexibility and operability of the scheme are increased, and on the other hand, the estimated value of the number of neighborhood edges can be accurately and intuitively displayed to the user, making it easier for the user to understand the details, thereby improving the practicality of the scheme.

[0258] The following is a detailed description of the neighborhood edge number estimation value determination device in this application, please refer to Fig.13 , Fig.13 This is a schematic diagram of an embodiment of a device for determining an estimated value of the number of neighboring edges in an embodiment of the present application. The device 20 for determining an estimated value of the number of neighboring edges includes:

[0259] An acquisition module 210 is used to acquire directed graph adjacency table data, wherein the directed graph adjacency table data represents adjacency table data corresponding to a directed graph, the directed graph includes K nodes, each piece of data in the directed graph adjacency table data includes a node number of a source node and a node number of a tail node corresponding to a directed edge, the directed graph adjacency table data is stored on at least two executors, and K is an integer greater than 1;

[0260] A generating module 220, configured to generate a node aggregation table according to the directed graph adjacency table data, wherein the node aggregation table includes a node number of a source node, a node number set of a tail node, and an edge number set, wherein the node number set includes at least one node number, and the edge number set includes at least one edge number;

[0261] A statistics module 230 is used to count the number of neighboring edges of each tail node according to the node aggregation table for each executor, so that each executor can respectively obtain the first round of update information of each tail node;

[0262] The sending module 240 is used to send the first round update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round update information of each tail node, wherein the first-order read counter stores the estimated value of the number of first-order neighborhood edges.

[0263] In an embodiment of the present application, a device for determining an estimated value of the number of neighborhood edges is provided. By adopting the above device and introducing a parameter server, not only can distributed storage of large-scale data be realized, but also operations such as parallel acquisition of data and update of data can be supported. Therefore, on the one hand, the problem of driver single-point network bottleneck can be eliminated and computing performance can be improved. On the other hand, there is no need to additionally create and store RDD for updating node count information, thereby saving computing resources.

[0264] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application,

[0265] The acquisition module 210 is specifically used to acquire undirected graph adjacency table data, wherein the undirected graph adjacency table data represents the adjacency table data corresponding to the undirected graph, the undirected graph includes K nodes, and each forward data in the undirected graph adjacency table data includes a node number of a source node and a node number of a tail node;

[0266] For each forward data in the undirected graph adjacency table data, generate reverse data corresponding to each forward data, wherein the source node in the reverse data is the tail node in the corresponding forward data, and the tail node in the reverse data is the source node in the corresponding forward data;

[0267] Directed graph adjacency table data is generated according to each piece of forward data and the reverse data corresponding to each piece of forward data, wherein both the forward data and the reverse data belong to the data in the directed graph adjacency table data.

[0268] In the embodiment of the present application, a device for determining the estimated value of the number of edges in a neighborhood is provided. The above device is used, considering that in practical applications, many data are expressed as undirected graphs. Therefore, in combination with the characteristics of the graph being undirected and directed, a conversion function between undirected graphs and directed graphs is designed. Thus, the application scenarios in which the input is compatible with both undirected graphs and directed graphs are achieved.

[0269] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application,

[0270] The generating module 220 is specifically used to encode the edge corresponding to each data according to the directed graph adjacency table data to obtain the edge number corresponding to each edge;

[0271] Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, and the edge number;

[0272] According to the node and edge relationship table, the node numbers of the same source node are aggregated to obtain a node aggregation table.

[0273] In an embodiment of the present application, a device for determining the estimated value of the number of neighborhood edges is provided. The device is used to aggregate the node numbers of the same source node using the source node as the clustering basis, thereby obtaining a node aggregation table. The node aggregation table can facilitate subsequent statistics and fusion, thereby improving the feasibility and operability of the solution.

[0274] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the neighborhood edge number estimation value determining device 20 provided by the embodiment of the present application, each piece of data in the directed graph adjacency table data further includes an edge label corresponding to the edge, and the node aggregation table further includes an edge label set, and the edge label set includes at least one edge label;

[0275] The generating module 220 is specifically used to encode the edge corresponding to each data according to the directed graph adjacency table data to obtain the edge number corresponding to each edge;

[0276] Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, the edge number and the edge label;

[0277] According to the node and edge relationship table, the node numbers of the same source nodes are aggregated to obtain a node aggregation table.

[0278] In the embodiment of the present application, a device for determining the estimated value of the number of neighborhood edges is provided. The above device only supports the case where the edges of the input graph are homogeneous and unlabeled, and is difficult to apply to heterogeneous graphs or input graphs with labeled edges. However, in scenarios such as financial payment transaction detection, graphs are often heterogeneous or edges are labeled. Therefore, adding edge label processing can expand the scope of application and adapt to more scenarios. In addition, the node aggregation table can facilitate subsequent statistics and fusion, thereby improving the feasibility and operability of the solution.

[0279] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application,

[0280] The statistics module 230 is specifically used to determine, for each executor, a counter corresponding to each neighbor edge in the neighbor edge set according to the node aggregation table, wherein each neighbor edge corresponds to a counter, and the counter stores an edge number and a cardinality estimate, and the cardinality estimate is 1;

[0281] For each executor, the counter corresponding to each neighborhood edge is fused with the counter of the corresponding tail node to obtain the first round of update information of each tail node.

[0282] In an embodiment of the present application, a device for determining an estimated value of the number of neighborhood edges is provided. By using the above device, each executor can respectively perform statistics on the update information of the first round, realize parallel processing and real-time processing, thereby improving data processing efficiency.

[0283] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application,

[0284] The statistics module 230 is specifically used to determine, for each executor, a counter corresponding to each neighbor edge in the neighbor edge set according to the node aggregation table, wherein each neighbor edge corresponds to a counter, and the counter stores the edge label, the edge number, and the cardinality estimate. When the edge label is a target edge label, the cardinality estimate is 1, and when the edge label is a non-target edge label, the cardinality estimate is 0;

[0285] For each executor, the counter corresponding to each neighborhood edge is fused with the counter of the corresponding tail node to obtain the first round of update information of each tail node.

[0286] In an embodiment of the present application, a device for determining the estimated value of the number of neighborhood edges is provided. By using the above device, on the one hand, each executor can respectively count the update information of the first round, realize parallel processing and real-time processing, thereby improving data processing efficiency. On the other hand, it only supports the situation where the edges of the input graph are homogeneous and unlabeled, which is difficult to apply to heterogeneous graphs or input graphs with labeled edges. However, in scenarios such as financial payment transaction detection, graphs are often heterogeneous or edges are labeled. Therefore, adding processing for edge labels can expand the scope of application and adapt to more scenarios. Furthermore, it is possible to support scenarios with edge labels, so that the estimated value of the number of neighborhood edges with specific labels in the node neighborhood edge set can be estimated.

[0287] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application,

[0288] The sending module 240 is specifically used to send the first round of update information of the tail node to the parameter server through each executor after the parameter server creates the counter matrix, so that the parameter server fuses the zero-order write counter of the corresponding node with the update information of the first round according to the number of each tail node and the update information of the first round, until the first round of iteration is completed, and the parameter server determines the corresponding first-order read counter according to the first-order write counter of each node in the K nodes stored in the counter matrix, wherein the initial cardinality estimate of the zero-order write counter is 0.

[0289] In an embodiment of the present application, a device for determining the estimated value of the number of neighborhood edges is provided. By using the above device, the parameter server supports operations such as obtaining and updating model parameters in parallel, and has the characteristics of high performance, scalability and fault tolerance. By introducing PS, the single-point network bottleneck of the Spark driver can be eliminated, and a large number of data shuffling operations and huge network overhead generated when calculating update information and updating counters can be avoided, thereby improving performance. The actual application scenarios of this application are wider, and the performance on ultra-large-scale graph data is superior.

[0290] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application, the apparatus 20 for determining the estimated value of the number of neighborhood edges further includes a determination module 250;

[0291] The acquisition module 210 is further configured to send the first-round update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first-round update information of each tail node, and then acquire the first-order read counter corresponding to the source node from the parameter server for each executor;

[0292] The determination module 250 is further configured to determine, for each executor, a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node;

[0293] The determination module 250 is further configured to determine, for each executor, the second round update information corresponding to each tail node by fusing the first-order read counters of all source nodes corresponding to each tail node in the tail node set;

[0294] The sending module 240 is also used to send the second-round update information of each tail node to the parameter server through each executor, so that the parameter server can fuse the first-order write counter of the corresponding node with the second-round update information according to the node number of each tail node and the second-round update information until the second round of iteration is completed. The parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, wherein the second-order read counter stores the estimated value of the second-order neighborhood edge number.

[0295] In an embodiment of the present application, a device for determining a neighborhood edge count estimate is provided. By using the above device, it can be seen that the method for determining a second-order neighborhood edge count estimate is different from the method for determining a first-order neighborhood edge count estimate. There is no need to count the edges again, but the counters corresponding to the nodes are directly merged, thereby improving processing efficiency.

[0296] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application,

[0297] The determination module 250 is specifically configured to obtain, for each executor, a first-order read counter of a source node corresponding to each tail node in the tail node set;

[0298] For each executor, the first-order read counters of the source nodes corresponding to each tail node are fused to obtain the second round update information corresponding to each tail node.

[0299] In an embodiment of the present application, a device for determining an estimated value of the number of neighborhood edges is provided. By using the above device, each executor can respectively perform statistics on the update information of the second round, realize parallel processing and real-time processing, thereby improving data processing efficiency.

[0300] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application,

[0301] The sending module 240 is specifically used to send the second-round update information of each tail node to the parameter server through each executor, so that the parameter server fuses the first-order write counter of the same tail node with the second-round update information according to the second-round update information of each tail node and the node number of each tail node to obtain a second-order write counter, until the second round of iterations is completed, and the second-order write counter of each node in the K nodes of the parameter server determines the corresponding second-order read counter.

[0302] In an embodiment of the present application, a device for determining the estimated value of the number of neighborhood edges is provided. By using the above device, the parameter server supports operations such as obtaining and updating model parameters in parallel, and has the characteristics of high performance, scalability and fault tolerance. By introducing PS, the single-point network bottleneck of the Spark driver can be eliminated, and a large number of data shuffling operations and huge network overhead generated when calculating update information and updating counters can be avoided, thereby improving performance. The actual application scenarios of this application are wider, and the performance on ultra-large-scale graph data is superior.

[0303] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the apparatus 20 for determining the estimated value of the number of neighborhood edges provided in the embodiment of the present application,

[0304] The acquisition module 210 is further used to send the second-round update information of each tail node to the parameter server through each executor, so that the parameter server performs a fusion process on the first-order write counter of the corresponding node and the second-round update information according to the number of each tail node and the second-round update information, until the second round of iteration is completed, and the parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, and then obtains the second-order read counter corresponding to the source node from the parameter server for each executor;

[0305] The determination module 250 is further configured to determine, for each executor, a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node;

[0306] The determination module 250 is further configured to determine, for each executor, the third round of update information corresponding to each tail node by fusing the second-order read counters of all source nodes corresponding to each tail node in the tail node set;

[0307] The sending module 240 is also used to send the third-round update information of each tail node to the parameter server through each executor, so that the parameter server fuses the second-order write counter of the corresponding node with the third-round update information according to the number of each tail node and the third-round update information until the third round of iteration is completed. The parameter server determines the corresponding third-order read counter according to the third-order write counter of each node in the K nodes, wherein the third-order read counter stores the estimated value of the third-order neighborhood edge number.

[0308] In an embodiment of the present application, a device for determining a neighborhood edge count estimate is provided. Using the above device, since the method for determining a third-order neighborhood edge count estimate is similar to the method for determining a second-order neighborhood edge count estimate, only one more round of calculation is needed to obtain the desired result, thereby improving the feasibility and operability of the solution.

[0309] Optionally, in the above Fig.13 On the basis of the corresponding embodiment, in another embodiment of the neighborhood edge number estimation value determining device 20 provided in the embodiment of the present application, the neighborhood edge number estimation value determining device 20 further includes a receiving module 260 and a sending module 270;

[0310] The acquisition module 210 is specifically used to receive a data upload instruction for a data upload control sent by a terminal device, wherein the data upload control is displayed on an interactive interface provided by the terminal device;

[0311] Obtain directed graph adjacency table data according to the data upload instruction;

[0312] A receiving module 260, configured to receive a tag selection instruction for at least one edge tag sent by a terminal device, wherein at least one selectable edge tag is displayed on an interactive interface provided by the terminal device;

[0313] A determination module 250, configured to determine a target edge label according to a label selection instruction;

[0314] The receiving module 260 is further configured to receive an order setting instruction for an order adjustment control sent by a terminal device, wherein the order adjustment control is displayed on an interactive interface provided by the terminal device;

[0315] The determination module 250 is further used to determine the target order according to the order setting instruction;

[0316] The receiving module 260 is further configured to receive a node viewing instruction for any node sent by a terminal device, wherein at least one node is displayed on an interactive interface provided by the terminal device;

[0317] A sending module 270, configured to send the estimated value of the number of first-order neighbor edges corresponding to any node to the terminal device if the node viewing instruction carries a first-order identifier, so that the terminal device displays the estimated value of the number of first-order neighbor edges corresponding to any node;

[0318] The sending module 270 is also used to send the estimated value of the second-order neighboring edges corresponding to any node to the terminal device if the node viewing instruction carries a second-order identifier, so that the terminal device displays the estimated value of the second-order neighboring edges corresponding to any node.

[0319] In an embodiment of the present application, a device for determining an estimated value of the number of neighborhood edges is provided. By using the above device, a user can trigger relevant instructions through a terminal device, and display corresponding content in combination with the instructions triggered by the user. Thus, on the one hand, the flexibility and operability of the scheme are increased, and on the other hand, the estimated value of the number of neighborhood edges can be intuitively displayed to the user, making it easier for the user to understand the details, thereby improving the practicality of the scheme.

[0320] Fig.14 3 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device 300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 322 (for example, one or more processors) and memory 332, and one or more storage media 330 (for example, one or more mass storage devices) storing application programs 342 or data 344. Among them, the memory 332 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device. Furthermore, the central processing unit 322 can be configured to communicate with the storage medium 330 to execute a series of instruction operations in the storage medium 330 on the computer device 300.

[0321] The computer device 300 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input and output interfaces 358, and / or one or more operating systems 341, such as Windows Server 2008, ... TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.

[0322] The steps performed by the computer device in the above embodiment can be based on the Fig.14 The computer device structure shown.

[0323] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method described in each of the aforementioned embodiments.

[0324] The embodiments of the present application also provide a computer program product including a program, which, when executed on a computer, enables the computer to execute the method described in each of the aforementioned embodiments.

[0325] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0326] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0327] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0328] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0329] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0330] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for determining an estimated value of the number of neighborhood edges, characterized in that: include: Obtaining directed graph adjacency table data, wherein the directed graph adjacency table data represents adjacency table data corresponding to a directed graph, the directed graph includes K nodes, each piece of data in the directed graph adjacency table data includes a node number of a source node and a node number of a tail node corresponding to a directed edge, the directed graph adjacency table data is stored on at least two executors, and K is an integer greater than 1; Generate a node aggregation table according to the directed graph adjacency table data, wherein each data of the node aggregation table includes a node number of a source node, a node number set of a tail node, and an edge number set, the node number set includes at least one node number, and the edge number set includes at least one edge number; For each executor, the number of neighboring edges of each tail node is counted according to the node aggregation table, so that each executor can respectively obtain the first round of update information of each tail node; The first round of update information of each tail node is sent to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round of update information of each tail node, wherein the first-order read counter stores the estimated value of the first-order neighborhood edge number.

2. The determination method according to claim 1, characterized in that: The step of obtaining the adjacency list data of the directed graph includes: Obtaining undirected graph adjacency table data, wherein the undirected graph adjacency table data represents adjacency table data corresponding to the undirected graph, the undirected graph includes the K nodes, and each forward data in the undirected graph adjacency table data includes a node number of a source node and a node number of a tail node; For each forward data in the undirected graph adjacency table data, generate reverse data corresponding to each forward data, wherein the source node in the reverse data is the tail node in the corresponding forward data, and the tail node in the reverse data is the source node in the corresponding forward data; The directed graph adjacency table data is generated according to each piece of forward data and the reverse data corresponding to each piece of forward data, wherein both the forward data and the reverse data belong to the data in the directed graph adjacency table data.

3. The determination method according to claim 1, characterized in that: The generating a node aggregation table according to the directed graph adjacency table data comprises: According to the directed graph adjacency table data, encode the edge corresponding to each piece of data to obtain an edge number corresponding to each edge; Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, and the edge number; According to the node-edge relationship table, aggregation processing is performed according to the node number of the source node to obtain the node aggregation table.

4. The determination method according to claim 1, characterized in that: Each piece of data in the directed graph adjacency table data further includes an edge label corresponding to the edge, and the node aggregation table further includes an edge label set, and the edge label set includes at least one edge label; The generating a node aggregation table according to the directed graph adjacency table data comprises: According to the directed graph adjacency table data, encode the edge corresponding to each piece of data to obtain an edge number corresponding to each edge; Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, the edge number and the edge label; According to the node-edge relationship table, node numbers of the same source nodes are aggregated to obtain the node aggregation table.

5. The determination method according to claim 1, characterized in that: For each executor, the number of neighboring edges of each tail node is counted according to the node aggregation table, so that each executor respectively obtains the first round of update information of each tail node by counting, including: For each of the executors, determine, according to the node aggregation table, a counter corresponding to each neighbor edge in the neighbor edge set, wherein each of the neighbor edges corresponds to a counter, the counter stores an edge number and a cardinality estimate, and the cardinality estimate is 1; For each executor, the counter corresponding to each neighborhood edge and the counter of the corresponding tail node are fused to obtain the first round of update information of each tail node.

6. The determination method according to claim 1, characterized in that: For each executor, the number of neighboring edges of each tail node is counted according to the node aggregation table, so that each executor respectively obtains the first round of update information of each tail node by counting, including: For each of the executors, according to the node aggregation table, determine the counter corresponding to each neighbor edge in the neighbor edge set, wherein each of the neighbor edges corresponds to a counter, the counter stores an edge label, an edge number, and a cardinality estimate, and the cardinality estimate is 1 when the edge label is a target edge label, and the cardinality estimate is 0 when the edge label is a non-target edge label; For each executor, the counter corresponding to each neighborhood edge and the counter of the corresponding tail node are fused to obtain the first round of update information of each tail node.

7. The determination method according to claim 1, characterized in that: The sending the first round update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round update information of each tail node, including: After the parameter server creates the counter matrix, the update information of the first round of the tail node is sent to the parameter server through each executor, so that the parameter server fuses the zero-order write counter of the corresponding node with the update information of the first round according to the number of each tail node and the update information of the first round, until the first round of iteration is completed, and the parameter server determines the corresponding first-order read counter according to the first-order write counter of each node in the K nodes stored in the counter matrix, wherein the initial cardinality estimate of the zero-order write counter is 0.

8. The determination method according to claim 1, characterized in that: After sending the first round update information of each tail node to the parameter server through each executor so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round update information of each tail node, the method further includes: For each of the executors, obtaining a first-order read counter corresponding to a source node from the parameter server; For each of the executors, determining a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node; For each of the executors, according to the first-order read counters of all source nodes corresponding to each tail node in the tail node set, the second round update information corresponding to each tail node is determined by fusing the first-order read counters of all corresponding source nodes; The second-round update information of each tail node is sent to the parameter server through each executor, so that the parameter server fuses the first-order write counter of the corresponding node with the second-round update information according to the number of each tail node and the second-round update information until the second round of iteration is completed. The parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, wherein the second-order read counter stores the estimated value of the second-order neighborhood edge number.

9. The determination method according to claim 8, characterized in that: The step of determining, for each executor, the second round of update information corresponding to each tail node according to the first-order read counter of the source node corresponding to each tail node in the tail node set includes: For each of the executors, obtaining a first-order read counter of a source node corresponding to each tail node in the tail node set; For each executor, a fusion process is performed on the first-order read counters of the source nodes corresponding to each tail node to obtain the second round update information corresponding to each tail node.

10. The determination method according to claim 8, characterized in that: The second round update information of each tail node is sent to the parameter server by each executor, so that the parameter server performs fusion processing on the first-order write counter of the corresponding node and the second round update information according to the node number of each tail node and the second round update information, until the second round of iteration is completed, and the parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, including: The second-round update information of each tail node is sent to the parameter server by each executor, so that the parameter server fuses the first-order write counter of the same tail node with the second-round update information according to the second-round update information of each tail node and the node number of each tail node to obtain a second-order write counter until the second round of iteration is completed. The parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes.

11. The determination method according to claim 8, characterized in that: The second round update information of each tail node is sent to the parameter server by each executor, so that the parameter server performs fusion processing on the first-order write counter of the corresponding node and the second round update information according to the number of each tail node and the second round update information, until the second round of iteration is completed, and the parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, the method also includes: For each of the executors, obtaining a second-order read counter corresponding to a source node from the parameter server; For each of the executors, determining a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node; For each of the executors, according to the second-order read counters of all source nodes corresponding to each tail node in the tail node set, the third round of update information corresponding to each tail node is determined by fusing the second-order read counters of all corresponding source nodes; The third round update information of each tail node is sent to the parameter server through each executor, so that the parameter server fuses the second-order write counter of the corresponding node with the third round update information according to the number of each tail node and the third round update information until the third round of iteration is completed. The parameter server determines the corresponding third-order read counter according to the third-order write counter of each node in the K nodes, wherein the third-order read counter stores the estimated value of the third-order neighborhood edge number.

12. The determination method according to any one of claims 1 to 11, characterized in that: The step of obtaining the directed graph adjacency table data comprises: Receiving a data upload instruction for a data upload control sent by a terminal device, wherein the data upload control is displayed on an interactive interface provided by the terminal device; Acquire the directed graph adjacency table data according to the data upload instruction; The method further comprises: receiving a label selection instruction for at least one edge label sent by the terminal device, wherein at least one selectable edge label is displayed on an interactive interface provided by the terminal device; Determine a target edge label according to the label selection instruction; The method further comprises: Receiving an order setting instruction for an order adjustment control sent by the terminal device, wherein the order adjustment control is displayed on an interactive interface provided by the terminal device; Determining a target order according to the order setting instruction; The method further comprises: Receiving a node viewing instruction for any node sent by the terminal device, wherein at least one node is displayed on an interactive interface provided by the terminal device; If the node viewing instruction carries a first-order identifier, sending the estimated value of the number of first-order neighboring edges corresponding to the arbitrary node to the terminal device, so that the terminal device displays the estimated value of the number of first-order neighboring edges corresponding to the arbitrary node; If the node viewing instruction carries a second-order identifier, the estimated value of the second-order neighboring edges corresponding to the arbitrary node is sent to the terminal device, so that the terminal device displays the estimated value of the second-order neighboring edges corresponding to the arbitrary node.

13. A device for determining an estimated value of the number of neighborhood edges, characterized in that: include: An acquisition module, configured to acquire directed graph adjacency table data, wherein the directed graph adjacency table data represents adjacency table data corresponding to a directed graph, the directed graph includes K nodes, each piece of data in the directed graph adjacency table data includes a node number of a source node and a node number of a tail node corresponding to a directed edge, the directed graph adjacency table data is stored on at least two executors, and K is an integer greater than 1; A generating module, configured to generate a node aggregation table according to the directed graph adjacency table data, wherein each data of the node aggregation table includes a node number of a source node, a node number set of a tail node, and an edge number set, wherein the node number set includes at least one node number, and the edge number set includes at least one edge number; A statistics module is used for counting the number of neighboring edges of each tail node according to the node aggregation table for each executor, so that each executor can respectively obtain the first round of update information of each tail node; A sending module is used to send the first round of update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first round of update information of each tail node, wherein the first-order read counter stores the estimated value of the first-order neighborhood edge number.

14. The device according to claim 13, characterized in that The acquisition module is specifically used to acquire undirected graph adjacency table data, wherein the undirected graph adjacency table data represents the adjacency table data corresponding to the undirected graph, the undirected graph includes the K nodes, and each forward data in the undirected graph adjacency table data includes the node number of the source node and the node number of the tail node; For each forward data in the undirected graph adjacency table data, generate reverse data corresponding to each forward data, wherein the source node in the reverse data is the tail node in the corresponding forward data, and the tail node in the reverse data is the source node in the corresponding forward data; The directed graph adjacency table data is generated according to each piece of forward data and the reverse data corresponding to each piece of forward data, wherein both the forward data and the reverse data belong to the data in the directed graph adjacency table data.

15. The device according to claim 13, characterized in that The generating module is specifically used to encode the edge corresponding to each piece of data according to the directed graph adjacency table data to obtain the edge number corresponding to each edge; Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, and the edge number; According to the node-edge relationship table, aggregation processing is performed according to the node number of the source node to obtain the node aggregation table.

16. The device according to claim 13, characterized in that Each piece of data in the directed graph adjacency table data also includes an edge label corresponding to the edge, and the node aggregation table also includes an edge label set, and the edge label set includes at least one edge label; The generating module is specifically used to encode the edge corresponding to each piece of data according to the directed graph adjacency table data to obtain the edge number corresponding to each edge; Generate a node-edge relationship table according to the directed graph adjacency table data and the edge number corresponding to each edge, wherein each data in the node-edge relationship table includes the node number of the source node, the node number of the tail node, the edge number and the edge label; According to the node-edge relationship table, node numbers of the same source nodes are aggregated to obtain the node aggregation table.

17. The device according to claim 13, characterized in that The statistical module is specifically used to determine, for each executor, a counter corresponding to each neighbor edge in the neighbor edge set according to the node aggregation table, wherein each neighbor edge corresponds to a counter, and the counter stores an edge number and a cardinality estimate, and the cardinality estimate is 1; For each executor, the counter corresponding to each neighborhood edge and the counter of the corresponding tail node are fused to obtain the first round of update information of each tail node.

18. The device according to claim 13, characterized in that The statistical module is specifically used to determine, for each executor, a counter corresponding to each neighbor edge in the neighbor edge set according to the node aggregation table, wherein each neighbor edge corresponds to a counter, the counter stores an edge label, an edge number, and a cardinality estimate, and the cardinality estimate is 1 when the edge label is a target edge label, and the cardinality estimate is 0 when the edge label is a non-target edge label; For each executor, the counter corresponding to each neighborhood edge and the counter of the corresponding tail node are fused to obtain the first round of update information of each tail node.

19. The device according to claim 13, characterized in that The sending module is specifically used to send the first round of update information of the tail node to the parameter server through each executor after the parameter server creates the counter matrix, so that the parameter server fuses the zero-order write counter of the corresponding node with the update information of the first round according to the number of each tail node and the update information of the first round, until the first round of iteration is completed, and the parameter server determines the corresponding first-order read counter according to the first-order write counter of each node in the K nodes stored in the counter matrix, wherein the initial cardinality estimate of the zero-order write counter is 0.

20. The device according to claim 13, characterized in that Also included is a determination module; The acquisition module is further configured to send the first-round update information of each tail node to the parameter server through each executor, so that the parameter server determines the first-order read counter corresponding to each node in the K nodes according to the first-round update information of each tail node, and then acquire the first-order read counter corresponding to the source node from the parameter server for each executor; The determination module is further configured to determine, for each executor, a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node; The determination module is further configured to determine, for each executor, the second round update information corresponding to each tail node by fusing the first-order read counters of all source nodes corresponding to each tail node in the tail node set; The sending module is also used to send the second round update information of each tail node to the parameter server through each executor, so that the parameter server fuses the first-order write counter of the corresponding node with the second round update information according to the number of each tail node and the second round update information until the second round of iteration is completed. The parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, wherein the second-order read counter stores the estimated value of the second-order neighborhood edge number.

21. The device according to claim 20, characterized in that The determination module is specifically configured to obtain, for each executor, a first-order read counter of a source node corresponding to each tail node in the tail node set; For each executor, a fusion process is performed on the first-order read counters of the source nodes corresponding to each tail node to obtain the second round update information corresponding to each tail node.

22. The device according to claim 20, characterized in that The sending module is specifically used to send the second round update information of each tail node to the parameter server through each executor, so that the parameter server fuses the first-order write counter of the same tail node with the second round update information according to the second round update information of each tail node and the node number of each tail node to obtain a second-order write counter until the second round of iteration is completed. The parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes.

23. The device according to claim 20, characterized in that The acquisition module sends the second-round update information of each tail node to the parameter server through each executor, so that the parameter server performs a fusion process on the first-order write counter of the corresponding node and the second-round update information according to the number of each tail node and the second-round update information until the second round of iteration is completed. After the parameter server determines the corresponding second-order read counter according to the second-order write counter of each node in the K nodes, the parameter server acquires the second-order read counter corresponding to the source node from the parameter server for each executor; The determination module is further configured to determine, for each executor, a tail node set corresponding to the source node, wherein the tail node set includes at least one tail node; The determination module is further configured to determine, for each executor, the third round of update information corresponding to each tail node by fusing the second-order read counters of all source nodes corresponding to each tail node in the tail node set; The sending module is also used to send the third round update information of each tail node to the parameter server through each executor, so that the parameter server fuses the second-order write counter of the corresponding node with the third round update information according to the number of each tail node and the third round update information until the third round of iteration is completed. The parameter server determines the corresponding third-order read counter according to the third-order write counter of each node in the K nodes, wherein the third-order read counter stores the estimated value of the third-order neighborhood edge number.

24. The device according to any one of claims 13 to 23, characterized in that Also includes a receiving module and a sending module; The acquisition module is specifically used to receive a data upload instruction for a data upload control sent by a terminal device, wherein the data upload control is displayed on an interactive interface provided by the terminal device; Acquire the directed graph adjacency table data according to the data upload instruction; The receiving module is configured to receive a label selection instruction for at least one edge label sent by the terminal device, wherein at least one selectable edge label is displayed on an interactive interface provided by the terminal device; The determination module is used to determine the target edge label according to the label selection instruction; The receiving module is further configured to receive an order setting instruction for an order adjustment control sent by the terminal device, wherein the order adjustment control is displayed on an interactive interface provided by the terminal device; The determination module is further used to determine the target order according to the order setting instruction; The receiving module is further configured to receive a node viewing instruction for any node sent by the terminal device, wherein at least one node is displayed on the interactive interface provided by the terminal device; The sending module is used for sending the estimated value of the number of first-order neighbor edges corresponding to the arbitrary node to the terminal device if the node viewing instruction carries the first-order identifier, so that the terminal device displays the estimated value of the number of first-order neighbor edges corresponding to the arbitrary node; The sending module is also used to send the estimated value of the second-order neighboring edges corresponding to any one of the nodes to the terminal device if the node viewing instruction carries a second-order identifier, so that the terminal device displays the estimated value of the second-order neighboring edges corresponding to any one of the nodes.

25. A computer device, characterized in that: include: Memory, processor, and bus system; Wherein, the memory is used to store programs; The processor is used to execute the program in the memory, and the processor is used to execute the determination method according to any one of claims 1 to 12 according to the instructions in the program code; The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.

26. A computer-readable storage medium comprising instructions, which, when executed on a computer, causes the computer to execute the determination method according to any one of claims 1 to 12.

27. A computer program product comprising a computer program and instructions, characterized in that When the computer program / instructions are executed by a processor, the determination method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Graph training system, data access method and device, electronic equipment and storage medium

    CN110751275A

  • Method, system and equipment for counting number of points and edges in graph database and storage medium

    CN112988827A