Graph data system and device

By storing multiple snapshots in the graph data system and including all points of the previous snapshot, the problem of inefficient incremental map loading and playback in the prior art is solved, and efficient data processing and convenient user operations are achieved.

CN120196589APending Publication Date: 2025-06-24HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311766204.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is inefficient in the loading and playback of incremental graphs, and requires frequent replacement of pages and recovery of incremental data, resulting in inconvenience to users.

Method used

Design a graph data system to realize efficient incremental graph loading playback by storing multiple snapshots, each snapshot including a point edge index table and an edge list. The system contains all points of the previous snapshot in the set of points for each snapshot, and uses the deletion list and the deletion list to efficiently apply the deletion operation.

Benefits of technology

It improves data processing efficiency, simplifies user operations, can quickly query effective points and effective edges, and significantly improves the convenience of graph analysis services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196589A_ABST
    Figure CN120196589A_ABST
Patent Text Reader

Abstract

According to the graph data system provided by the invention, based on the scheme provided by the invention, a plurality of snapshots are stored in the graph data system, and the point set corresponding to each snapshot comprises all points in the point set corresponding to the previous snapshot, so that the operation of a user is facilitated, and the data processing efficiency is improved. And in a loading playback stage, the deleted points and the deleted edges can be efficiently applied, so that whether a certain point and a certain edge are effective or not can be quickly inquired when a subsequent graph algorithm obtains a certain point associated edge set in each snapshot, and graph analysis business can be conveniently carried out. On the whole, according to the scheme provided by the invention, the service experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of graph analysis, and more specifically, to a graph data system and apparatus. Background Art

[0002] In recent years, the global big data has entered an accelerated development period, and the data volume has grown exponentially. The data generated by the association relationships between different individuals in big data is presented in the form of a "graph". Here, the "graph" does not refer to a literal picture or image, but rather in terms of "graph theory" in mathematics, and can be understood as a data structure composed of "points" and "edges". Vertices are equivalent to nodes, and the association relationships between vertices are called "edges". And "incremental data" can be understood as the newly added "points" and "edges" in the graph data structure.

[0003] Loading the file of the stored incremental graph into memory and through a certain method, making the incremental changes of the graph data reflected in the organizational form in memory and being able to be recognized by subsequent graph analysis processes. This process can be understood as "incremental graph loading and playback". Existing methods need to replace pages during the incremental graph loading and playback process and need to repeatedly restore the incremental data through the continue flag between multiple old pages, with low efficiency and inconvenient for user operation. Therefore, a solution is needed that can more conveniently perform graph analysis operations and improve data processing efficiency. Summary of the Invention

[0004] This application provides a graph data system in which user operation is facilitated and data processing efficiency is improved.

[0005] In a first aspect, a graph data system is provided. The graph data system may include multiple units, modules, apparatuses for implementing graph data processing functions. Alternatively, the graph data system may also include constituent components (such as chips or circuits) for implementing graph data processing functions, and this is not limited.

[0006] The graph data system stores M snapshots. Each of the M snapshots includes a point-edge index table and an edge list. The point-edge index table in the i-th snapshot among the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot. Among them, the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot. The i-th point set includes the points in the (i - 1)-th point set and the newly added points in the incremental data corresponding to the i-th snapshot. The (i - 1)-th point set is the point set corresponding to the point-edge index table in the (i - 1)-th snapshot among the M snapshots. i is an integer greater than 1, and M is an integer greater than 1.

[0007] It can also be understood that in this application, the points in the point set in the point-edge index table of the i-th snapshot actually include the points in the (i-1)-th snapshot and the newly added points in the i-th snapshot. In other words, in this application, the points corresponding to the i-th snapshot actually include the points in all previous snapshots and the newly added points in the i-th snapshot.

[0008] For example, the graph data system can be a graph database or a file system. In this application, the name of the graph data system is not limited, and any device or module that can implement the graph data processing function belongs to the protection scope of this application.

[0009] Exemplarily, the representation method of snapshots in this application can be the compressed sparse row (CSR) representation method.

[0010] In the solution of this application, the point set corresponding to each snapshot includes all the points in the point set corresponding to the previous snapshot, which is convenient for user operation and improves data processing efficiency. And in the loading and playback stage, the deleted points and deleted edges can be efficiently applied, which is convenient for subsequent graph algorithms to quickly query whether a certain point or a certain edge is valid when obtaining the associated edge set of a certain point in each snapshot, and it is more convenient to perform graph analysis services. Generally speaking, the solution provided by this application improves the user's business experience.

[0011] In combination with the first aspect, in one implementation, the i-th snapshot further includes a deleted point list and / or a deleted edge list, where the deleted point list is used to indicate the points deleted in the incremental data corresponding to the i-th snapshot; the deleted edge list is used to indicate the edges deleted in the incremental data corresponding to the i-th snapshot, and the j-th snapshot where the deleted edges are located, where the j-th snapshot is one of the M snapshots, and j is an integer less than i.

[0012] The "snapshot" in this application is used to indicate the updated points and updated edges in the graph. Therefore, the snapshot can also include a deleted point list and / or a deleted edge list, which is used to indicate at least one of the deleted points, deleted edges, and newly added edges.

[0013] In combination with the first aspect, in one implementation, the edge list in the first snapshot among the M snapshots is used to indicate the original edges in the original data corresponding to the first snapshot, and the point-edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot, and the first point set includes the original points in the original data.

[0014] It should be noted that in this application, the first snapshot can be understood as the original snapshot. That is, it can be understood as the original snapshot corresponding to the original data. The relationships between the points and edges indicated in this snapshot can be understood as the point-edge relationships in the original graph. The edge list in the first snapshot is used to indicate the original edges in the original data corresponding to the first snapshot, and the point-edge index table in the first snapshot is used to indicate the index relationships between the points in the first point set corresponding to the first snapshot and the edges in the edge list of the first snapshot. The first point set includes the original points in the original data. At this time, there is no point deletion list and edge deletion list in the first snapshot. Or, it can also be understood that in the first snapshot, the contents of the point deletion list and the edge deletion list are blank.

[0015] In combination with the first aspect, in one implementation, the graph data system is further configured to receive a second request message from the first module. The second request message is used to request the graph data system to export an incremental record, and the incremental record includes at least one of the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges. The graph data system is configured to generate a snapshot corresponding to the incremental record according to the incremental record. The snapshot includes at least one of the following: a point-edge index table, an edge list, a point deletion list, and an edge deletion list.

[0016] It should be noted that in this application, the "incremental record" records a list of point and / or edge updates. For example, the incremental record includes the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges. An incremental snapshot corresponding to the incremental record can be generated through the data in the incremental record, and the incremental snapshot can be used to describe the changes of points and edges in the incremental record. Exemplarily, the incremental record is in the form of a list. For example, a list of newly added points, a list of newly added edges, a list of deleted points, and a list of deleted edges. Exemplarily, the snapshot includes: a point-edge index table, an edge list, a point deletion list, and an edge deletion list.

[0017] In one possible implementation, the graph data system can receive a second request message from the first module. The second request message is used to request the graph data system to export an incremental record, and subsequently, the graph database can generate a snapshot corresponding to the incremental record according to the incremental record.

[0018] In one possible implementation, the first module can trigger the graph data system regularly or quantitatively, so that the graph data system generates an "incremental record". For example, the first module can be an "incremental trigger module"; or, for example, the first module can be an incremental monitor module. In this application, the specific name of the first module is not limited.

[0019] In combination with the first aspect, in one implementation, the graph data system further includes: receiving a first request message from a user, where the first request message is used to request graph analysis of a target graph using a target algorithm; performing a point query and / or an edge query on the target graph according to M snapshots; sending a first response message, where the first response message carries the analysis result of the target algorithm on the target graph, and the analysis result is generated based on the point query and / or the edge query.

[0020] For example, the user can indicate the target graph and the target algorithm.

[0021] It should be noted that generally, in the process of the target algorithm performing graph analysis on the target graph, the target algorithm needs to perform a point query and / or an edge query on the replayed graph multiple times, and finally obtain the analysis result based on the point query and / or the edge query result.

[0022] Exemplarily, the graph data system is used to receive a first request message from a user, where the first request message is used to request graph analysis of a target graph using a target algorithm; the graph data system is used to perform a point query and / or an edge query on the target graph according to M snapshots; send a first response message, where the first response message carries the analysis result of the target algorithm on the target graph, and the analysis result is generated based on the point query and / or the edge query.

[0023] In this application, before performing graph analysis, it is necessary to first complete the loading and replay of the graph. When performing graph loading and replay, it is necessary to maintain the point bitmap and the edge bitmap. For example, the following implementation can be adopted when specifically maintaining the point bitmap and the edge bitmap. In one possible implementation, the graph data system can read and scan the snapshots, obtain the snapshot number 2, the total number of points L p as 7, and the number of edges [7, 3] corresponding to each snapshot, and successively establish an array b p of the point bitmap with a length of L p = 7, and an array b e of the edge bitmap. Among them, the length of b e [1] is 7, the length of b e [2] is 3, and all bit positions of these bitmaps are set to 0. The graph data system can scan the snapshots again, and apply the information on the delete point list and the delete edge list of each snapshot to the point bitmap and the edge bitmap, that is, set the bit at the corresponding position in the corresponding point bitmap and edge bitmap for the records that appear in these two types of lists to 1. The above steps complete the replay of the graph.

[0024] In combination with the first aspect, in one implementation, the graph data system is used to perform a point query on the target graph according to the M snapshots, including: the graph data system is used to determine the set of existing points according to the Mth point set corresponding to the point-edge index table in the Mth snapshot among the M snapshots; the graph data system is used to determine the deleted points according to the list of deleted points included in each of the M snapshots; the graph data system is used to determine the set of valid points according to the set of existing points and the deleted points.

[0025] The above implementation introduces the specific implementation of querying the set of valid points during the point query process. Specifically, since each snapshot includes a point-edge index table, all the points that already exist in the current snapshot can actually be obtained based on this point-edge index table. Then, based on the list of deleted points in this snapshot, the points deleted in the current snapshot are determined, so as to determine the set of valid points in the current snapshot.

[0026] In combination with the first aspect, in one implementation, the graph data system is used to perform a point query on the target graph according to the M snapshots, including: the graph data system is used to determine a point bitmap according to the point-edge index table in the Mth snapshot and the list of deleted points in each of the snapshots. The point bitmap is used to indicate the respective states of all the points included in the M snapshots, and the states include the deleted state and the undeleted state. Among them, the bit length corresponding to the point bitmap is the same as the number of points in the Mth point set corresponding to the Mth snapshot, and the bit value corresponding to the deleted points in the point bitmap is different from the bit value corresponding to the undeleted points; the graph data system is used to determine the set of valid points according to the point bitmap.

[0027] In combination with the first aspect, in one implementation, the graph data system is used to perform an edge query on the target graph according to the M snapshots, including: the graph data system is used to determine the set of existing edges corresponding to each valid point according to the edge list in each of the M snapshots; the graph data system is used to determine the deleted edges corresponding to each valid point according to the list of deleted edges in each of the M snapshots; the graph data system is used to determine the set of valid edges corresponding to each valid point according to the set of existing edges corresponding to each valid point and the deleted edges corresponding to each valid point.

[0028] In combination with the first aspect, in one implementation, the graph data system is used to perform edge queries on a target graph based on M snapshots, including: the graph data system is used to determine, according to the edge list and the edge deletion list in each of the M snapshots, an edge bitmap corresponding to each snapshot, where the edge bitmap is used to indicate the status of the edges in each snapshot, and the status includes the deleted status and the undeleted status. Among them, the bit length of the bitmap of the edges corresponding to each snapshot is the same as the number of edges in the edge list of the corresponding snapshot, and the bit value corresponding to the deleted edge in each edge bitmap is different from the bit value corresponding to the undeleted edge; the graph data system is used to determine, according to the edge bitmap, the set of valid edges corresponding to each valid point.

[0029] The scenarios of querying the set of valid points and the scenarios of querying the set of valid edges corresponding to each valid point are introduced above respectively. In another possible application scenario, the graph data system can query whether a certain point specified by the user is a valid point. If it is a valid point, the graph data system will further query the valid edges corresponding to the valid point.

[0030] In the second aspect, a graph data system device is provided, and this device is used to execute the method in any one of the possible implementation manners in the first aspect above. Specifically, the device may include units and / or modules for executing the method in any one of the possible implementation manners in the first aspect, such as a transceiver unit and / or a processing unit.

[0031] In one implementation, the device is a computing device, and the communication unit may be a transceiver, or an input / output interface; the processing unit may be at least one processor. Optionally, the transceiver may be a transceiver circuit. Optionally, the input / output interface may be an input / output circuit.

[0032] In another implementation, the device is a chip, a chip system or a circuit of a computing device for implementing graph analysis functions. When the device is a chip, a chip system or a circuit for implementing graph analysis functions, the communication unit may be an input / output interface, an interface circuit, an output circuit, an input circuit, a pin or a related circuit, etc. on the chip, the chip system or the circuit; the processing unit may be at least one processor, a processing circuit or a logic circuit, etc.

[0033] In the third aspect, a device of a graph data system is provided, and this device includes: at least one processor, which is used to execute the computer program or instruction stored in the memory to execute the method in any one of the possible implementation manners in the first aspect above. Optionally, the device further includes a memory for storing the computer program or instruction. Optionally, the device further includes a communication interface, and the processor reads the computer program or instruction stored in the memory through the communication interface.

[0034] Fourth aspect, the present application provides a processor, including: an input circuit, an output circuit, and a processing circuit. The processing circuit is configured to receive signals through the input circuit and transmit signals through the output circuit, so that the processor executes the method in any possible implementation manner of the first aspect.

[0035] In a specific implementation process, the above-mentioned processor may be one or more chips, the input circuit may be an input port, the output circuit may be an output port, and the processing circuit may be transistors, gate circuits, flip-flops, and various logic circuits, etc. The input signals received by the input circuit may be received and input by, for example, but not limited to, a transceiver, and the signals output by the output circuit may be output to, for example, but not limited to, a transmitter and transmitted by the transmitter, and the input circuit and the output circuit may be the same circuit, which is used as the input circuit and the output circuit at different times respectively. The embodiments of the present application do not limit the specific implementation manners of the processor and various circuits.

[0036] For operations such as sending and obtaining / receiving involved in the processor, if there is no special description, or if it does not conflict with its actual function or internal logic in the relevant description, it may be understood as the output and reception, input, etc. operations of the processor, or it may also be understood as the sending and receiving operations performed by the radio frequency circuit and the antenna. The present application does not limit this.

[0037] Fifth aspect, a processing device is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory, and may receive signals through a transceiver and transmit signals through a transmitter to execute the method in any possible implementation manner of the first aspect.

[0038] Optionally, the processor is one or more, and the memory is one or more.

[0039] Optionally, the memory may be integrated with the processor, or the memory and the processor are separately arranged.

[0040] In a specific implementation process, the memory may be a non-transitory memory, such as a read only memory (ROM), which may be integrated with the processor on the same chip or may be separately arranged on different chips. The embodiments of the present application do not limit the type of the memory and the setting manner of the memory and the processor.

[0041] It should be understood that the relevant data interaction process, such as sending indication information, can be a process of outputting indication information from the processor, and receiving capability information can be a process of the processor receiving input capability information. Specifically, the data output by the processor can be output to the transmitter, and the input data received by the processor can come from the transceiver. Among them, the transmitter and the transceiver can be collectively referred to as the transceiver.

[0042] The processing device in the above fifth aspect can be one or more chips. The processor in the processing device can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading the software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.

[0043] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable medium stores program code for a device to execute. The program code includes methods for executing any possible implementation manner in the above first aspect.

[0044] In a seventh aspect, a computer program product including instructions is provided. When the computer program product runs on a computer, the computer is caused to execute the methods in any possible implementation manner in the above first aspect.

[0045] In an eighth aspect, a chip system is provided, including a processor for calling and running a computer program from a memory, so that a device equipped with the chip system executes the methods in the various implementation manners in the above first aspect. Description of the Drawings

[0046] Figure 1 is a schematic flowchart of full export and incremental update in graph analysis provided by this application.

[0047] Figure 2 is a schematic block diagram of a system architecture applicable to this application.

[0048] Figure 3 is another schematic block diagram of a system architecture applicable to this application.

[0049] Figure 4 is a schematic diagram of snapshot #1 at time T1 provided by this application.

[0050] Figure 5 is a schematic diagram of snapshot #2 at time T2 provided by this application.

[0051] Figure 6 is a schematic diagram of a dot map and an edge map provided by this application.

[0052] Figure 7 It is a schematic flowchart of loading and playback provided by this application.

[0053] Figure 8 It is a schematic flowchart of the graph analysis method 800 provided by this application.

[0054] Figure 9 It is a schematic block diagram of a device 900 of a graph data system proposed by this application. Detailed implementation manners

[0055] Next, the technical solutions in this application will be described in conjunction with the accompanying drawings.

[0056] To facilitate the understanding of the technical solutions in this application, some individual technical terms in this application will be briefly introduced first.

[0057] 1. Graph

[0058] In recent years, the global big data has entered an accelerated development period, and the data volume has grown exponentially. The data generated by the association relationships between different individuals in big data is presented in the form of a "graph". The "graph" here does not refer to the literal meaning of pictures or images, but in terms of "graph theory" in mathematics, and can be understood as a data structure composed of "points" and "edges". Vertices are equivalent to nodes, and the association relationships between vertices are called "edges". And "incremental data" can be understood as the newly added "points" and "edges" in the graph data structure. For example: three people sitting in the office, and these three people are three points. The relationships between the three people are called edges, such as: colleague relationship, junior sister relationship, project cooperation relationship, and so on.

[0059] 2. Graph analysis

[0060] Graph analysis is based on graph-based methods to analyze the connected data. The main contents of graph analysis can be, for example: querying graph data, using basic statistical information, visually exploring the graph, displaying the graph, or preprocessing the graph information and then merging it into machine learning tasks. "Graph query" is usually used for the analysis of local data, while graph computing usually involves the entire graph and iterative analysis. Graph analysis focuses on analyzing the strength and direction of the relationships between entities in graph data, so as to discover insights and assist in decision-making. Graph analysis obtains the results of a certain evaluation or information extraction of the input graph data by running algorithms that can identify the graph data structure.

[0061] Graph analysis relies on many different graph analysis algorithms to achieve many different types of analysis. For example: association analysis algorithms, path analysis algorithms, classification algorithms, clustering algorithms, time series analysis algorithms, etc.

[0062] 3. Graph analysis application scenarios

[0063] (1) Social network analysis scenario: Social networks are a very common type of graph data, representing social relationships between various individuals or organizations. Graph data can present complex social network relationships, making it easy for users to conduct further analysis. For example, in a typical social network, there are often "who knows whom, who attended what school, and who lives where permanently". Social relationships can be managed through graph analysis to achieve friend recommendations, etc.

[0064] (2) E-commerce application scenario: E-commerce is a core business on the Internet. In this scenario, nodes are divided into two categories: users and products. The existing relationships include browsing, collecting, purchasing, etc. There can be multiple relationships between users and products, such as both a collection relationship and a purchase relationship. Such complex data scenarios can be easily described using property graphs. E-commerce has given rise to a well-known technology application - "recommendation systems". The interaction relationships between users and products reflect users' shopping preferences.

[0065] (3) Transportation network application scenario: Transportation networks have various forms. For example, in a subway network, each station is regarded as a node, and the connectivity between stations is regarded as an edge. Usually, in transportation networks, we are more concerned about problems related to path planning: such as the shortest path problem. Another example is that we use traffic flow as an attribute of nodes in the network to predict changes in future traffic flow. A typical application scenario is map navigation.

[0066] 4. Offline graph analysis: Export the graph data stored in the graph database and perform graph analysis in a non-online form.

[0067] 5. Graph partitioning: The process of dividing large-scale graph data into several slices, enabling multiple computer processes to simultaneously run the same graph algorithms on these slices.

[0068] 6. Incremental graph: Based on the existing graph partitioning, the incremental data exported from the database is partitioned through graph partitioning and attached to the original graph partitioning result as an incremental data structure. The so-called "incremental graph" can be understood as real-time data. Incremental graph analysis only needs to recalculate the changed part of the data, and then fuse the incremental result with the original graph calculation result through some algorithms to efficiently obtain the new calculation result.

[0069] 7. Incremental graph loading and playback: Load the file storing the incremental graph into memory and, through a certain method, make the incremental changes in the graph data reflected in the organizational form in memory and be recognizable by subsequent graph analysis processes. It can also be understood that the graph data represented in memory after playback is consistent with the graph data in the database.

[0070] Figure 1A schematic flowchart of full-amount export and incremental update in graph analysis is shown. Among them, for "full-amount export", it is necessary to export full-amount data based on the data in the graph database (the "full amount" can be understood as historical data), then perform distributed graph partitioning, then perform graph loading, and restore the full-amount graph in memory before it can be used as the input for the algorithm. In the scenario of data update, especially for large-scale graphs with tens of billions of nodes and edges, the time proportion required for the database to re-export the full-amount graph and perform graph partitioning will be very large. The incremental update mechanism for offline graph analysis can retain the topology and attribute data of the completed graph partitioning to the greatest extent possible, avoiding re-partitioning the entire graph. However, in the incremental update scenario, incremental data can be exported through the business system to obtain incremental snapshot #1, which can be used to describe the incremental data. Then, the original graph data, that is, incremental snapshot #0, can be determined through the incremental graph storage. The incremental snapshot #0 and the incremental snapshot #1 are loaded, and the incremental graph is loaded and played back in memory to restore the graph. Finally, the restored graph is analyzed by the graph analysis algorithm.

[0071] Generally, graph incremental update needs to consider the efficiency of graph storage, graph loading and playback, and graph algorithm's access to the graph after playback. The existing method requires replacing pages during the incremental graph loading and playback process, and needs to repeatedly restore the incremental data through the continue flag between multiple old pages, with low efficiency and inconvenience for user operation. In view of this, the present application provides a method for graph analysis, providing a simple and efficient incremental graph loading and playback method, avoiding complex jump recovery, being convenient for user operation, and improving the efficiency of data processing.

[0072] Figure 2 is a schematic block diagram of a system architecture applicable to the present application, as Figure 2 shown, the system architecture includes a front-end system 210, a business system 220, a graph partitioning system 230, and a graph analysis system 240. In a possible scenario, the user initiates an incremental graph analysis request for a certain graph on the terminal side. The server of the front-end system 210 receives the request and drives the business system 220 to generate an incremental record corresponding to this graph, and imports the incremental record into the graph partitioning system 230 for incremental update, thereby generating a new incremental snapshot corresponding to the incremental record. The new incremental snapshot will be stored in the incremental graph storage device. In addition, the incremental graph storage device also stores the previous existing snapshot set. The graph analysis system 240 performs graph loading and playback based on the existing snapshot set and the new incremental snapshot, and the graph analysis algorithm starts the graph algorithm based on the graph data in the memory after playback, outputs the graph analysis result, and feeds back the graph analysis result to the user.

[0073] It should be noted that in this application, the "incremental record" records a list of point and / or edge updates. For example, the incremental record includes the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges. The incremental snapshot corresponding to the incremental record can be generated through the data in the incremental record, and the incremental snapshot can be used to describe the changes of points and edges in the incremental record. Exemplarily, the incremental record is in the form of a list. For example, a list of newly added points, a list of newly added edges, a list of deleted points, and a list of deleted edges. Exemplarily, the snapshot includes: a point-edge index table, an edge list, a deleted point list, and a deleted edge list.

[0074] In a possible implementation manner, the graph data system can receive a second request message from the first module, and the second request message is used to request the graph data system to export an incremental record. Subsequently, the graph database can generate a snapshot corresponding to the incremental record according to the incremental record.

[0075] In a possible implementation manner, the first module can trigger the graph data system regularly or quantitatively, so that the graph data system generates an "incremental record". For example, the first module can be an "incremental trigger module" ( Figure 2 not shown in the figure); for another example, the first module can be an incremental monitor module. In this application, the specific name of the first module is not limited.

[0076] Figure 3 is a schematic block diagram of an incremental update system shown in this application, Figure 3 which can be understood as a more specific description of the Figure 2 process. As Figure 3 shown, for a specific graph, the graph database (or, file system) 310 first exports the incremental record corresponding to the graph, and the graph segmentation module 320 performs graph segmentation operations on it to generate an original snapshot (for example, it can be a memory object or a file combination), and the incremental graph storage module 340 saves the original snapshot. Exemplarily, the original snapshot can be represented in the structure of an incremental graph. Subsequently, if there are changes in the graph structure, or point data, or edge data in the graph database, at this time, the graph database 310 can export the incremental records corresponding to these changes, and the incremental update module 330 performs graph segmentation of the incremental data, and stores the newly generated incremental snapshot into the incremental graph storage module 340. The graph loading module 350 can load and replay the existing snapshot set (including incremental snapshots and original snapshots) provided by the incremental graph storage, and use the replay result as the input of the graph algorithm, and the graph analysis algorithm module 360 performs graph analysis.

[0077] It should be noted that in this application Figure 2 and Figure 3Each module in [[ ]] may be multiple hosts (e.g., a computer cluster) that undertake corresponding functions, may also be multiple processes in a computer, or may be multiple modules in multiple processes, etc., without limitation. Any device, apparatus, or chip that can implement the above functions falls within the protection scope of this application.

[0078] The following introduces the representation method of the incremental graph in this application. For example, taking the compressed sparse row (CSR) representation method as an example, the representation method of the incremental graph is introduced. It should be noted that in this application, the representation method of the incremental graph is not limited to the description method in the CSR format, and other description methods can also be used for representation, without limitation. The description method in the CSR format here is only an example.

[0079] Figure 4 shows the snapshot at time T1, and the snapshot at time T1 is denoted as snapshot #1, as Figure 4 shown in (a) in [[ ]]. Snapshot #1 includes a vertex-edge index table and an edge list. From the vertex-edge index table and the edge list of snapshot #1, the vertex-edge relationship described by this snapshot #1 can be obtained as Figure 4 the graph shown in (b) in [[ ]]. Figure 4 In (a) in [[ ]], it can be understood that vertices 0 to 5 mean there are 6 vertices in this snapshot, which are "0", "1", "2", "3", "4", and "5" respectively. Each number in the edge list has the same meaning as these 6 vertices. For example, "1" in the edge list refers to vertex 1, and "5" in the edge list refers to vertex 5. Among them, the "vertex-edge index table" in the CSR format describes the edge information corresponding to a certain vertex (which can also be understood as "out-edge", that is, each edge corresponding to this vertex starts from this vertex, and the destination vertex is another vertex). The "edge list" describes the information of the other vertex of the edge corresponding to a certain vertex (in the CSR format, it can also be understood as the destination vertex of a certain edge corresponding to this vertex). For example, from the "vertex-edge index table", the value corresponding to vertex i is n, and the value corresponding to vertex (i - 1) is m. Then the number of the other vertices of the edges corresponding to vertex i is (m to n - 1) (that is, the number of edges corresponding to vertex i is also (m to n - 1)). Among them, the first vertex of the edges corresponding to vertex i is the value corresponding to vertex m in the edge list, and the last vertex of the edges corresponding to vertex i is the value corresponding to vertex (n - 1) in the edge list.

[0080] The above CSR representation method can also be understood as follows: The "point-edge index table" in Snapshot #1 is used to indicate the index relationship between the points in the first point set and the edges in the "edge list" in Snapshot #1, where the "edge list" in Snapshot #1 is used to indicate the original edges in the original data corresponding to the first snapshot. The "point-edge index relationship" in this application actually indicates the relationship between the current points, or it can also be understood that for each of the current existing points, which edges it has and which point is the destination point of these edges. The following takes Figure 4 as an example to specifically introduce the CSR method.

[0081] Exemplarily, as Figure 4 shown in (a) of, in the "point-edge index table", the value corresponding to point 0 is "2", and the value corresponding to point 1 is "4", that is, it can be determined that there are two edges corresponding to point 1 (i.e., 2, (4 - 1)), and the destination points of these two edges should be found by looking at the values corresponding to points 2 and 3 in the edge list respectively. In the edge list, the value corresponding to point "2" is "4", so it can be determined that one edge of point 1 is from point 1 → point 4; in the edge list, the value corresponding to point 3 is 3, so it can be determined that the other edge of point 1 is from point 1 → point 3.

[0082] Exemplarily, as Figure 4 shown in (a) of, in the "point-edge index table", the value corresponding to point 2 is "5", and the value corresponding to point 1 is "4", that is, it can be determined that there is one edge corresponding to point 2 (i.e., 4), and the destination point of this one edge should be found by looking at the value corresponding to point 4 in the edge list. In the edge list, the value corresponding to point "4" is "4", so it can be determined that one edge of point 2 is from point 2 → point 4.

[0083] Exemplarily, as Figure 4 shown in (a) of, in the "point-edge index table", the value corresponding to point 3 is "7", and the value corresponding to point 2 is "5", that is, it can be determined that there are two edges corresponding to point 3 (i.e., 5, (7 - 1)), and the destination points of these two edges should be found by looking at the values corresponding to points 5 and 6 (not shown in the figure) in the edge list respectively. In the edge list, the value corresponding to point "5" is "4", so it can be determined that one edge of point 3 is from point 3 → point 4; in the edge list, the value corresponding to point 6 (not shown in the figure) is 5, so it can be determined that the other edge of point 3 is from point 3 → point 5.

[0084] To more clearly show the Figure 4 point-edge relationship described in (a) of, based on Figure 4 the (a) of, Figure 4 the (b) of is drawn. As mentioned above, in the CSR format, the number of edges corresponding to a certain point refers to the number of edges starting from that point. Based on Figure 4In (a) of [reference], there are two edges corresponding to point 0, which are from point 0 to point 1 and from point 0 to point 2 respectively; there are two edges corresponding to point 1, which are from point 1 to point 4 and from point 1 to point 3 respectively; there is one edge corresponding to point 2, which is from point 2 to point 4; there are two edges corresponding to point 3, which are from point 3 to point 4 and from point 3 to point 5 respectively; there are no corresponding edges for point 4 and point 5.

[0085] In addition, it can also be seen that the length of the edge list in snapshot #1 actually corresponds to the number of edges in the graph. In Figure 4 the length of the edge list in (a) of [reference] is 7, corresponding to a total of 7 edges in (b) of Figure 4 [reference].

[0086] The above mainly introduces a method for looking up a table in CSR format. Next, the solution of this application will continue to be introduced by taking the CSR format as an example.

[0087] This application proposes that: M snapshots are stored in the graph data system, and each of the M snapshots includes a point-edge index table and an edge list. The point-edge index table in the i-th snapshot among the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot. Among them, the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot. The i-th point set includes the points in the (i - 1)-th point set and the newly added points in the incremental data corresponding to the i-th snapshot. The (i - 1)-th point set is the point set corresponding to the point-edge index table in the (i - 1)-th snapshot among the M snapshots. i is an integer greater than 1, and M is an integer greater than 1.

[0088] The "graph data system" in this application can be, for example, a graph database, or it can be, for example, a file system, and so on. It should be noted that the "graph data system" in this application can also refer to other database systems. Any device, apparatus, or module that can use the method provided in this application is within the protection scope of this application.

[0089] It can also be understood that in this application, the points in the point set in the point-edge index table in the i-th snapshot actually include the points in the (i - 1)-th snapshot and the newly added points in the i-th snapshot. In other words, in this application, the points corresponding to the i-th snapshot actually include the points in all previous snapshots and the newly added points in the i-th snapshot.

[0090] In a possible scenario, the i-th snapshot may further include a point deletion list and / or an edge deletion list. Among them, the point deletion list is used to indicate the points deleted in the incremental data corresponding to the i-th snapshot; the edge deletion list is used to indicate the edges deleted in the incremental data corresponding to the i-th snapshot, and the j-th snapshot where the deleted edge is located. Among them, the j-th snapshot is one of the M snapshots, and j is an integer less than i.

[0091] It should be noted that the first snapshot in this application can be understood as the original snapshot. That is, it can be understood as the original snapshot corresponding to the original data. The relationships between the points and edges indicated in this snapshot can be understood as the point-edge relationships in the original graph. The edge list in the first snapshot is used to indicate the original edges in the original data corresponding to the first snapshot, and the point-edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot. The first point set includes the original points in the original data. At this time, there is no deleted point list and deleted edge list in the first snapshot. Or, it can also be understood that in the first snapshot, the contents of the deleted point list and the deleted edge list are blank.

[0092] Next Figure 5 is a possible schematic diagram of the solution proposed in this application. Figure 5 Shows the snapshot at time T2. Assume that time T2 is a snapshot later than time T1, and the data in the graph at time T1 is updated at time T2. For example, the snapshot at time T2 is denoted as snapshot #2. As Figure 5 shown in (a) of [], snapshot #2 includes a point-edge index table, an edge list (which can also be understood as the newly added edge list), a deleted point list, and a deleted edge list, respectively describing the newly added points, newly added edges, deleted points, and deleted edges at time T2.

[0093] From Figure 5 in (a), it can be seen that a new point is added. Assume that this point is denoted as point "6". In this application, exemplarily, the indexes corresponding to the newly added points are sorted in sequence with the indexes corresponding to the previous points. For example, the indexes corresponding to the original points and the indexes corresponding to the newly added points can be arranged in ascending order in sequence. In addition, Figure 5 also includes a newly added edge list. From this list, it can be found that three new edges are added at time T1. The destination points of these three edges are point 2, point 6, and point 5 respectively. According to Figure 4As can be seen from the look-up table method introduced in (a) of [reference], in snapshot #2, in the "point-edge index table", the value corresponding to point 1 is "2", and the value corresponding to point 0 is "0". That is, it can be determined that two new edges are added to point 1 (i.e., 0, (2 - 1)). The destination points of these two edges should be found according to the values corresponding to point 0 and point 1 in the edge list respectively. In the edge list, the value corresponding to point 0 is "2", so it can be determined that one of the new edges added to point 1 is from point 1 → point 2; in the new edge list, the value corresponding to point 1 is 6, so it can be determined that the other new edge added to point 1 is from point 1 → point 6. In the point-edge index table, the value corresponding to point "6" is "3", and the value corresponding to point 5 is "2", so it can be determined that one new edge is added to point 6 (i.e., 2). In the edge list, the value corresponding to point 2 is "5", so it can be determined that one new edge added to point 6 is from point 1 → point 5.

[0094] From Figure 5 From the deleted point list in (a) of [reference], it can be seen that one point 4 is deleted. In this application, in the deleted point list, the index of the point is the index corresponding to a certain point in the point-edge index table. In other words, in this application, the corresponding index of the deleted point is the same as the index corresponding to the point in the above-mentioned point-edge index table, representing the same meaning.

[0095] From Figure 5 As can be seen from the deleted edge list in (a) of [reference], four edges in snapshot #0 are deleted. In this application, the deleted edge list includes the index of the snapshot where the deleted edge is located, and the index of the point corresponding to the edge list in that snapshot. As Figure 5 shown in (a) of [reference], "0:0" in the deleted edge list means that in snapshot #0, the edge corresponding to point 0 in the edge list is deleted, that is, the edge with the destination point "1" in the edge list of snapshot #0 is deleted. Since the edge corresponding to the destination point "1" in the edge list is actually an edge starting from point "0", that is, the edge from point 0 → point 1 in snapshot #1 is deleted. "0:2" means that in snapshot #0, the edge corresponding to point 2 in the edge list is deleted, that is, the edge with the destination point "4" in the edge list of snapshot #0 is deleted. Since one of the edges corresponding to point "1" in the edge list has the destination point 4, that is, the edge from point 1 → point 4 in snapshot #1 is deleted; "0:4" means that in snapshot #0, the edge corresponding to point 4 in the edge list is deleted, that is, the edge with the destination point "4" in the edge list of snapshot #0 is deleted. Since one of the edges corresponding to point "2" in the edge list has the destination point 4, that is, the edge from point 2 → point 4 in snapshot #0 is deleted. "0:5" means that in snapshot #0, the edge corresponding to point 5 in the edge list is deleted, that is, the edge with the destination point "4" in the edge list of snapshot #0 is deleted. Since one of the edges corresponding to point "3" in the edge list has the destination point 4, that is, the edge from point 3 → point 4 in snapshot #0 is deleted. To show more clearly Figure 5The newly added points, newly added edges, deleted points, and deleted edges described in (a) of Figure 5 are drawn based on (a) in Figure 5 (b) in Figure 5 and (c) in

[0096] to facilitate understanding of the technical solution provided by this application. Figure 4 and Figure 5 From the schematic diagrams of

[0097] It can be seen that in a possible scenario, the snapshot in this application includes an original snapshot and multiple incremental snapshots. Each incremental snapshot represents the newly added points and edges at a certain moment, as well as the deletion status of the points and edges included in the previous snapshot. In this application, each snapshot includes a point-edge index table, and the length of the point-edge index table is the sum of the number of points in the previous snapshot and the newly added points in this snapshot. Exemplarily, the arrangement order of the original points in the nth snapshot is the same as that in the (n + 1)th snapshot. In this application, each incremental snapshot also includes a list of newly added edges, a point deletion list, and an edge deletion list. Among them, the elements of the point deletion list contain the index (which can also be called "number") of a certain point, and the edge deletion list indicates the index of the snapshot where the deleted edge is located and the deleted edge.

[0097] The above Figure 5 mainly introduces the representation method of snapshots in the graph data system proposed in this application. Next, it mainly introduces how to perform graph loading and playback based on the stored snapshots. For example, in a possible scenario, the graph data system can read in all snapshots, and the graph data system is also used to maintain a point bitmap for marking deleted points and an edge bitmap for marking deleted edges. Applying the deleted points and deleted edges in all snapshots to the point bitmap and edge bitmap, at this time, it can be understood that the playback of a specific graph is completed.

[0098] Figure 6 shows the schematic diagrams of the point bitmap and edge bitmap determined based on the snapshots of Figure 4 and Figure 5 Among them, Figure 6 (a) in Figure 6 is the schematic diagram of the point bitmap,

[0099] As shown in (a) of Figure 6 , this point bitmap can also be understood as the "bitmap of deleted points". In the point bitmap, the bit length corresponding to the point bitmap is the same as the number of points in the point set corresponding to the last snapshot among the M snapshots. Taking Figure 5For example, assume there are two snapshots in total, namely snapshot #1 and snapshot #2. Then the bit length of the bitmap should be the number of points in the point set corresponding to snapshot #2. Since the point set corresponding to snapshot #2 is {0, 1, 2, 3, 4, 5, 6}, that is, there are 7 points in total, so the bit length of the bitmap is 7. Further, since the snapshot #2 includes a point deletion list, which indicates that the point "4" has been deleted. Therefore, in the bitmap, the value of the bit corresponding to the deleted point (e.g., point "4") needs to be set to "1", while the values of the bits corresponding to the remaining undeleted points are set to "0". Or, the value of the bit corresponding to the deleted point can also be set to "0", and the values of the bits corresponding to the undeleted points are set to "1".

[0100] As Figure 6 shown in (b) of Figure 4 , it can be seen that in the edge bitmap, each snapshot corresponds to an edge bitmap respectively. For example, assume there are two snapshots, namely snapshot #1 and snapshot #2. Then there should be two edge bitmaps, that is, snapshot #1 corresponds to an edge bitmap, and snapshot #2 corresponds to an edge bitmap. Among them, the bit length of the edge bitmap corresponding to snapshot #1 is the same as the number of edges in snapshot #1. It can be seen from the edge list of snapshot #1 in

[0101] that snapshot #1 includes a total of 7 edges. Therefore, the length of the edge list corresponding to snapshot #1 is 7. Further, it can be seen from snapshot #2 that 4 edges in snapshot #1 have been deleted, and the positions of the deleted edges correspond to the indices of the points. For example, based on the edge deletion list in snapshot #2, the bit value of the corresponding bit position can be set to "1", and the values of the bits corresponding to the remaining undeleted edges are set to "0". For snapshot #2, there are 3 newly added edges in the edge list of snapshot #2. Therefore, the length of the edge bitmap corresponding to snapshot #2 is 3. Since there are no deleted edges in snapshot #2, for example, the bit values of the edge bitmap corresponding to snapshot #2 can all be set to "0". p For example, when specifically maintaining the bitmap and the edge bitmap, the following implementation method can be adopted. In a possible implementation method, the graph data system can read and scan the snapshots, obtain the number of snapshots as 2, the total number of points L p as 7, and the number of edges of each snapshot [7, 3]. Then an array b p of bitmaps with a length of L e = 7 is established in sequence, as well as an array b e [1] with a length of 7, and b eThe length of [2] is 3, and all bits of these bitmaps are set to 0. The graph data system can scan the snapshot again and apply the information on the deletion point list and deletion edge list of each snapshot to the point bitmap and edge bitmap, that is, for the records appearing in these two types of lists, the bits at the corresponding positions on the corresponding point bitmap and edge bitmap are set to 1. The above steps complete the playback of the graph.

[0102] Figure 7 is a schematic flowchart showing the loading and playback of snapshots in this application. As Figure 7 shown, first scan all M snapshots and obtain the total number of points L in the M snapshots p , and save the number of edges corresponding to each snapshot into the array b e . Then, set all bits of the point bitmap corresponding to the length L p to 0. Next, set all bits of the edge bitmap corresponding to each snapshot to 0. Then, starting from the first snapshot, according to the points marked in the deletion point list in each snapshot, set the bits of the corresponding points in L p to 1. Then, according to the edges marked in the deletion edge list in each snapshot, set the bits of the corresponding edges in b e to 1. Apply the deletion point list and deletion edge list in each snapshot to the bitmap for each snapshot until the update of the point bitmap and edge bitmap is completed for all M snapshots. This process can be understood as completing the loading and playback of the graph

[0103] It should be noted that during the loading and playback process, a certain point may have corresponding edges in multiple snapshots. However, since these edge sets have been loaded and exist in memory, subsequent algorithms need to query the edge sets corresponding to the indexes of the corresponding points in these snapshots and check the status of the corresponding points and edges in the point bitmap and edge bitmap to determine whether they are valid points and valid edges.

[0104] For the user's graph analysis request, the graph data system can perform a graph query on a specific graph based on the playback result, and generate an analysis result of the specific graph based on the graph query result. The "graph query" process will be introduced in detail below in combination with different scenarios.

[0105] In a possible implementation manner, the graph data system may receive a first request message from the user. This first request message is used to request graph analysis on the target graph using the target algorithm. The graph data system can run a specific graph analysis algorithm based on these M snapshots, perform point queries and / or edge queries on the target graph according to the graph analysis algorithm, and perform operations defined by the algorithm on the query results to obtain the graph analysis result; the graph data system sends a first response message, and this first response message carries the analysis result of the target algorithm on the target graph, where the analysis result is generated based on point queries and / or edge queries.

[0106] Scenario 1

[0107] In a possible application scenario, the graph data system performs a point query on the target graph based on the M snapshots, including: the graph data system determines the set of existing points according to the Mth point set corresponding to the point-edge index table in the Mth snapshot among the M snapshots; determines the deleted points according to the point deletion list included in each of the M snapshots; and determines the set of valid points according to the set of existing points and the deleted points. The above scenario can also be understood as a scenario for querying valid points. Exemplarily, it is assumed that the user can instruct the graph database to query the valid points of the current target graph.

[0108] Exemplarily, in the above Scenario 1, the graph data system can perform a point query on the target graph according to the M snapshots, including: the graph data system can determine the point bitmap according to the point-edge index table in the Mth snapshot and the point deletion list in each snapshot, where the point bitmap is used to indicate the respective status of all points included in the M snapshots, and the status includes the deleted status and the undeleted status. Among them, the bit length of the point bitmap corresponds to the number of points in the Mth point set corresponding to the Mth snapshot, and the bit value corresponding to the deleted point in the point bitmap is different from the bit value corresponding to the undeleted point; determines the set of valid points according to the point bitmap.

[0109] Scenario 2

[0110] In another possible application scenario, the graph data system performs an edge query on the target graph according to the M snapshots, including: the graph data system determines the set of existing edges corresponding to each valid point according to the edge list in each of the M snapshots; determines the deleted edges corresponding to each valid point according to the edge deletion list in each of the M snapshots; and determines the set of valid edges corresponding to each valid point according to the set of existing edges corresponding to each valid point and the deleted edges corresponding to each valid point. This scenario can also be understood as a scenario for querying valid edges. Exemplarily, it is assumed that the user can instruct the graph database to query the valid edges of the current target graph.

[0111] Exemplarily, in the above Scenario 2, the graph data system performs an edge query on the target graph according to the M snapshots, including: determines the edge bitmap corresponding to each snapshot according to the edge list and the edge deletion list in each of the M snapshots, where the edge bitmap is used to indicate the status of the edges in each snapshot, and the status includes the deleted status and the undeleted status. Among them, the bit length of the edge bitmap corresponding to each snapshot corresponds to the number of edges in the edge list of the respective snapshot, and the bit value corresponding to the deleted edge in each edge bitmap is different from the bit value corresponding to the undeleted edge; determines the set of valid edges corresponding to each valid point according to the edge bitmap.

[0112] Scenario 3

[0113] In yet another possible application scenario, the graph data system can query whether a certain point specified by the user is a valid point. If it is a valid point, the graph data system will further query the valid edge corresponding to the valid point.

[0114] As described above in conjunction with Figures 4 to 7 the technical solution of the present application has been introduced in detail. Below, the technical solution of the present application will be described again in conjunction with Figure 2 and described as a whole. Taking Figure 2 as an example, the "graph data system" in the present application can be understood, for example, as the front-end system 210, the business system 220, the graph partitioning system 230, and the graph analysis system 240 in the graph.

[0115] Next, a graph analysis method 800 proposed by the present application will be introduced. Figure 8 FIG. shows a schematic flowchart of method 800. For example, this method can be executed by the graph data system, or can be jointly executed by each component module in the graph data system, which is not limited. Below, taking Figure 2 as an example to illustrate this method, the method 800 includes:

[0116] 810, the front-end system 210 receives a first request message from the user. The first request message is used to request to perform graph analysis on the target graph using the target algorithm.

[0117] Exemplarily, the user can indicate the target graph and the target algorithm.

[0118] 820, the business system 220 triggers the graph database or the file system to generate incremental records regularly or quantitatively.

[0119] In one possible implementation, the graph database or the file system can receive a second request message from a monitor module (not shown in the figure). The second request message is used to request the graph database or the file system to export incremental records. Exemplarily, the monitor module can trigger the graph database to export incremental records of the target graph regularly. Exemplarily, the monitor module can monitor the data volume of the graph database or the file system. When it exceeds a certain threshold, it will trigger the export of incremental records.

[0120] 830, the graph partitioning system 230 generates M snapshots according to the incremental records.

[0121] In this application, each snapshot includes a point-edge index table and an edge list. The point-edge index table in the i-th snapshot among the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot. Among them, the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot. The i-th point set includes the points in the (i - 1)-th point set and the newly added points in the incremental data corresponding to the i-th snapshot. The (i - 1)-th point set is the point set corresponding to the point-edge index table in the (i - 1)-th snapshot among the M snapshots.

[0122] In this application, i is an integer greater than 1, and M is an integer greater than 1.

[0123] In a possible implementation, the i-th snapshot further includes a point deletion list and / or an edge deletion list. Among them, the point deletion list is used to indicate the deleted points in the incremental data corresponding to the i-th snapshot; the edge deletion list is used to indicate the deleted edges in the incremental data corresponding to the i-th snapshot, and the j-th snapshot where the deleted edges are located, where the j-th snapshot is one of the M snapshots, and j is an integer less than i.

[0124] 840, the graph analysis system 240 performs point queries and / or edge queries on the target graph according to the M snapshots.

[0125] For example, the graph analysis system 240 can run a specific graph analysis algorithm, perform point queries and / or edge queries on the target graph according to the graph analysis algorithm, and perform operations defined by the algorithm on the query results to obtain a graph analysis result.

[0126] Specifically, the processes of point query and edge query can be referred to the above introduction and will not be elaborated here.

[0127] 850, the graph analysis system 240 sends a first response message to the user. The first response message carries the analysis result of the target algorithm on the target graph, where the analysis result is generated based on point queries and / or edge queries.

[0128] It should be noted that in this application, the representation method of snapshots is not limited to the CSR method, and other representation methods can also be used. As long as all the points in the corresponding point set of the previous snapshot are included in the point set corresponding to each snapshot proposed in this application, it falls within the protection scope required by this application. For example, the snapshot representation method can also be: using a map <i:e>, which is used to map the number i of the newly added points and the set e of associated edges.

[0129] In addition, in the illustrations of Figure 4 and Figure 5 , it is shown that the associated edges of the points are "out-edges". In fact, in the representation method of the snapshot, the associated edges of the points can also be represented as "in-edges" (that is, the associated edges of the points must be the edges with this point as the destination vertex), or the associated edges of the points include both in-edges and out-edges, and so on. This application does not limit the associated edges of the points to be "out-edges" and / or "in-edges".

[0130] Based on the solution proposed in this application, the set of points corresponding to each snapshot includes all the points in the set of points corresponding to the previous snapshot, which is convenient for users to operate and improves the data processing efficiency. And in the loading and playback stage, the deleted points and deleted edges can be efficiently applied, which is convenient for subsequent graph algorithms to quickly query whether a certain point and a certain edge are valid when obtaining the associated edge set of a certain point in each snapshot, and it is more convenient to perform graph analysis operations. Generally speaking, the solution provided in this application improves the user's business experience.

[0131] It can be understood that the term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and back associated objects.

[0132] Those skilled in the art should be able to realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.

[0133] The embodiments of this application can divide the functional modules of the computing device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other division methods in actual implementation. The following takes the division of each functional module corresponding to each function as an example for illustration.

[0134] In this application, the schematic diagrams of the various subsystems and sub-modules that the apparatus of the graph data system may include can be referred to the above Figure 2 , Figure 3 for understanding. Each of the subsystems and sub-modules can execute each step in the above method 800, and will not be repeated here.

[0135] Figure 9 FIG. is a schematic block diagram of another apparatus 900 of the graph data system provided by an embodiment of this application. As shown in the figure, the apparatus includes: at least one processor 920. The processor 920 is coupled to a memory and is configured to execute instructions stored in the memory to send signals and / or receive signals. Optionally, the apparatus 900 further includes a memory 930 for storing instructions. Optionally, the apparatus 900 further includes a transceiver 910, and the processor 920 controls the transceiver 910 to send signals and / or receive signals.

[0136] It should be understood that the above processor 920 and memory 930 can be integrated into a processing device, and the processor 920 is configured to execute program code stored in the memory 930 to implement the above functions. Specifically, in implementation, the memory 930 can also be integrated in the processor 920 or be independent of the processor 920.

[0137] It should also be understood that the transceiver 910 can include a transceiver (or, a receiver) and a transmitter (or, a transmitter). The transceiver can further include antennas, and the number of antennas can be one or more. The transceiver 910 can be a communication interface or an interface circuit.

[0138] As a solution, the apparatus is configured to implement the steps in the embodiment of the above method 800. For example, the processor 920 is configured to execute a computer program or instructions stored in the memory 930 to implement each step in the above method 800.

[0139] It should be understood that the specific processes of each transceiver and processor executing the above corresponding steps have been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.

[0140] In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in software form. The steps of the method disclosed in combination with the embodiments of this application can be directly embodied as being executed and completed by the hardware processor, or be executed and completed by a combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0141] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments may be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0142] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch-link DRAM (SLDRAM), and direct ram-bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0143] According to the method provided by the embodiments of the present application, the present application also provides a computer program product. Computer program code is stored on the computer program product. When the computer program code runs on a computer, the computer is caused to execute the steps in the embodiments of method 800.

[0144] According to the method provided by the embodiments of the present application, the present application also provides a computer-readable medium. The computer-readable medium stores program code. When the program code runs on a computer, the computer is caused to execute the steps in the embodiments of the above-mentioned method 800.

[0145] For the explanations and beneficial effects of the relevant content in any of the above-mentioned devices, reference can be made to the corresponding method embodiments provided above, and details are not described herein again.

[0146] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-density digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.

[0147] In each of the above device embodiments, the corresponding steps are executed by the corresponding module or unit. For example, the transceiver unit (transceiver) executes the receiving or sending steps in the method embodiments, and the other steps except for sending and receiving can be executed by the processing unit (processor). The functions of the specific units can refer to the corresponding method embodiments. Among them, the processor can be one or more.

[0148] As used in this specification, the terms "component", "module", "system", etc. are used to denote computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside in a process and / or an execution thread, and a component can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media having various data structures stored thereon. A component can communicate, for example, by signals according to one or more data packets (e.g., data from two components interacting with another component among a local system, a distributed system, and / or a network, such as via the Internet interacting with other systems by signals) through local and / or remote processes.

[0149] Those of ordinary skill in the art will appreciate that the elements and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0150] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0151] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0152] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0153] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0154] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0155] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.< / i:e>

Claims

1. A graph data system, characterized in that, including: The graph data system stores M snapshots, each of the M snapshots includes a point-edge index table and an edge list. The point-edge index table in the i-th snapshot of the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot. Wherein, the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot. The i-th point set includes the points in the (i - 1)-th point set and the newly added points in the incremental data corresponding to the i-th snapshot. The (i - 1)-th point set is the point set corresponding to the point-edge index table in the (i - 1)-th snapshot of the M snapshots. i is an integer greater than 1, and M is an integer greater than 1.

2. The graph data system according to claim 1, wherein The i-th snapshot further includes a deleted point list and / or a deleted edge list, wherein, the deleted point list is used to indicate the points deleted in the incremental data corresponding to the i-th snapshot; the deleted edge list is used to indicate the edges deleted in the incremental data corresponding to the i-th snapshot, and the j-th snapshot where the deleted edges are located, where the j-th snapshot is one of the M snapshots, and j is an integer less than i.

3. The graph data system according to claim 1, wherein including: The edge list in the first snapshot of the M snapshots is used to indicate the original edges in the original data corresponding to the first snapshot. The point-edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot. The first point set includes the original points in the original data.

4. The graph data system according to any one of claims 1 to 3, characterized in that, The graph data system further includes: Receiving a first request message from a user, the first request message is used to request graph analysis on a target graph using a target algorithm; Performing a point query and / or an edge query on the target graph according to the M snapshots; Sending a first response message, the first response message carries the analysis result of the target algorithm on the target graph, wherein the analysis result is generated based on the point query and / or the edge query.

5. The graph data system according to claim 4, wherein The performing a point query on the target graph according to the M snapshots includes: Determining a set of existing points according to the M-th point set corresponding to the point-edge index table in the M-th snapshot of the M snapshots; Determining the deleted points according to the deleted point list included in each of the M snapshots; Determining the set of valid points according to the set of existing points and the deleted points.

6. The figure data system according to claim 4 or 5, characterized in that, The performing a point query on the target graph according to the M snapshots includes: Determining a point bitmap according to the point-edge index table in the M-th snapshot and the deleted point list in each snapshot. The point bitmap is used to indicate the respective states of all the points included in the M snapshots. The states include a deleted state and an undeleted state. Wherein, the bit length of the point bitmap is the same as the number of points in the M-th point set corresponding to the M-th snapshot. The bit value corresponding to the deleted points in the point bitmap is different from the bit value corresponding to the undeleted points; Determining the set of valid points according to the point bitmap.

7. The graph data system according to claim 4, wherein Performing edge queries on the target graph according to the M snapshots includes: Determining, according to the edge lists in each of the M snapshots, a set of existing edges corresponding to each valid point; Determining, according to the edge deletion lists in each of the M snapshots, the deleted edges corresponding to each valid point; Determining, according to the set of existing edges corresponding to each valid point and the deleted edges corresponding to each valid point, a set of valid edges corresponding to each valid point.

8. The graph data system according to claim 4 or 7, characterized in that, Performing edge queries on the target graph according to the M snapshots includes: Determining, according to the edge lists and the edge deletion lists in each of the M snapshots, an edge bitmap corresponding to each snapshot, where the edge bitmap is used to indicate the status of the edges in each snapshot, and the status includes a deleted status and an undeleted status. Wherein, the bit length of the edge bitmap corresponding to each snapshot is the same as the number of edges in the edge list in the corresponding snapshot, and the bit value corresponding to the deleted edge in each edge bitmap is different from the bit value corresponding to the undeleted edge; Determining, according to the edge bitmap, a set of valid edges corresponding to each valid point.

9. The graph data system according to any one of claims 1 to 8, characterized in that The graph data system further includes: Receiving a second request message from a first module, where the second request message is used to request the graph data system to export incremental records, and the incremental records include at least one of the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges; Generating, according to the incremental records, a snapshot corresponding to the incremental records, where the snapshot includes at least one of the following: the point-edge index table, the edge list, the point deletion list, and the edge deletion list.

10. An apparatus for a graph data system, characterized in that, Including: A processor and a memory, where the processor is used to execute computer programs or instructions stored in the memory. Wherein, The memory stores M snapshots, and each of the M snapshots includes a point-edge index table and an edge list. The point-edge index table in the i-th snapshot among the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot. Wherein, the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot, the i-th point set includes the points in the (i - 1)-th point set and the newly added points in the incremental data corresponding to the i-th snapshot, the (i - 1)-th point set is the point set corresponding to the point-edge index table in the (i - 1)-th snapshot among the M snapshots, i is an integer greater than 1, and M is an integer greater than 1.

11. The device according to claim 10, characterized in that, The i-th snapshot further includes a point deletion list and / or an edge deletion list. Wherein, The point deletion list is used to indicate the deleted points in the incremental data corresponding to the i-th snapshot; The edge deletion list is used to indicate the deleted edges in the incremental data corresponding to the i-th snapshot and the j-th snapshot where the deleted edges are located, where the j-th snapshot is one of the M snapshots, and j is an integer less than i.

12. The device according to claim 10, characterized in that, Including: The edge list in the first snapshot among the M snapshots is used to indicate the original edges in the original data corresponding to the first snapshot, and the point-edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot. The first point set includes the original points in the original data.

13. The device according to any one of claims 10 to 12, characterized in that, The apparatus further includes a transceiver, wherein, the transceiver is configured to receive a first request message for requesting graph analysis on a target graph using a target algorithm; the processor is configured to perform point queries and / or edge queries on the target graph according to the M snapshots; send a first response message, where the first response message carries the analysis result of the target algorithm on the target graph, and the analysis result is generated based on the point queries and / or edge queries.

14. The device according to claim 13, characterized in that, The processor is configured to perform point queries on the target graph according to the M snapshots, including: the processor is configured to determine the set of existing points according to the Mth point set corresponding to the point-edge index table in the Mth snapshot among the M snapshots; the processor is configured to determine the deleted points according to the point deletion list included in each of the M snapshots; the processor is configured to determine the set of valid points according to the set of existing points and the deleted points.

15. The device according to claim 13 or 14, characterized in that, The processor is configured to perform point queries on the target graph according to the M snapshots, including: the processor is configured to determine a point bitmap according to the point-edge index table in the Mth snapshot and the point deletion list in each snapshot. The point bitmap is used to indicate the status of each of all the points included in the M snapshots, and the status includes a deleted status and an undeleted status. Wherein, the bit length of the point bitmap is the same as the number of points in the Mth point set corresponding to the Mth snapshot, and the bit value corresponding to the deleted points in the point bitmap is different from the bit value corresponding to the undeleted points; the processor is configured to determine the set of valid points according to the point bitmap.

16. The device according to claim 13, characterized in that, The processor is configured to perform edge queries on the target graph according to the M snapshots, including: the processor is configured to determine the set of existing edges corresponding to each valid point according to the edge list in each of the M snapshots; the processor is configured to determine the deleted edges corresponding to each valid point according to the edge deletion list in each of the M snapshots; the processor is configured to determine the set of valid edges corresponding to each valid point according to the set of existing edges corresponding to each valid point and the deleted edges corresponding to each valid point.

17. The device according to claim 13 or 16, characterized in that, The processor is configured to perform edge queries on the target graph according to the M snapshots, including: The processor is configured to determine, according to the edge list and the edge deletion list in each of the M snapshots, an edge bitmap corresponding to each snapshot, where the edge bitmap is used to indicate the status of edges in each snapshot, and the status includes a deleted status and an undeleted status. Wherein, the bit length of the edge bitmap corresponding to each snapshot is the same as the number of edges in the edge list in the corresponding snapshot, and the bit value corresponding to the deleted edge in each edge bitmap is different from the bit value corresponding to the undeleted edge; The processor is configured to determine, according to the edge bitmap, a set of valid edges corresponding to each valid point.

18. The device according to any one of claims 10 to 17, characterized in that, The apparatus further includes a transceiver, wherein, The transceiver is configured to receive a second request message for requesting the graph data system to export incremental records, where the incremental records include at least one of the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges; The processor is configured to generate, according to the incremental records, a snapshot corresponding to the incremental records, where the snapshot includes at least one of the following: the point-edge index table, the edge list, the point deletion list, and the edge deletion list.

19. A computer-readable storage medium, characterized in that, Instructions are stored on the computer-readable storage medium, and when the instructions are run on a computer, the computer is caused to perform the actions performed by the graph data system according to any one of claims 1 to 9.

20. A computer program product, characterized in that, Includes instructions that, when run on a computer, cause the computer to perform the actions performed by the graph data system according to any one of claims 1 to 9.