Graph data system and apparatus

By storing multiple snapshots in the graph data system, each snapshot contains a point edge index table and an edge list, the problem of inefficient loading and playback of incremental graphs is solved, and efficient data processing and convenient graph analysis are achieved.

WO2025130796A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139450
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-21
Filing Date
2024-12-16
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The prior art is inefficient in the loading and playback of incremental graphs, and requires frequent replacement of pages and recovery of incremental data, resulting in inconvenience to users.

Method used

By storing M snapshots in the graph data system, each snapshot includes a point edge index table and an edge list, ensuring that the point set of each snapshot contains all points and new points of the previous snapshot, and supports efficient incremental data loading and playback.

Benefits of technology

It improves data processing efficiency, simplifies graph analysis services, reduces the complexity of user operations, and realizes convenient graph analysis and data query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139450_26062025_PF_FP_ABST
    Figure CN2024139450_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application is a graph data system. On the basis of the solution provided in the present application, the graph data system stores a plurality of snapshots, and each point set corresponding to a snapshot comprises all points in a point set corresponding to the former snapshot, thereby facilitating operations of a user, and improving the data processing efficiency. In addition, in a loading and playback stage, deleted points and deleted edges can be efficiently applied, such that when incident edge sets of certain points in snapshots are acquired by using a graph algorithm, whether a certain point or a certain edge is valid can be quickly queried, and thus a graph analysis service is performed conveniently. Generally speaking, the solution provided in the present application improves the service experience of users.
Need to check novelty before this filing date? Find Prior Art

Description

Graph data system and device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 21, 2023, with application number 202311766204.1, and priority to the Chinese patent application entitled “A Graph Data System and Device”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of graph analysis, and more specifically, to a graph data system and apparatus. Background Art

[0003] In recent years, global big data has entered a period of accelerated development, with data volumes growing exponentially. The data generated by the relationships between different individuals in big data is presented in the form of "graphs." "Graphs" here do not refer to literal pictures or images, but rather refer to the mathematical concept of "graph theory," where they can be understood as data structures composed of "vertices" and "edges." Vertices are equivalent to nodes, and the relationships between vertices are called "edges." "Incremental data" can be understood as the newly added "vertices" and "edges" to the graph data structure.

[0004] The stored incremental graph files are loaded into memory and, through certain methods, the incremental changes in the graph data are reflected in the changes in the organization of the graph data in memory, making it recognizable by subsequent graph analysis. This process can be understood as "incremental graph loading and playback." Existing methods require replacing pages during incremental graph loading and playback, and require multiple continuation markers between multiple old pages to restore the incremental data. This is inefficient and inconvenient for users. Therefore, a solution is needed that can more conveniently perform graph analysis and improve data processing efficiency. Summary of the Invention

[0005] The present application provides a graph data system, in which user operations can be facilitated and data processing efficiency can be improved.

[0006] In a first aspect, a graph data system is provided. The graph data system may include multiple units, modules, and devices that implement graph data processing functions. Alternatively, the graph data system may also include components (e.g., chips or circuits) that implement graph data processing functions, without limitation.

[0007] M snapshots are stored in the graph data system, each of the M snapshots includes a point-edge index table and an edge list, the point-edge index table in the i-th snapshot among the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot, wherein the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot, the i-th point set includes the points in the i-1-th point set and the newly added points in the incremental data corresponding to the i-th snapshot, the i-1-th point set is the point set corresponding to the point-edge index table in the i-1-th snapshot among the M snapshots, i is an integer greater than 1, and M is an integer greater than 1.

[0008] It can also be understood that, in this application, the points in the point set in the point edge index table in the i-th snapshot actually include the points in the (i-1)-th snapshot and the points newly added in the i-th snapshot. In other words, in this application, the points corresponding to the i-th snapshot actually include the points in all previous snapshots and the points newly added in the i-th snapshot.

[0009] For example, the graph data system can be a graph database or a file system. This application does not limit the name of the graph data system. As long as the device or module can realize the graph data processing function, it falls within the scope of protection of this application.

[0010] Exemplarily, the snapshot representation method in the present application may be a compressed sparse row (CSR) representation method.

[0011] In the solution of this application, the point set corresponding to each snapshot includes all points in the point set corresponding to the previous snapshot, which facilitates user operations and improves data processing efficiency. Furthermore, during the loading and playback phase, deleted points and edges can be efficiently applied, making it easier for subsequent graph algorithms to quickly query the validity of a point or edge when obtaining the edge set associated with a point in each snapshot, making graph analysis more convenient. Overall, the solution provided by this application improves the user experience.

[0012] In combination with the first aspect, in one implementation, the i-th snapshot also includes a point deletion list and / or an edge deletion list, wherein the point deletion list is used to indicate the points deleted in the incremental data corresponding to the i-th snapshot; the edge deletion list is used to indicate the edges deleted in the incremental data corresponding to the i-th snapshot, and the j-th snapshot where the deleted edges are located, wherein the j-th snapshot is one of the M snapshots, and j is an integer less than i.

[0013] In this application, a "snapshot" is used to indicate updated points or edges in a graph. Therefore, a snapshot may also include a deleted point list and / or a deleted edge list, which indicates at least one of deleted points, deleted edges, and newly added edges.

[0014] In combination with the first aspect, in one implementation, the edge list in the first snapshot of M snapshots is used to indicate the original edges in the original data corresponding to the first snapshot, and the point edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot, and the first point set includes the original points in the original data.

[0015] It should be noted that the first snapshot in this application can be understood as the original snapshot. That is, it can be understood as the original snapshot corresponding to the original data. The relationship between the points and edges indicated in the snapshot can be understood as the relationship between the points and edges in the original graph. The edge list in the first snapshot is used to indicate the original edges in the original data corresponding to the first snapshot, and the point-edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot. The first point set includes the original points in the original data. At this time, there is no deleted point list and deleted edge list in the first snapshot. Alternatively, it can also be understood that in the first snapshot, the contents of the deleted point list and deleted edge list are blank.

[0016] In combination with the first aspect, in one implementation, the graph data system is further used to receive a second request message from the first module, and the second request message is used to request the graph data system to export incremental records, and the incremental records include at least one of the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges; the graph data system is used to generate a snapshot corresponding to the incremental records based on the incremental records, and the snapshot includes at least one of the following items: a point-edge index table, an edge list, a deleted point list, and a deleted edge list.

[0017] It should be noted that in this application, the "incremental record" records a list of updated points and / or edges. For example, the incremental record includes the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges. The incremental snapshot corresponding to the incremental record can be generated by the data in the incremental record, and the incremental snapshot can be used to describe the changes in points and / or edges in the incremental record. Exemplarily, the incremental record is in the form of a list. For example, a list of newly added points, a list of newly added edges, a list of deleted points, and a list of deleted edges. Exemplarily, the snapshot includes: a point-edge index table, an edge list, a deleted point list, and a deleted edge list.

[0018] In one possible implementation, the graph data system may receive a second request message from the first module, where the second request message is used to request the graph data system to export incremental records, and the graph database may subsequently generate a snapshot corresponding to the incremental records based on the incremental records.

[0019] In one possible implementation, the first module can trigger the graph data system in a timed or quantitative manner, causing the graph data system to generate "incremental records." For example, the first module can be an "incremental trigger module"; in another example, the first module can be an incremental monitor module. The specific name of the first module is not limited in this application.

[0020] In combination with the first aspect, in one implementation, the graph data system further includes: receiving a first request message from a user, the first request message being used to request graph analysis of a target graph using a target algorithm; performing point queries and / or edge queries on the target graph based on M snapshots; and sending a first response message, the first response message carrying the analysis results of the target graph by the target algorithm, wherein the analysis results are generated based on the point queries and / or edge queries.

[0021] For example, the user can indicate a target graph and a target algorithm.

[0022] It should be noted that, generally, in the process of graph analysis of the target algorithm on the target graph, the target algorithm needs to perform point query and / or edge query on the replayed graph multiple times, and finally obtain analysis results based on the point query and / or edge query results.

[0023] Exemplarily, the graph data system is used to receive a first request message from a user, where the first request message is used to request graph analysis of a target graph using a target algorithm; the graph data system is used to perform point queries and / or edge queries on the target graph based on M snapshots; and send a first response message, where the first response message carries the analysis results of the target graph by the target algorithm, where the analysis results are generated based on the point queries and / or edge queries.

[0024] In this application, before performing graph analysis, it is necessary to complete the loading and playback of the graph. When loading and replaying the graph, it is necessary to maintain the point map and edge map. For example, the following implementation method can be used when specifically maintaining the point map and edge map. In one possible implementation method, the graph data system can read in and scan the snapshots to obtain the number of snapshots 2, the total number of points Lp is 7, and the number of edges corresponding to each snapshot is [7,3]. An array bp of a point map with a length of Lp=7 and an array be of an edge map are successively established. Among them, the length of be[1] is 7, the length of be[2] is 3, and all bits of these bitmaps are set to 0. The graph data system can scan the snapshots again and apply the information on the point deletion list and edge deletion list of each snapshot to the point map and edge map, that is, the bits of the corresponding positions of the records in these two lists on the corresponding point map and edge map are set to 1. The above steps complete the playback of the graph.

[0025] In combination with the first aspect, in one implementation, the graph data system is used to perform a point query on the target graph based on the M snapshots, including: the graph data system is used to determine the set of existing points based on the M-th point set corresponding to the point edge index table in the M-th snapshot among the M snapshots; the graph data system is used to determine the deleted points based on the deleted point list included in each of the M snapshots; the graph data system is used to determine the set of valid points based on the set of existing points and the deleted points.

[0026] The above implementation describes how to query the set of valid points during a point query. Specifically, since each snapshot includes a point-edge index table, this table can be used to retrieve all existing points in the current snapshot. Then, based on the deleted point list in that snapshot, the points deleted from the current snapshot can be determined, thereby determining the set of valid points in the current snapshot.

[0027] In combination with the first aspect, in one implementation, the graph data system is used to perform point queries on the target graph based on the M snapshots, including: the graph data system is used to determine a point bitmap based on the point edge index table in the M-th snapshot and the deleted point list in each snapshot, the point bitmap is used to indicate the corresponding status of all points included in the M snapshots, the status including a deleted state and a non-deleted state, wherein the bit length corresponding to the point bitmap is the same as the number of points in the M-th point set corresponding to the M-th snapshot, and the bit values ​​corresponding to the deleted points in the point bitmap are different from the bit values ​​corresponding to the non-deleted points; the graph data system is used to determine a set of valid points based on the point bitmap.

[0028] In combination with the first aspect, in one implementation, the graph data system is used to perform edge queries on the target graph based on M snapshots, including: the graph data system is used to determine the set of existing edges corresponding to each valid point based on the edge list in each of the M snapshots; the graph data system is used to determine the deleted edges corresponding to each valid point based on the deleted edge list in each of the M snapshots; the graph data system is used to determine the set of valid edges corresponding to each valid point based on the set of existing edges corresponding to each valid point and the deleted edges corresponding to each valid point.

[0029] In combination with the first aspect, in one implementation, a graph data system is used to perform edge queries on a target graph based on M snapshots, including: the graph data system is used to determine an edge bitmap corresponding to each snapshot based on an edge list and a deleted edge list in each of the M snapshots, the edge bitmap being used to indicate the state of the edges in each snapshot, the states including a deleted state and an undeleted state, wherein the bit length of the edge bitmap corresponding to each snapshot is the same as the number of edges in the edge list in the corresponding snapshot, and the bit values ​​corresponding to the deleted edges in each edge bitmap are different from the bit values ​​corresponding to the undeleted edges; the graph data system is used to determine, based on the edge bitmap, the set of valid edges corresponding to each valid point.

[0030] The above describes the scenarios of querying the set of valid points and querying the set of valid edges corresponding to each valid point. In another possible application scenario, the graph data system can query whether a user-specified point is a valid point. If it is a valid point, the graph data system will further query the valid edges corresponding to the valid point.

[0031] In a second aspect, a graph data system apparatus is provided, the apparatus being configured to execute the method of any possible implementation of the first aspect. Specifically, the apparatus may include units and / or modules, such as a transceiver unit and / or a processing unit, configured to execute the method of any possible implementation of the first aspect.

[0032] In one implementation, the apparatus is a computing device, the communication unit may be a transceiver or an input / output interface, and the processing unit may be at least one processor. Alternatively, the transceiver may be a transceiver circuit. Alternatively, the input / output interface may be an input / output circuit.

[0033] In another implementation, the apparatus is a chip, chip system, or circuit of a computing device for implementing graph analysis functionality. When the apparatus is a chip, chip system, or circuit for implementing graph analysis functionality, the communication unit may be an input / output interface, interface circuit, output circuit, input circuit, pin, or related circuit on the chip, chip system, or circuit; and the processing unit may be at least one processor, processing circuit, or logic circuit.

[0034] In a third aspect, a graph data system apparatus is provided, the apparatus comprising: at least one processor configured to execute a computer program or instructions stored in a memory to perform the method of any possible implementation of the first aspect. Optionally, the apparatus further comprises a memory configured to store the computer program or instructions. Optionally, the apparatus further comprises a communication interface, through which the processor reads the computer program or instructions stored in the memory.

[0035] In a fourth aspect, the present application provides a processor, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is configured to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method in any possible implementation of the first aspect.

[0036] In a specific implementation, the processor may be one or more chips, the input circuit may be an input port, the output circuit may be an output port, and the processing circuit may be a transistor, a gate circuit, a trigger, or various logic circuits. The input signal received by the input circuit may be, for example, but not limited to, received and input by a transceiver, and the signal output by the output circuit may be, for example, but not limited to, output to and transmitted by a transmitter. The input circuit and the output circuit may be the same circuit, which functions as an input circuit and an output circuit at different times. The embodiments of the present application do not limit the specific implementation of the processor and various circuits.

[0037] For the operations such as sending and acquiring / receiving involved in the processor, unless otherwise specified, or if they do not conflict with their actual functions or internal logic in the relevant descriptions, they can be understood as processor output, reception, input and other operations, and can also be understood as sending and receiving operations performed by the radio frequency circuit and antenna. This application does not limit this.

[0038] In a fifth aspect, a processing device is provided, comprising a processor and a memory. The processor is configured to read instructions stored in the memory and receive signals via a transceiver and transmit signals via a transmitter to execute the method of any possible implementation of the first aspect.

[0039] Optionally, there are one or more processors and one or more memories.

[0040] Optionally, the memory may be integrated with the processor, or the memory may be provided separately from the processor.

[0041] In the specific implementation process, the memory can be a non-transitory memory, such as a read-only memory (ROM), which can be integrated with the processor on the same chip or can be set on different chips. The embodiments of the present application do not limit the type of memory and the setting method of the memory and the processor.

[0042] It should be understood that related data interaction processes, such as sending indication information, can be the process of outputting indication information from the processor, and receiving capability information can be the process of receiving input capability information from the processor. Specifically, data output by the processor can be output to the transmitter, and input data received by the processor can be received from the transceiver. The transmitter and transceiver can be collectively referred to as a transceiver.

[0043] The processing device in the fifth aspect may be one or more chips. The processor in the processing device may be implemented in hardware or software. When implemented in hardware, the processor may be a logic circuit, an integrated circuit, or the like; when implemented in software, the processor may be a general-purpose processor implemented by reading software code stored in a memory, which may be integrated into the processor or located independently of the processor.

[0044] In a sixth aspect, a computer-readable storage medium is provided, which stores a program code for execution by a device, wherein the program code includes a method for executing any possible implementation of the first aspect.

[0045] In a seventh aspect, a computer program product comprising instructions is provided, which, when run on a computer, enables the computer to execute the method in any possible implementation of the first aspect.

[0046] In an eighth aspect, a chip system is provided, comprising a processor for calling and running a computer program from a memory, so that a device equipped with the chip system executes the methods in each implementation of the above-mentioned first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] FIG1 is a schematic flowchart of full export and incremental update in graph analysis provided in this application.

[0048] FIG2 is a schematic block diagram of a system architecture applicable to the present application.

[0049] FIG3 is another schematic block diagram of a system architecture applicable to the present application.

[0050] FIG4 is a schematic diagram of a snapshot #1 at time T1 provided by this application.

[0051] FIG5 is a schematic diagram of snapshot #2 at time T2 provided by this application.

[0052] FIG6 is a schematic diagram of a point map and an edge map provided in this application.

[0053] FIG7 is a schematic flowchart of loading and playback provided by this application.

[0054] FIG8 is a schematic flowchart of a graph analysis method 800 provided in this application.

[0055] FIG9 is a schematic block diagram of a device 900 of a graph data system proposed in this application. DETAILED DESCRIPTION

[0056] The technical solution in this application will be described below with reference to the accompanying drawings.

[0057] In order to facilitate understanding of the technical solution of this application, the following first briefly introduces some professional terms in this application.

[0058] 1. Picture

[0059] In recent years, global big data has entered a period of accelerated development, with data volumes growing exponentially. The data generated by the relationships between different individuals in big data is presented in the form of "graphs." "Graph" here doesn't refer to a literal picture or image, but rather refers to the mathematical concept of "graph theory," where it can be understood as a data structure composed of "vertices" and "edges." Vertices are equivalent to nodes, and the relationships between vertices are called "edges." "Incremental data" can be understood as newly added "vertices" and "edges" to the graph data structure. For example, three people sitting in an office are three vertex points. The relationships between these three people are called edges, such as those between colleagues, junior colleagues, and project partners.

[0060] 2. Graph Analysis

[0061] Graph analytics uses graph-based methods to analyze connected data. Graph analytics can involve querying graph data, applying basic statistics, visually exploring graphs, displaying graphs, or preprocessing graph information and incorporating it into machine learning tasks. Graph queries are typically used for analyzing localized data, while graph computations typically involve analyzing the entire graph and performing iterative analysis. Graph analytics focuses on analyzing the strength and direction of relationships between entities in graph data to uncover insights and aid decision-making. Graph analytics evaluates or extracts information from the input graph data by running algorithms that recognize the structure of the graph.

[0062] Graph analysis relies on a variety of graph analysis algorithms to implement many different types of analysis, such as association analysis algorithms, path analysis algorithms, classification algorithms, clustering algorithms, and time series analysis algorithms.

[0063] 3. Graph Analysis Application Scenarios

[0064] (1) Social network analysis scenario: Social networks are a very common type of graph data that represents the social relationships between various individuals or organizations. Graph data can present complex social network relationships, making it easier for users to conduct further analysis. For example, in a typical social network, there are often questions like "who knows whom, who went to what school, who lives where", and graph analysis can be used to manage social relationships, implement friend recommendations, and so on.

[0065] (2) E-shopping application scenarios: E-shopping is a core business on the Internet. In this scenario, nodes are divided into two categories: users and products, and the relationships that exist include browsing, collecting, and purchasing. Multiple relationships can exist between users and products, such as both a collection relationship and a purchase relationship. This complex data scenario can be easily described using attribute graphs. E-shopping has spawned a well-known technical application - the "recommendation system." The interactive relationship between users and products reflects users' shopping preferences.

[0066] (3) Transportation network application scenarios: Transportation networks have various forms. For example, in a subway network, each station is considered a node, and the connectivity between stations is considered an edge. In transportation networks, we usually focus on path planning related problems, such as the shortest path problem. Another example is that we use traffic flow as an attribute of nodes in the network to predict future changes in traffic flow. A typical application scenario is map navigation.

[0067] 4. Offline graph analysis: By exporting the graph data stored in the graph database, graph analysis can be performed in an offline form.

[0068] 5. Graph sharding: The process of dividing large-scale graph data into several shards, so that multiple computer processes can run the same graph algorithm on these shards simultaneously.

[0069] 6. Incremental Graph: Based on the existing graph segmentation, incremental data exported from the database is segmented and appended to the original graph segmentation results to create an incremental data structure. The so-called "incremental graph" can be understood as real-time data. Incremental graph analysis only requires recalculating the changed parts of the data. Then, using some algorithms, the incremental results are integrated with the original graph calculation results to efficiently obtain new calculation results.

[0070] 7. Incremental Graph Load and Replay: Load the incremental graph file into memory and, through a specific method, ensure that the incremental changes in the graph data are reflected in the changes in the in-memory organization, making them recognizable in subsequent graph analysis. This means that the graph data represented in memory after replay is consistent with the graph data in the database.

[0071] Figure 1 shows a schematic flowchart of the full export and incremental update process in graph analysis. A full export requires exporting the full data ("full" can be understood as historical data) from the graph database, performing distributed graph partitioning, loading the graph, and restoring the full graph in memory before it can be used as algorithm input. In scenarios involving data updates, particularly for large-scale graphs with tens of billions of nodes and edges, the time required to re-export the full graph and perform graph partitioning can be significant. The incremental update mechanism used for offline graph analysis maximizes the preservation of the topology and attribute data of the completed graph partitions, avoiding the need to re-partition the entire graph. However, in incremental update scenarios, the business system can export incremental data to generate incremental snapshot #1, which can be used to describe the incremental data. The original graph data, incremental snapshot #0, is then determined through incremental graph storage. Incremental snapshot #0 and incremental snapshot #1 are loaded, and the incremental graph load and playback are performed in memory to restore the graph. Finally, graph analysis is performed on the restored graph using the graph analysis algorithm.

[0072] Normally, incremental graph updates need to take into account graph storage, graph loading and playback, and the efficiency of graph algorithms for accessing the graph after playback. The existing method requires replacing pages during incremental graph loading and playback, and requires multiple iterations of continued mark recovery between multiple old pages to recover incremental data, which is inefficient and inconvenient for user operation. In view of this, the present application provides a graph analysis method, a simple and efficient incremental graph loading and playback method, which avoids complex jump recovery, facilitates user operation, and improves data processing efficiency.

[0073] Figure 2 is a schematic block diagram of a system architecture applicable to the present application. As shown in Figure 2, the system architecture includes a front-end system 210, a business system 220, a graph segmentation system 230, and a graph analysis system 240. In a possible scenario, a user initiates an incremental graph analysis request for a certain graph on the end side. The server of the front-end system 210 receives the request and drives the business system 220 to generate an incremental record corresponding to the graph, and imports the incremental record into the graph segmentation system 230 for incremental update, thereby generating a new incremental snapshot corresponding to the incremental record. The new incremental snapshot will be stored in the incremental graph storage device. In addition, the incremental graph storage device also stores the previously existing snapshot set. The graph analysis system 240 loads and replays the graph based on the existing snapshot set and the new incremental snapshot, and the graph analysis algorithm starts the graph algorithm based on the graph data in the replayed memory, outputs the graph analysis results, and feeds back the graph analysis results to the user.

[0074] It should be noted that in this application, the "incremental record" records a list of updated points and / or edges. For example, the incremental record includes the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges. The incremental snapshot corresponding to the incremental record can be generated by the data in the incremental record, and the incremental snapshot can be used to describe the changes in points and / or edges in the incremental record. Exemplarily, the incremental record is in the form of a list. For example, a list of newly added points, a list of newly added edges, a list of deleted points, and a list of deleted edges. Exemplarily, the snapshot includes: a point-edge index table, an edge list, a deleted point list, and a deleted edge list.

[0075] In one possible implementation, the graph data system may receive a second request message from the first module, where the second request message is used to request the graph data system to export incremental records, and the graph database may subsequently generate a snapshot corresponding to the incremental records based on the incremental records.

[0076] In one possible implementation, the first module can trigger the graph data system in a timed or quantitative manner, causing the graph data system to generate "incremental records." For example, the first module can be an "incremental trigger module" (not shown in Figure 2); for another example, the first module can be an incremental monitor module. The specific name of the first module is not limited in this application.

[0077] Figure 3 is a schematic block diagram of an incremental update system shown in this application. Figure 3 can be understood as a more specific description of the process in Figure 2. As shown in Figure 3, for a specific graph, the graph database (or file system) 310 first exports the incremental records corresponding to the graph, the graph segmentation module 320 performs a graph segmentation operation on it, and generates an original snapshot (for example, it can be a memory object or a file combination), and the incremental graph storage module 340 saves the original snapshot. Exemplarily, the original snapshot can be represented by the structure of the incremental graph. Subsequently, if the graph database has a change in the graph structure, or point data, or edge data, at this time, the graph database 310 can export the incremental records corresponding to these changes, and the incremental update module 330 performs graph segmentation of the incremental data, and stores the generated new incremental snapshot in the incremental graph storage module 340. The graph loading module 350 can load and replay the existing snapshot set (including incremental snapshots and original snapshots) provided by the incremental graph storage, and use the playback results as input to the graph algorithm. The graph analysis algorithm module 360 ​​performs graph analysis.

[0078] It should be noted that the modules in Figures 2 and 3 of this application may be multiple hosts (e.g., a computer cluster) that perform corresponding functions, multiple processes in a computer, or multiple modules in multiple processes, etc., without limitation. Any device, apparatus, or chip that can implement the above functions falls within the scope of protection of this application.

[0079] The following is an introduction to the representation method of the incremental graph in this application. For example, the representation method of the incremental graph is introduced by taking the compressed sparse row (CSR) representation method as an example. It should be noted that in this application, the representation method of the incremental graph is not limited to the description method of the CSR format and can also be represented by other description methods. It is not limited to this. The description method of the CSR format here is only an example.

[0080] Figure 4 shows a snapshot at time T1, denoted as snapshot #1. As shown in Figure 4(a), snapshot #1 includes a vertex-edge index table and an edge list. Using the vertex-edge index table and edge list of snapshot #1, we can derive the vertex-edge relationship described by snapshot #1 as the graph shown in Figure 4(b). Figure 4(a) shows points 0 through 5, which can be understood as representing six points in this snapshot: "0," "1," "2," "3," "4," and "5." The numbers in the edge list have the same meaning as these six points. For example, "1" in the edge list refers to point 1, and "5" in the edge list refers to point 5. The "vertex-edge index table" in the CSR format describes the edge information corresponding to a particular vertex (also understood as an "outgoing edge," meaning that each edge corresponding to a particular vertex starts at that vertex and ends at another vertex). The "edge list" describes the information of another vertex on an edge corresponding to a particular vertex (also understood in the CSR format as the destination point of an edge corresponding to that vertex). For example, from the "point-edge index table", the value corresponding to point i is n, and the value corresponding to point (i-1) is m. Then the number of other points on the edge corresponding to point i is (m~n-1) (that is, the number of edges corresponding to point i is also (m~n-1)). Among them, the first point of the edge corresponding to point i is the value corresponding to point m in the edge list, and the last point of the edge corresponding to point i is the value corresponding to point (n-1) in the edge list.

[0081] The above CSR representation method can also be understood as follows: the "point-edge index table" in snapshot #1 is used to indicate the index relationship between the points in the first point set and the edges in the "edge list" in snapshot #1, wherein the edge list in snapshot #1 is used to indicate the original edges in the original data corresponding to the first snapshot. The "point-edge index relationship" in this application actually indicates the relationship between the current points, or it can be understood as which edges each of the current points has and which points are the destination points of these edges. The following is a detailed introduction to the CSR method using Figure 4 as an example.

[0082] For example, as shown in (a) of Figure 4, in the "Vertex-Edge Index Table," the value corresponding to point 0 is "2," and the value corresponding to point 1 is "4." This means that point 1 has two edges (2 and (4-1)). The destinations of these two edges should look up the values ​​corresponding to points 2 and 3 in the edge list. In the edge list, the value corresponding to point 2 is "4," so one edge of point 1 can be determined to be from point 1 to point 4. In the edge list, the value corresponding to point 3 is 3, so another edge of point 1 can be determined to be from point 1 to point 3.

[0083] For example, as shown in (a) of Figure 4, in the "Vertex-Edge Index Table," the value corresponding to point 2 is "5," and the value corresponding to point 1 is "4." This means that there is an edge corresponding to point 2 (i.e., 4). The destination of this edge should look for the value corresponding to point 4 in the edge list. In the edge list, the value corresponding to point "4" is "4," so it can be determined that the edge for point 2 is from point 2 to point 4.

[0084] For example, as shown in (a) of Figure 4, in the "Point-Edge Index Table," the value corresponding to point 3 is "7," and the value corresponding to point 2 is "5." This means that there are two edges corresponding to point 3 (i.e., 5 and (7-1)). The destinations of these two edges should look up the values ​​corresponding to points 5 and 6 (not shown) in the edge list. In the edge list, the value corresponding to point "5" is "4," so one edge of point 3 can be determined to be from point 3 to point 4. In the edge list, the value corresponding to point 6 (not shown) is 5, so another edge of point 3 can be determined to be from point 3 to point 5.

[0085] To more clearly illustrate the vertex and edge relationships described in Figure 4 (a), Figure 4 (b) is drawn based on Figure 4 (a). As previously mentioned, according to the CSR format, the number of edges corresponding to a point refers to the number of edges starting from that point. Based on Figure 4 (a), we can see that point 0 has two edges corresponding to it, from point 0 to point 1 and from point 0 to point 2; point 1 has two edges corresponding to it, from point 1 to point 4 and from point 1 to point 3; point 2 has one edge corresponding to it, from point 2 to point 4; point 3 has two edges corresponding to it, from point 3 to point 4 and from point 3 to point 5; points 4 and 5 have no corresponding edges.

[0086] In addition, it can be seen that the length of the edge list in snapshot #1 actually corresponds to the number of edges in the graph. In Figure 4 (a), the length of the edge list is 7, which corresponds to a total of 7 edges in Figure 4 (b).

[0087] The above mainly introduces a table lookup method in CSR format. The following continues to use CSR format as an example to introduce the solution of this application.

[0088] The present application proposes: M snapshots are stored in a graph data system, and each of the M snapshots includes a point-edge index table and an edge list. The point-edge index table in the i-th snapshot among the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot, wherein the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot, the i-th point set includes the points in the i-1-th point set and the newly added points in the incremental data corresponding to the i-th snapshot, the i-1-th point set is the point set corresponding to the point-edge index table in the i-1-th snapshot among the M snapshots, i is an integer greater than 1, and M is an integer greater than 1.

[0089] The "graph data system" in this application can be, for example, a graph database, or a file system, etc. It should be noted that the "graph data system" in this application can also refer to other database systems. Any device, apparatus, or module that can use the method provided in this application is within the scope of protection of this application.

[0090] It can also be understood that, in this application, the points in the point set in the point edge index table in the i-th snapshot actually include the points in the (i-1)-th snapshot and the points newly added in the i-th snapshot. In other words, in this application, the points corresponding to the i-th snapshot actually include the points in all previous snapshots and the points newly added in the i-th snapshot.

[0091] In one possible scenario, the i-th snapshot may also include a point deletion list and / or an edge deletion list, wherein the point deletion list is used to indicate the points deleted in the incremental data corresponding to the i-th snapshot; the edge deletion list is used to indicate the edges deleted in the incremental data corresponding to the i-th snapshot, and the j-th snapshot where the deleted edge is located, wherein the j-th snapshot is one of the M snapshots, and j is an integer less than i.

[0092] It should be noted that the first snapshot in this application can be understood as the original snapshot. That is, it can be understood as the original snapshot corresponding to the original data. The relationship between the points and edges indicated in the snapshot can be understood as the relationship between the points and edges in the original graph. The edge list in the first snapshot is used to indicate the original edges in the original data corresponding to the first snapshot, and the point-edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot. The first point set includes the original points in the original data. At this time, there is no deleted point list and deleted edge list in the first snapshot. Alternatively, it can also be understood that in the first snapshot, the contents of the deleted point list and deleted edge list are blank.

[0093] Figure 5 below is a possible schematic diagram of the solution proposed in this application. Figure 5 shows a snapshot at time T2, assuming that time T2 is a snapshot later than time T1, and the data in the graph at time T1 is updated at time T2. For example, the snapshot at time T2 is recorded as snapshot #2. As shown in (a) in Figure 5, snapshot #2 includes a point-edge index table, an edge list (which can also be understood as a newly added edge list), a deleted point list, and a deleted edge list, which respectively describe the newly added points, newly added edges, deleted points, and deleted edges at time T2.

[0094] It can be seen from (a) in Figure 5 that a new point has been added, assuming that the point is recorded as point "6". In this application, for example, the indexes corresponding to the new points are sorted in sequence with the indexes corresponding to the previous points. For example, the corresponding indexes of the original points and the indexes corresponding to the new points can be arranged in order from small to large. In addition, Figure 5 also includes a new edge list, from which it can be found that three new edges were added at time T1. The destination points of the three edges are point 2, point 6, and point 5, respectively. According to the table lookup method introduced in (a) in Figure 4, it can be seen that in snapshot #2, in the "point edge index table", the value corresponding to point 1 is "2", and the value corresponding to point 0 is "0", which means that two new edges have been added to point 1 (i.e.: 0, (2-1)). The destination points of these two edges should look up the values ​​corresponding to point 0 and point 1 in the edge list. In the edge list, the value corresponding to point 0 is "2", so we can determine that the new edge added to point 1 is from point 1 to point 2; in the newly added edge list, the value corresponding to point 1 is 6, so we can determine that the other new edge added to point 1 is from point 1 to point 6. In the vertex-edge index table, the value corresponding to point "6" is "3", and the value corresponding to point 5 is "2", so we can determine that a new edge (i.e., 2) has been added to point 6. In the edge list, the value corresponding to point 2 is "5", so we can determine that the new edge added to point 6 is from point 1 to point 5.

[0095] From the deleted point list in Figure 5 (a), we can see that point 4 has been deleted. In this application, the index of a point in the deleted point list is the index corresponding to a point in the point-edge index table. In other words, in this application, the corresponding index of the deleted point is consistent with the index corresponding to the point in the aforementioned point-edge index table, and represents the same meaning.

[0096] It can be seen from the deleted edge list in (a) of Figure 5 that the four edges in snapshot #0 have been deleted. In this application, the deleted edge list includes the index corresponding to the snapshot where the deleted edge is located, and the index of the point corresponding to the edge list in the snapshot. As shown in (a) of Figure 5, "0:0" in the deleted edge list means that the edge corresponding to the midpoint 0 in the edge list in snapshot #0 has been deleted, that is, the edge with the destination point "1" in the edge list in snapshot #0 has been deleted. Since the destination point "1" in the edge list actually corresponds to an edge with the starting point "0", that is, the edge from point 0 to point 1 in snapshot #1 has been deleted. "0:2" means that the edge corresponding to the midpoint 2 in the edge list in snapshot #0 has been deleted, that is, the edge with the destination point "4" in the edge list in snapshot #0 has been deleted. Since the destination point of the edge corresponding to midpoint "1" in the edge list is 4, the edge from point 1 to point 4 in snapshot #1 has been deleted. "0:4" indicates that the edge corresponding to midpoint 4 in the edge list in snapshot #0 has been deleted. That is, the edge with a destination point of "4" in the edge list in snapshot #0 has been deleted. Since the destination point of the edge corresponding to midpoint "2" in the edge list is 4, the edge from point 2 to point 4 in snapshot #0 has been deleted. "0:5" indicates that the edge corresponding to midpoint 5 in the edge list in snapshot #0 has been deleted. That is, the edge with a destination point of "4" in the edge list in snapshot #0 has been deleted. Since the destination point of the edge corresponding to midpoint "3" in the edge list is 4, the edge from point 3 to point 4 in snapshot #0 has been deleted. To more clearly illustrate the newly added points, newly added edges, deleted points, and deleted edges described in Figure 5 (a), Figures 5 (b) and 5 (c) are drawn based on Figure 5 (a) to facilitate understanding of the technical solutions provided by this application.

[0097] Based on the schematic diagrams of Figures 4 and 5, it can be seen that in a possible scenario, the snapshot in this application includes an original snapshot and multiple incremental snapshots, and each incremental snapshot represents the newly added points and edges at a certain moment, as well as the deletion status of the points and edges contained in the previous snapshot. In this application, each snapshot contains a point-edge index table, the length of which is the sum of the number of points already in the previous snapshot and the number of newly added points in this snapshot. For example, the arrangement order of the original points in the nth snapshot is consistent with that in the n+1th snapshot. In this application, each incremental snapshot also contains a list of newly added edges, a point deletion list, and an edge deletion list. Among them, the point deletion list element contains the index of a point (also called a "number"), and the edge deletion list indicates the index of the snapshot where the deleted edge is located and the edge to be deleted.

[0098] The above Figure 5 mainly introduces the representation method of snapshots in the graph data system proposed in this application. The following mainly introduces how to load and replay the graph based on the stored snapshots in this application. For example, in one possible scenario, the graph data system can read in all snapshots. The graph data system is also used to maintain a point bitmap for marking deleted points and an edge bitmap for marking deleted edges. The deleted points and deleted edges in all snapshots are applied to the point bitmap and edge bitmap. At this time, it can be understood that the replay of a specific graph has been completed.

[0099] FIG6 shows a schematic diagram of a point map and an edge map determined based on the snapshots of FIG4 and FIG5 , wherein FIG6(a) is a schematic diagram of a point map, and FIG6(b) is a schematic diagram of an edge map.

[0100] As shown in (a) of Figure 6, this point bitmap can also be understood as a "bitmap of deleted points." In the point bitmap, the bit length corresponding to the point bitmap is the same as the number of points in the point set corresponding to the last of the M snapshots. Taking Figure 5 as an example, assuming there are two snapshots, snapshot #1 and snapshot #2, the bit length of the point bitmap should be the number of points in the point set corresponding to snapshot #2. Since the point set corresponding to snapshot #2 is {0, 1, 2, 3, 4, 5, 6}, that is, a total of 7 points, the bit length of the point bitmap is 7. Furthermore, since snapshot #2 includes a deleted point list indicating that point "4" has been deleted, in the point bitmap, the bit value corresponding to the deleted point (e.g., point "4") needs to be set to "1", while the bit values ​​corresponding to the remaining undeleted points are all set to "0". Alternatively, the bit value corresponding to the deleted point can be set to "0", while the bit values ​​corresponding to the remaining points can be set to "1".

[0101] As shown in (b) of Figure 6, it can be seen that in the edge bitmap, each snapshot corresponds to an edge bitmap. For example, assuming there are two snapshots, snapshot #1 and snapshot #2, there should be two edge bitmaps, that is, snapshot #1 corresponds to an edge bitmap, and snapshot #2 corresponds to an edge bitmap. Among them, the bit length of the edge bitmap corresponding to snapshot #1 is the same as the number of edges in snapshot #1. From the edge list of snapshot #1 in Figure 4, it can be seen that snapshot #1 includes a total of 7 edges, so the length of the edge list corresponding to snapshot #1 is 7. Furthermore, it can be seen in snapshot #2 that 4 edges in snapshot #1 have been deleted, and the positions of the deleted edges correspond to the index of the points. For example, based on the edge deletion list in snapshot #2, the bit value of the corresponding bit position can be set to "1", and the values ​​of the bit positions corresponding to the remaining undeleted edges can be set to "0". For snapshot #2, there are 3 newly added edges in the edge list of snapshot #2, so the length of the edge bitmap corresponding to snapshot #2 is 3. Since there are no deleted edges in snapshot #2, for example, the values ​​of the bits of the edge bitmap corresponding to snapshot #2 can all be set to "0".

[0102] For example, the following implementation can be used when maintaining point bitmaps and edge bitmaps. In one possible implementation, the graph data system can read and scan snapshots to obtain the number of snapshots 2 and the total number of points L. p is 7, and the number of edges of each snapshot is [7,3], which in turn establishes a length L p =7 bitmap array b p , and the edge bitmap array b e . Where b e The length of [1] is 7, b e The length of [2] is 3, and all bits of these bitmaps are set to 0. The graph data system can scan the snapshot again and apply the information on the point deletion list and edge deletion list of each snapshot to the point bitmap and edge bitmap. That is, the records appearing in these two lists are all set to 1 in the corresponding positions of the corresponding point bitmap and edge bitmap. The above steps complete the playback of the graph.

[0103] FIG7 is a schematic flow chart of loading and replaying snapshots shown in the present application. As shown in FIG7 , first, all M snapshots are scanned and the total number of points L in the M snapshots is obtained. p , and save the number of edges corresponding to each snapshot into array b e Then, the length L p The corresponding bitmap bits are all set to 0. Then, the bitmap bits corresponding to each snapshot are all set to 0. Then, starting from the first snapshot, according to the points marked in the deletion list of each snapshot, L p The corresponding bit of the point in is set to 1. Then, according to the edges marked in the edge deletion list in each snapshot, b eThe corresponding edge bit in is set to 1. For each snapshot, the deletion list of each snapshot and the deletion list of each snapshot are applied to the bitmap until the update of the point bitmap and edge bitmap is completed for all M snapshots. This process can be understood as completing the loading and playback of the graph.

[0104] It should be noted that during the loading and playback process, a certain point may have corresponding edges in multiple snapshots, but because these edge sets have been loaded and stored in memory, the subsequent algorithm needs to query the edge sets of the indexes of the corresponding points in these snapshots and check the status of the point and the corresponding edge in the point bitmap and edge bitmap to determine whether they are valid points and valid edges.

[0105] In response to user graph analysis requests, the graph data system can perform graph queries on specific graphs based on the playback results and generate analysis results for the specific graphs based on the graph query results. The following describes the graph query process in detail, combining different scenarios.

[0106] In one possible implementation, a graph data system may receive a first request message from a user requesting that a target graph be analyzed using a target algorithm. The graph data system may then run a specific graph analysis algorithm based on the M snapshots, perform point queries and / or edge queries on the target graph based on the graph analysis algorithm, and perform algorithm-defined operations on the query results to obtain a graph analysis result. The graph data system then sends a first response message carrying the analysis result of the target graph using the target algorithm, wherein the analysis result is generated based on the point queries and / or edge queries.

[0107] Scene 1

[0108] In one possible application scenario, the graph data system performs a point query on the target graph based on the M snapshots, including: the graph data system determines the set of existing points based on the Mth point set corresponding to the point-edge index table in the Mth snapshot of the M snapshots; determines the points to be deleted based on the list of deleted points included in each of the M snapshots; and determines the set of valid points based on the set of existing points and the deleted points. The above scenario can also be understood as a scenario of querying valid points. For example, it is assumed that the user can indicate to the graph database that they need to query the valid points of the current target graph.

[0109] Exemplarily, in the above scenario one, the graph data system can perform point queries on the target graph based on M snapshots, including: the graph data system can determine a point bitmap based on the point edge index table in the Mth snapshot and the deleted point list in each snapshot, the point bitmap is used to indicate the corresponding status of all points included in the M snapshots, which status includes a deleted state and a non-deleted state, wherein the bit length corresponding to the point bitmap is the same as the number of points in the Mth point set corresponding to the Mth snapshot, and the bit values ​​corresponding to the deleted points in the point bitmap are different from the bit values ​​corresponding to the non-deleted points; according to the point bitmap, a set of valid points is determined.

[0110] Scene 2

[0111] In another possible application scenario, the graph data system performs edge queries on the target graph based on M snapshots, including: the graph data system determines the set of existing edges corresponding to each valid point based on the edge list in each of the M snapshots; the graph data system determines the deleted edges corresponding to each valid point based on the deleted edge list in each of the M snapshots; the graph data system determines the set of valid edges corresponding to each valid point based on the set of existing edges corresponding to each valid point and the deleted edges corresponding to each valid point. This scenario can also be understood as a scenario of querying valid edges. For example, assuming that the user can indicate to the graph database that he needs to query the valid edges of the current target graph

[0112] Exemplarily, in the above-mentioned scenario 2, the graph data system performs an edge query on the target graph based on the M snapshots, including: determining the edge bitmap corresponding to each snapshot based on the edge list and deleted edge list in each of the M snapshots, the edge bitmap being used to indicate the state of the edge in each snapshot, which state includes a deleted state and an undeleted state, wherein the bit length of the edge bitmap corresponding to each snapshot is the same as the number of edges in the edge list in the corresponding snapshot, and the bit value corresponding to the deleted edge in each edge bitmap is different from the bit value corresponding to the undeleted edge; according to the edge bitmap, determining the set of valid edges corresponding to each valid point.

[0113] Scenario 3

[0114] In another possible application scenario, the graph data system can query whether a point specified by the user is a valid point. If it is a valid point, the graph data system will further query the valid edges corresponding to the valid point.

[0115] The technical solution of this application has been described in detail above with reference to Figures 4 to 7 . Now, the technical solution of this application will be described in its entirety with reference to Figure 2 . Taking Figure 2 as an example, the "graph data system" in this application can be understood as the front-end system 210 , business system 220 , graph segmentation system 230 , and graph analysis system 240 in the figure.

[0116] The following describes a graph analysis method 800 proposed in this application. FIG8 shows a schematic flow chart of method 800. For example, this method can be executed by a graph data system or by various components of the graph data system. The method 800 includes the following steps:

[0117] 810. The front-end system 210 receives a first request message from a user, where the first request message is used to request that a target algorithm be used to perform graph analysis on a target graph.

[0118] For example, the user may indicate a target graph and a target algorithm.

[0119] 820. The business system 220 triggers the graph database or file system to generate incremental records periodically or quantitatively.

[0120] In one possible implementation, a graph database or file system may receive a second request message from a monitor module (not shown in the figure), and the second request message is used to request the graph database or file system to export incremental records. Exemplarily, the monitor module may periodically trigger the graph database to export incremental records of the target graph. Exemplarily, the monitor module may monitor the amount of data in the graph database or file system, and when it exceeds a certain threshold, it will trigger the export of incremental records.

[0121] 830. The graph segmentation system 230 generates M snapshots based on the incremental records.

[0122] In the present application, each snapshot includes a point-edge index table and an edge list, and the point-edge index table in the i-th snapshot among M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot, wherein the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot, the i-th point set includes the points in the i-1-th point set and the newly added points in the incremental data corresponding to the i-th snapshot, and the i-1-th point set is the point set corresponding to the point-edge index table in the i-1-th snapshot among the M snapshots.

[0123] In the present application, i is an integer greater than 1, and M is an integer greater than 1.

[0124] In one possible implementation, the i-th snapshot also includes a point deletion list and / or an edge deletion list, wherein the point deletion list is used to indicate the points deleted in the incremental data corresponding to the i-th snapshot; the edge deletion list is used to indicate the edges deleted in the incremental data corresponding to the i-th snapshot, and the j-th snapshot where the deleted edges are located, wherein the j-th snapshot is one of the M snapshots, and j is an integer less than i.

[0125] 840 , the graph analysis system 240 performs point query and / or edge query on the target graph based on the M snapshots.

[0126] For example, the graph analysis system 240 may run a specific graph analysis algorithm, perform point query and / or edge query on the target graph according to the graph analysis algorithm, and perform algorithm-defined operations on the query results to obtain graph analysis results.

[0127] Specifically, the process of point query and edge query can be referred to the above introduction and will not be repeated here.

[0128] 850. The graph analysis system 240 sends a first response message to the user. The first response message carries the analysis result of the target algorithm on the target graph, wherein the analysis result is generated based on the point query and / or edge query.

[0129] It should be noted that the snapshot representation method in this application is not limited to the CSR method, and other representation methods can also be used. As long as the point set corresponding to each snapshot proposed in this application includes all points in the point set corresponding to the previous snapshot, it falls within the scope of protection required by this application. For example, the snapshot representation method can also be: using a map <i:e>, used to map the number of the newly added point i and the set of associated edges e.

[0130] In addition, in the diagrams of Figures 4 and 5, the edges associated with a point are shown as "outgoing edges." In fact, in the snapshot representation method, the edges associated with a point can also be represented as "incoming edges" (that is, the edges associated with a point must be edges with the point as the destination vertex), or the edges associated with a point can include both incoming and outgoing edges, etc. This application does not limit the edges associated with a point to being "outgoing edges" and / or "incoming edges."

[0131] Based on the solution proposed in this application, the point set corresponding to each snapshot includes all points in the point set corresponding to the previous snapshot, which facilitates user operations and improves data processing efficiency. Furthermore, during the loading and playback phase, deleted points and edges can be efficiently applied, making it easier for subsequent graph algorithms to quickly query the validity of a point or edge when obtaining the edge set associated with a point in each snapshot, making graph analysis more convenient. Overall, the solution provided by this application improves the user experience.

[0132] It should be understood that the term "and / or" in this document simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0133] Those skilled in the art should be aware that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is performed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0134] In the embodiment of the present application, the computing device can be divided into functional modules according to the above method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation. The following is an example of dividing each functional module corresponding to each function.

[0135] In this application, the schematic diagrams of the various subsystems and submodules that the device of the graph data system may include can be understood by referring to Figures 2 and 3 above. Each subsystem and submodule can execute each step in the above method 800, which will not be repeated here.

[0136] FIG9 is a schematic block diagram of another apparatus 900 of a graph data system provided in an embodiment of the present application. As shown in the figure, the apparatus includes: at least one processor 920. The processor 920 is coupled to a memory and is configured to execute instructions stored in the memory to send and / or receive signals. Optionally, the apparatus 900 also includes a memory 930 for storing instructions. Optionally, the apparatus 900 also includes a transceiver 910, and the processor 920 controls the transceiver 910 to send and / or receive signals.

[0137] It should be understood that the processor 920 and memory 930 may be combined into one processing device, and the processor 920 is configured to execute the program code stored in the memory 930 to implement the above functions. In specific implementations, the memory 930 may also be integrated into the processor 920 or independent of the processor 920.

[0138] It should also be understood that the transceiver 910 may include a transceiver (or receiver) and a transmitter (or transmitter). The transceiver may further include an antenna, and the number of antennas may be one or more. The transceiver 910 may also be a communication interface or interface circuit.

[0139] As a solution, the apparatus is used to implement the steps in the embodiment of the above method 800. For example, the processor 920 is used to execute the computer program or instructions stored in the memory 930 to implement the various steps in the above method 800.

[0140] It should be understood that the specific process of each transceiver and processor executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.

[0141] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.

[0142] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0143] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous-link DRAM (SLDRAM), and direct RAM-bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0144] According to the method provided in the embodiment of the present application, the present application also provides a computer program product, which stores computer program code. When the computer program code runs on a computer, the computer executes the steps in the embodiment of method 800.

[0145] According to the method provided in the embodiment of the present application, the present application also provides a computer-readable medium, which stores program code. When the program code runs on a computer, the computer executes the steps in the embodiment of the above method 800.

[0146] The explanation of the relevant contents and beneficial effects of any of the above-mentioned devices can be referred to the corresponding method embodiments provided above, which will not be repeated here.

[0147] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a high-density digital video disc (DVD)), or a semiconductor medium (eg, a solid state disc (SSD)).

[0148] In each of the above-mentioned device embodiments, the corresponding modules or units perform the corresponding steps. For example, the transceiver unit (transceiver) performs the receiving or sending steps in the method embodiments, and other steps except sending and receiving can be performed by the processing unit (processor). The functions of the specific units can be referred to in the corresponding method embodiments. There can be one or more processors.

[0149] As used in this specification, the terms "component", "module", "system", etc. are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, through local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component across a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0150] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0151] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0152] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0153] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0154] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0155] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0156] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.< / i:e>

Claims

1. A graph data system, characterized in that: include: M snapshots are stored in the graph data system, each of the M snapshots includes a point-edge index table and an edge list, the point-edge index table in the i-th snapshot among the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot, wherein the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot, the i-th point set includes the points in the i-1-th point set and the newly added points in the incremental data corresponding to the i-th snapshot, the i-1-th point set is the point set corresponding to the point-edge index table in the i-1-th snapshot among the M snapshots, i is an integer greater than 1, and M is an integer greater than 1.

2. The graph data system according to claim 1, characterized in that: The i-th snapshot also includes a point deletion list and / or an edge deletion list, wherein: The deleted point list is used to indicate the points deleted in the incremental data corresponding to the i-th snapshot; The edge deletion list is used to indicate the edge deleted in the incremental data corresponding to the i-th snapshot and the j-th snapshot where the deleted edge is located, wherein the j-th snapshot is one of the M snapshots, and j is an integer less than i.

3. The graph data system according to claim 1, characterized in that: include: The edge list in the first snapshot of the M snapshots is used to indicate the original edges in the original data corresponding to the first snapshot, and the point edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot, and the first point set includes the original points in the original data.

4. The graph data system according to any one of claims 1 to 3, characterized in that: The graph data system further includes: receiving a first request message from a user, wherein the first request message is used to request to perform graph analysis on a target graph using a target algorithm; Performing point query and / or edge query on the target graph according to the M snapshots; A first response message is sent, where the first response message carries an analysis result of the target algorithm on the target graph, wherein the analysis result is generated based on the point query and / or edge query.

5. The graph data system according to claim 4, characterized in that: The performing point query on the target graph according to the M snapshots includes: Determine a set of existing points according to an Mth point set corresponding to the point-edge index table in an Mth snapshot among the M snapshots; Determine a point to be deleted according to the deletion point list included in each of the M snapshots; The set of valid points is determined according to the set of existing points and the deleted points.

6. The graph data system according to claim 4 or 5, characterized in that: The performing point query on the target graph according to the M snapshots includes: Determine a point bitmap according to the point edge index table in the Mth snapshot and the deleted point list in each snapshot, wherein the point bitmap is used to indicate states corresponding to all points included in the M snapshots, the states including a deleted state and a non-deleted state, wherein a bit length corresponding to the point bitmap is the same as the number of points in the Mth point set corresponding to the Mth snapshot, and a bit value corresponding to the deleted point in the point bitmap is different from a bit value corresponding to the non-deleted point; According to the point map, the set of valid points is determined.

7. The graph data system according to claim 4, characterized in that: The performing edge query on the target graph according to the M snapshots includes: Determine, according to the edge list in each of the M snapshots, a set of existing edges corresponding to each valid point; Determine, according to the edge deletion list in each of the M snapshots, the deleted edge corresponding to each valid point; The set of valid edges corresponding to each valid point is determined according to the set of existing edges corresponding to each valid point and the deleted edges corresponding to each valid point.

8. The graph data system according to claim 4 or 7, characterized in that: The performing edge query on the target graph according to the M snapshots includes: Determine, according to the edge list and the deleted edge list in each of the M snapshots, an edge bitmap corresponding to each snapshot, the edge bitmap being used to indicate a state of an edge in each snapshot, the state including a deleted state and a non-deleted state, wherein a bit length of the edge bitmap corresponding to each snapshot is the same as the number of edges in the edge list in the corresponding snapshot, and a bit value corresponding to the deleted edge in each edge bitmap is different from a bit value corresponding to the non-deleted edge; According to the edge bitmap, a set of valid edges corresponding to each of the valid points is determined.

9. The graph data system according to any one of claims 1 to 8, characterized in that: The graph data system further includes: Receive a second request message from the first module, the second request message is used to request the graph data system to export an incremental record, the incremental record including at least one of the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges; The snapshot corresponding to the incremental record is generated according to the incremental record, wherein the snapshot includes at least one of the following items: the point-edge index table, the edge list, the point deletion list, and the edge deletion list.

10. A device for a graph data system, characterized in that: include: A processor and a memory, wherein the processor is used to execute a computer program or instruction stored in the memory, wherein: M snapshots are stored in the memory, each of the M snapshots includes a point-edge index table and an edge list, the point-edge index table in the i-th snapshot among the M snapshots is used to indicate the index relationship between the points in the i-th point set and the edges in the edge list in the i-th snapshot, wherein the edge list in the i-th snapshot is used to indicate the newly added edges in the incremental data corresponding to the i-th snapshot, the i-th point set includes the points in the i-1-th point set and the newly added points in the incremental data corresponding to the i-th snapshot, the i-1-th point set is the point set corresponding to the point-edge index table in the i-1-th snapshot among the M snapshots, i is an integer greater than 1, and M is an integer greater than 1.

11. The device according to claim 10, characterized in that The i-th snapshot also includes a point deletion list and / or an edge deletion list, wherein: The deleted point list is used to indicate the points deleted in the incremental data corresponding to the i-th snapshot; The edge deletion list is used to indicate the edge deleted in the incremental data corresponding to the i-th snapshot and the j-th snapshot where the deleted edge is located, wherein the j-th snapshot is one of the M snapshots, and j is an integer less than i.

12. The device according to claim 10, characterized in that include: The edge list in the first snapshot of the M snapshots is used to indicate the original edges in the original data corresponding to the first snapshot, and the point edge index table in the first snapshot is used to indicate the index relationship between the points in the first point set corresponding to the first snapshot and the edges in the edge list in the first snapshot, and the first point set includes the original points in the original data.

13. The device according to any one of claims 10 to 12, characterized in that The device also includes a transceiver, wherein The transceiver is used to receive a first request message, where the first request message is used to request to perform graph analysis on a target graph using a target algorithm; The processor is used for performing point query and / or edge query on the target graph according to the M snapshots; A first response message is sent, where the first response message carries an analysis result of the target algorithm on the target graph, wherein the analysis result is generated based on the point query and / or edge query.

14. The device according to claim 13, characterized in that The processor is used to perform a point query on the target graph according to the M snapshots, including: The processor is used to determine a set of existing points according to an Mth point set corresponding to the point-edge index table in an Mth snapshot among the M snapshots; The processor is configured to determine a point to be deleted according to the deletion point list included in each of the M snapshots; The processor is used to determine the set of valid points according to the set of existing points and the deleted points.

15. The device according to claim 13 or 14, characterized in that The processor is used to perform a point query on the target graph according to the M snapshots, including: The processor is used to determine a point bitmap according to the point edge index table in the Mth snapshot and the deleted point list in each snapshot, wherein the point bitmap is used to indicate the states corresponding to all the points included in the M snapshots, the states including a deleted state and a non-deleted state, wherein the bit length corresponding to the point bitmap is the same as the number of points in the Mth point set corresponding to the Mth snapshot, and the bit values ​​corresponding to the deleted points in the point bitmap are different from the bit values ​​corresponding to the non-deleted points; The processor is used to determine the set of valid points according to the point map.

16. The device according to claim 13, characterized in that The processor is used to perform edge query on the target graph according to the M snapshots, including: The processor is used to determine a set of existing edges corresponding to each valid point according to the edge list in each of the M snapshots; The processor is configured to determine the deleted edge corresponding to each valid point according to the deleted edge list in each of the M snapshots; The processor is used to determine the set of valid edges corresponding to each valid point according to the set of existing edges corresponding to each valid point and the deleted edges corresponding to each valid point.

17. The device according to claim 13 or 16, characterized in that The processor is used to perform edge query on the target graph according to the M snapshots, including: The processor is used to determine, according to the edge list and the deleted edge list in each of the M snapshots, an edge bitmap corresponding to each snapshot, the edge bitmap being used to indicate the state of the edge in each snapshot, the state including a deleted state and a non-deleted state, wherein the bit length of the edge bitmap corresponding to each snapshot is the same as the number of edges in the edge list in the corresponding snapshot, and the bit value corresponding to the deleted edge in each edge bitmap is different from the bit value corresponding to the non-deleted edge; The processor is used to determine, according to the edge bitmap, a set of valid edges corresponding to each of the valid points.

18. The device according to any one of claims 10 to 17, characterized in that The device also includes a transceiver, wherein The transceiver is used to receive a second request message, wherein the second request message is used to request the graph data system to export an incremental record, wherein the incremental record includes at least one of the following data: data of newly added points, data of newly added edges, data of deleted points, and data of deleted edges; The processor is used to generate the snapshot corresponding to the incremental record according to the incremental record, and the snapshot includes at least one of the following items: the point-edge index table, the edge list, the point deletion list, and the edge deletion list.

19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the actions executed by the graph data system as described in any one of claims 1 to 9.

20. A computer program product, characterized in that The invention comprises instructions which, when executed on a computer, cause the computer to execute the actions executed by the graph data system as claimed in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Object storage management method and system for billion-level node scale knowledge graph based on Ceph

    CN111639082A

  • Large-scale time-varying graph storage method and system based on snapshot similarity

    CN114064982A

  • Hybrid distributed graph data storage and calculation method

    CN117112692A

  • Fast and memory efficient in-memory columnar graph updates while preserving analytical performance

    US20220284056A1