A data writing method, device, apparatus and storage medium
Patent Information
- Application Number
- CN202211713253.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-12-29
AI Technical Summary
[0015]本申请实施例提供的一种数据写入方法、装置、设备及存储介质,通过为图谱结构的顶点和边分别设置不同的属性信息,并根据属性信息与每次写入单个图数据库节点的数据数量的对应关系,确定每个数据表每次写入图数据库集群的数据数量,可以使得不同顶点的数据以及不同边的数据分别依次导入,不会因为每次写入图数据库集群的数据数量相同而造成的数据的堆积或者无法充分利用硬件资源的问题。
Smart Images

Figure CN116049277B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of graph database technology, and in particular to a data writing method, apparatus, device and storage medium. Background Technology
[0002] In recent years, with the widespread application of digital technologies such as the Internet of Things and artificial intelligence, enterprise data has experienced explosive growth, and the complexity of relationships between data has also surged. Graph databases, due to their ability to better clarify and utilize relationships between data, have a significant advantage in handling complex problems and are therefore widely used in the field of data management.
[0003] Data import is a crucial step in building a graph database. Currently, the data import method involves: determining the attribute items mapped to each node and edge in the graph structure; extracting the attribute information of each node from the data to be written based on the node mapping attribute items, and extracting the attribute information of each edge from the data to be written based on the edge mapping attribute items; generating a triplet object (vertex-edge-vertex) file based on the attribute information of each node, the attribute information of each edge, and the association relationship between nodes and edges in the sample graph structure, and importing this triplet object file into the graph database. In this data import method, the data corresponding to nodes and edges need to be imported simultaneously in the same batch, which is unsuitable for scenarios where only points can be imported first, followed by edges. Furthermore, when the number of attributes associated with a point is much greater than the number of attributes associated with an edge, importing data in the same batch may lead to excessively large batches, resulting in data accumulation in the graph database, or insufficient batches, preventing full utilization of the graph database's hardware resources. Summary of the Invention
[0004] This application provides a data writing method, apparatus, device, and storage medium, which enables the import of data for different vertices and different edges of a graph structure in separate batches, ensuring full utilization of hardware resources without causing data accumulation.
[0005] In a first aspect, embodiments of this application provide a data writing method, the method comprising: Obtain multiple data tables to be written to the graph database cluster. The categories of the data tables include entity data tables that describe entity attributes and relationship data tables that describe the relationship between a pair of entity data table attributes. For each data table, based on the category and table structure of the data table, query the target attribute information that matches from multiple predefined attribute information describing the graph structure; Based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, the number of data written to each data table to the graph database cluster each time is determined; Each data table is written to the graph database cluster in sequence according to the number of data items corresponding to each data table.
[0006] In one possible implementation, for each data table, based on the category and table structure of the data table, matching attribute information is queried from multiple predefined attribute information describing the graph structure, including: If the data table belongs to the category of entity data table, according to the field describing entity attributes in the entity data table, query the target attribute information that matches the entity data table from multiple predefined attribute information describing the vertex attributes of the graph structure; If the data table belongs to the category of relational data table, then based on the field in the relational data table that describes the relationship between the attributes of a pair of entity data tables, the target attribute information that matches the relational data table is queried from multiple attribute information of predefined description graph structure edge attributes.
[0007] In one possible implementation, determining the number of data tables written to the graph database cluster each time, based on a pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, includes: Based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, determine the number of data written to a single graph database node for each data table each time; The amount of data written to each data table in each graph database cluster is determined by multiplying the amount of data written to a single graph database node each time by the number of nodes in the graph database cluster and the maximum concurrency supported by a single graph database node.
[0008] In one possible implementation, before sequentially writing each data table into the graph database cluster according to the number of data items corresponding to each data table, the method further includes: After negotiating with the server to determine the vertex identifiers for generating the graph structure, the generated vertex identifiers are sent to the server's graph database cluster, wherein the number of vertex identifiers matches the number of data in all entity data tables to be written to the graph database cluster.
[0009] In one possible implementation, the vertex identifier is in hash format, and each vertex in the graph structure has a different vertex identifier, with the vertex identifier having no more than a preset number of bits.
[0010] In one possible implementation, before sequentially writing each data table into the graph database cluster according to the number of data items corresponding to each data table, the method further includes: After negotiating with the server to determine which data tables to be written to the graph database cluster should be cleaned up, duplicate data in the data tables to be written to the graph database cluster should be cleaned up.
[0011] In one possible implementation, before sequentially writing each data table into the graph database cluster according to the number of data items corresponding to each data table, the method further includes: For each target attribute information, if the description of the target attribute information includes a data consistency parameter, then after negotiating with the server to determine whether to lock the data, the data in the corresponding data table will be locked.
[0012] Secondly, embodiments of this application provide a data writing device, the device comprising: The acquisition module is used to acquire multiple data tables to be written to the graph database cluster. The categories of the data tables include entity data tables that describe entity attributes and relation data tables that describe the relationship between a pair of entity data table attributes. The query module is used to query the target attribute information that matches the data table from multiple predefined attribute information describing the graph structure, based on the category and table structure of the data table. The determination module is used to determine the number of data tables written to the graph database cluster each time, based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time. The write module is used to write each data table sequentially into the graph database cluster according to the number of data corresponding to each data table.
[0013] Thirdly, embodiments of this application provide a data writing device, the device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method as described in the first aspect above.
[0014] Fourthly, embodiments of this application provide a computer storage medium storing a computer program for causing a computer to perform the method described in the first aspect above.
[0015] This application provides a data writing method, apparatus, device, and storage medium. By setting different attribute information for the vertices and edges of the graph structure, and determining the number of data tables written to the graph database cluster each time based on the correspondence between the attribute information and the number of data written to a single graph database node each time, the data of different vertices and different edges can be imported sequentially, avoiding the problems of data accumulation or insufficient utilization of hardware resources caused by the same number of data written to the graph database cluster each time. Attached Figure Description
[0016] Figure 1 This application provides a schematic diagram of a graph database storage backend. Figure 2 This is a schematic diagram illustrating an application scenario of a data writing method provided in an embodiment of this application. Figure 3 This is a schematic flowchart of a data writing method provided in an embodiment of this application; Figure 4 A schematic diagram of a spectral structure provided in an embodiment of this application; Figure 5 This application provides a schematic diagram of a specific process for a data writing method applied to a client. Figure 6 This application provides a schematic diagram of a specific process for a data writing method applied to a server. Figure 7 A schematic diagram of a data writing device provided in an embodiment of this application; Figure 8 This is a schematic diagram of a data writing device provided in an embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0018] First, some concepts involved in the embodiments of this application will be introduced: 1. Graph: A graph consists of two elements: nodes and relationships. Each node represents an entity (land, thing, category, or other data), and each relationship represents the way two nodes are associated.
[0019] 2. Property Graph: Defines a graph model. A property graph is a directed graph consisting of vertices, edges, labels, and properties. Vertices are also called nodes, and edges are also called relationships.
[0020] 3. JanusGraph, a graph database engine, uses property graphs for modeling. JanusGraph's modular architecture allows it to adapt to various storage and indexing backends. Storage backends refer to the databases that actually store the graph's vertex and edge data, such as Cassandra and HBase; indexing backends refer to systems that leverage their own features to support more retrieval functions such as fuzzy search, geographic coordinate search, and full-text search. Optional indexing backend systems include ElasticSearch, Solr, and Lucene. This invention uses HBase as the storage backend and ElasticSearch (ES) as the external indexing backend as an example.
[0021] 4. Graph storage: such as Figure 1 As shown, JanusGraph stores data in an edge-cutting manner. Taking HBase as the storage backend as an example, the HBase Rowkey is used as the vertex identifier, and the attributes or edges associated with the vertex are stored as units. The graph database provided in this embodiment can also be Neo4j, Arangodb, Orientdb, etc., without specific limitations here.
[0022] like Figure 2 The diagram illustrates an application scenario of a data writing method provided in this embodiment. This scenario includes graph database nodes (servers 201_1, 201_2, and 201_N as shown) and a client 202, used to send data tables (202_1, 202_2, and 202_N as shown). Multiple graph database nodes form a graph database cluster, which can be deployed across multiple servers or within a single server; this embodiment does not impose specific limitations. Each graph database node receives and stores the data tables uploaded by the client 202.
[0023] To solve the problem in the prior art that different vertex data and different edge data cannot be uploaded separately, which further causes data accumulation in the graph database due to an excessively large batch setting and fails to fully utilize the hardware resources of the graph database when the batch setting is too small, embodiments of the present application provide a data writing method, as Figure 3 shown, the method includes: S301: Obtain a plurality of data tables to be written into the graph database cluster.
[0024] Entity data tables store entity names, and relationship data tables store associations between entities, Figure 4 is a schematic diagram of the relationship between entities provided in the embodiments of the present application. For vertex 1: the "arrow" connected to vertex 1 is its out-edge (outE), and vertex 2 is the in-vertex (inV) of this out-edge; for vertex 2: the "arrow" connected to vertex 1 is its in-edge (inE), and vertex 1 is the out-vertex (outV) of this in-edge.
[0025] The data tables provided in the embodiments of the present application include two categories: (1) Entity data table.
[0026] An entity data table is a table describing attributes of a certain entity, as shown in Table 1 (vehicle data).
[0027] Table 1 Vehicle
[0028] For data tables describing different entities, the table structures (number and content of fields) may be different; in addition, the table structures (number and content of fields) describing the same entity may also be different.
[0029] (2) Relationship data table.
[0030] A relationship data table is a table describing the relationship between attributes of a pair of entity data tables. As shown in Table 2, "1" is used to identify that there is an ownership relationship between the vehicle owner and the vehicle, for example, vehicle owner A owns the vehicle with license plate number "Jing XX", and vehicle owner C owns the vehicle with license plate number "Ji XX".
[0031] Table 2
[0032] The table structure of the relationship data table can also be as shown in Table 3: Table 3
[0033] S302: For each data table, query matched target attribute information from a plurality of pre-defined attribute information describing the graph structure according to the category and table structure of the data table.
[0034] In this embodiment of the application, multiple attribute information is predefined for the graph structure. Different types of vertices (e.g., car owner, car) and different types of edges (e.g., car owner-car owner relationship, car owner-car relationship) have different attribute information.
[0035] In one possible implementation, for each data table, based on the category and structure of the data table, matching attribute information is queried from multiple predefined attribute information describing the graph structure, including the following two types: (1) If the data table belongs to the category of entity data table, according to the field describing entity attributes in the entity data table, query the target attribute information that matches the entity data table from multiple attribute information that predefined the description of the vertex attributes of the graph structure.
[0036] For example, the attribute information set according to Table 1 is shown in Table 4 (Target Attribute Information of Vertices). Here, the "cardinality" feature indicates that attribute "attribute 1" is unique; composite index indicates an equal-value index; for example, indexing directly by attribute 1 will retrieve the user corresponding to attribute 1; and mixed index indicates a fuzzy index; for example, indexing by address will retrieve other information about that address. Furthermore, if some data in the entity data table is not needed, it can be left undefined. For example, Table 4 only defines attributes 1, 2, and 3, while other attributes are not defined. When importing data into the graph database cluster later, the data corresponding to other attributes will not be imported.
[0037] Table 4
[0038] (2) If the data table belongs to the category of relational data table, then according to the field in the relational data table that describes the relationship between the attributes of a pair of entity data tables, the target attribute information that matches the relational data table is queried from the multiple attribute information of the predefined graph structure edge attributes.
[0039] For example, attribute information can be set according to Table 2 as shown in Table 5 (target attribute information of edges).
[0040] Table 5
[0041] If the relational data table does not explicitly specify which pair of entities are related to it, then the attribute information can define the entity data table that is associated with the relational data table, as shown in Table 6.
[0042] Table 6
[0043] S303: Based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, determine the number of data tables written to the graph database cluster each time.
[0044] To prevent slow data import due to excessively small batches and service timeouts due to excessively large batches, batches are differentiated based on the memory usage of vertex-related attribute information and edge-related attribute information. The number of data items (tokens) imported into the graph database cluster each time is set to maximize the utilization of graph database hardware resources and improve write performance.
[0045] Generally, increasing the batch size for data import can reduce network interactions between the client and server, thereby improving data import performance. However, excessively large batches can increase memory usage, impacting server latency and ultimately reducing import performance. Furthermore, due to the characteristics of the security industry, the memory usage of vertices and edges varies depending on their attribute information, leading to differences in import performance. For example, the "car owner" node has more attributes, while the "car" node has fewer attributes, resulting in different numbers of HBase entries and Elasticsearch indexes imported, with the "car" node showing better import performance. Therefore, depending on the actual data, a larger batch size can be set for "cars".
[0046] Based on the pre-established correspondence between the target attribute information and the amount of data written to a single graph database node each time, the amount of data written to each data table to the graph database cluster each time is determined, including: Based on the pre-established correspondence between the target attribute information and the number of data (batch size) written to a single graph database node each time, determine the number of data tables written to a single graph database node each time. The amount of data written to each data table in each graph database cluster is determined by multiplying the amount of data written to a single graph database node each time by the number of nodes in the graph database cluster and the maximum concurrency supported by a single graph database node.
[0047] As shown in Tables 4 (Target Attribute Information of Vertices) and 5 (Target Attribute Information of Edges) above, if the memory size occupied by the target attribute information of vertices is much larger than that occupied by the target attribute information of edges, then based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, the number of data entries for the target attribute information of vertices written to a single graph database node each time is determined, such as 2000; and the number of data entries for the target attribute information of vertices written to a single graph database node each time is determined, such as 5000. Alternatively, the target attribute information of vertices and edges can be combined to set a unified batch, for example, 1500 batches of entity table data corresponding to vertices + 3000 batches of relation table data corresponding to edges = 4500 batches.
[0048] In the prior art, during the process of importing a data table into a graph database cluster, the server generates vertex identifiers for the imported data, performs verification (index type verification, cardinality type verification), cleans the imported data, and locks the data, resulting in slow data entry. However, in this embodiment, the generation of vertex identifiers, data cleanup, and data locking are performed on the client side to improve import efficiency.
[0049] Therefore, before writing each data table sequentially into the graph database cluster according to the number of data items corresponding to each data table, the method further includes: After negotiating with the server to determine the vertex identifiers for generating the graph structure, the generated vertex identifiers are sent to the server's graph database cluster, wherein the number of vertex identifiers matches the number of data in all entity data tables to be written to the graph database cluster.
[0050] On the server side, a first switch for generating vertex identifiers can be pre-set. When it is agreed with the server to generate vertex identifiers, the first switch on the server side is turned off, indicating that the vertex identifiers will be generated by the client and the server will not perform the vertex identifier generation step. On the other hand, if the negotiation fails, the vertex identifiers will be generated by the server and the first switch will be in the open state.
[0051] The vertex identifier must meet the following characteristics: the vertex identifier is in hash format, each vertex in the graph structure has a different vertex identifier, and the number of bits in the vertex identifier is no greater than a preset number of bits.
[0052] The purpose of hashing is to prevent the accumulation of data in the entity tables corresponding to each vertex when all points are stored on a single RegionServer, which would lead to a decrease in data import or query performance, especially when the storage backend is HBase. Different points having the same vertex identifier will cause data overwriting and compromise data integrity; therefore, it is essential to maintain the uniqueness of vertex identifiers. Furthermore, fewer bits can reduce the memory or disk storage overhead of the graph database.
[0053] After negotiating with the server to determine which data tables to be written to the graph database cluster should be cleaned up, duplicate data in the data tables to be written to the graph database cluster should be cleaned up.
[0054] A second data cleanup switch can be pre-configured on the server side. Once data cleanup is agreed upon after negotiation with the server, the second switch on the server side will be turned off, indicating that the data cleanup will be performed by the client and the server will not perform the data cleanup step. On the other hand, if the negotiation fails, the data cleanup will be performed by the server and the second switch will be turned on.
[0055] Data cleanup should be performed on the client side to avoid the graph database cluster using locking mechanisms to ensure data consistency, which could reduce data import performance. Scenarios requiring cleanup include situations where a unique index has been created for a field, and the imported data must ensure that this field is unique within the graph.
[0056] Additionally, some validation rules can be set on the client side to validate the data table to be written, thereby reducing the server-side deduplication validation logic of a large number of graph database clusters.
[0057] For each target attribute information, if the description of the target attribute information includes a data consistency parameter, then after negotiating with the server to determine whether to lock the data, the data in the corresponding data table will be locked.
[0058] A third switch for data locking can be pre-configured on the server side. Once data locking is agreed upon after negotiation with the server, the third switch on the server side is turned off, indicating that the data locking is performed by the client and the server does not perform the data locking step. On the other hand, if the negotiation fails, the data locking is performed by the server and the third switch is turned on.
[0059] Data consistency means that the same data is identical after being imported into the storage backend and the index backend. For data that requires consistency, it can be defined in the attribute information, as shown in Table 7.
[0060] Table 7
[0061] By locking the corresponding data (attribute 1) based on the attribute information in Table 7, it can be ensured that the data will not be changed or lost when imported into the graph database cluster.
[0062] S304: Write each data table into the graph database cluster sequentially according to the number of data corresponding to each data table.
[0063] When importing the graph database cluster, each data table is imported sequentially according to the amount of data written to the graph database cluster each time, as obtained in S303. For example, the data is imported in the order of Table 4-Table 1-Table 2. This application embodiment does not specifically limit the import order of the data tables.
[0064] Additionally, the target attribute information of vertices and edges can be combined and set into a single batch, importing both entity data tables and relational data tables simultaneously (e.g., importing in the order of Table 4 + Table 1 - Table 2). For example, 1500 batches of vertices corresponding to Table 1 data + 3000 batches of edges corresponding to Table 4 data = 4500 batches. In other words, when importing into a graph database cluster, 1500 batches can be imported simultaneously. Number of nodes in the graph database Table 1 contains data on the vertices corresponding to the maximum concurrency supported by a single node in the graph database, plus 3000 batches. Number of nodes in the graph database Table 4 shows the data corresponding to the maximum number of edges that a single node in the graph database can support for concurrency.
[0065] The following is through Figure 5 and Figure 6 This application provides a detailed description of a data writing method according to an embodiment.
[0066] (1) such as Figure 5 The diagram shows the specific flow of the data writing method applied to the client.
[0067] Step 1: Query the target attribute information corresponding to the data tables to be written to the graph database cluster; Step 2: Generate vertex identifiers; Step 3: Clean up duplicate data in each data table; Step 4: Based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, determine the number of data tables written to a single graph database node each time; Step 5: Determine the amount of data written to each data table in each graph database cluster based on the amount of data written to a single graph database node in a single transaction and the number of nodes in the graph database cluster; Step 6: Write each data table into the graph database cluster in sequence.
[0068] (2) such as Figure 6 The diagram shows the specific process of the data writing method applied to the server.
[0069] S601: Determine whether the first switch for generating vertex identifiers is turned on. If yes, execute S602; otherwise, execute S603. S602: Generate distributed vertex identifiers; S603: Determine whether the imported data table is the one corresponding to the vertex. If yes, execute S604; otherwise, execute S605. S604: Bind undefined attribute information. If the imported data table is the data table corresponding to the edge, then it is not necessary to bind undefined attribute information. S605: Determine whether the second switch for data cleanup is turned on. If yes, execute S606; otherwise, execute S607. S606: Attribute validation, unique index validation, data cleanup; S607: Determine whether the third switch for data locking is open. If yes, execute S608; otherwise, execute S609. S608: Data locking; S609: Assemble the data of the storage backend and index backend, and serialize the data; S610: Determine if there is an error in writing data. If so, execute S611; otherwise, end the process after writing the data. S611: Transaction rollback; For the JanusGraph graph database, all interactions with JanusGraph are associated with a transaction. JanusGraph transactions are multi-threaded concurrent operations; the same data is first written to the storage backend and then to the index backend. When the write to the storage backend fails (power outage, etc.), the transaction is rolled back to before the write to the storage backend. If the write to the storage backend is normal but the write to the index backend fails, the transaction is rolled back to before the write to the index backend.
[0070] S612: Set a scheduled task to rewrite the data.
[0071] Additionally, attribute information can be validated on the server side. When the attribute information of a certain data table does not exist, the server can automatically infer the corresponding attribute information of the data table and create the attribute information that matches the data table.
[0072] In this embodiment, by determining different batches based on the attribute information corresponding to different data tables, it is possible to prevent slow data entry due to excessively small batches and service timeouts due to excessively large batches caused by message backlog. The client generates unique identifiers in advance and avoids duplicate data entry, reducing the time spent by the server in generating vertex identifiers and performing deduplication logic verification, thereby improving the data writing speed. Based on the characteristics of data in the security field, where different types of attribute information occupy different amounts of memory and have significant differences in import performance, the client writes data to the graph database cluster in real time according to the number of tokens, making full use of graph database hardware resources and improving data import performance.
[0073] Based on the same inventive concept, this application also provides a data writing device 700, such as... Figure 7As shown, the device includes: The acquisition module 701 is used to acquire multiple data tables to be written to the graph database cluster. The categories of the data tables include entity data tables that describe entity attributes and relation data tables that describe the relationship between a pair of entity data table attributes. The query module 702 is used to query the target attribute information that matches the data table from multiple attribute information of the predefined description graph structure, based on the category and table structure of the data table. The determination module 703 is used to determine the number of data tables written to the graph database cluster each time based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time. The write module 704 is used to write each data table sequentially into the graph database cluster according to the number of data corresponding to each data table.
[0074] In one possible implementation, the query module 702 is used to, for each data table, query matching attribute information from a predefined set of attribute information describing the graph structure, based on the category and table structure of the data table, including: If the data table belongs to the category of entity data table, according to the field describing entity attributes in the entity data table, query the target attribute information that matches the entity data table from multiple predefined attribute information describing the vertex attributes of the graph structure; If the data table belongs to the category of relational data table, then based on the field in the relational data table that describes the relationship between the attributes of a pair of entity data tables, the target attribute information that matches the relational data table is queried from multiple attribute information of predefined description graph structure edge attributes.
[0075] In one possible implementation, the determining module 703 is used to determine the number of data tables written to the graph database cluster each time, based on a pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time. Based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, determine the number of data written to a single graph database node for each data table each time; The amount of data written to each data table in each graph database cluster is determined by multiplying the amount of data written to a single graph database node each time by the number of nodes in the graph database cluster and the maximum concurrency supported by a single graph database node.
[0076] In one possible implementation, the apparatus further includes a negotiation module, used to write each data table sequentially to the graph database cluster according to the number of data items corresponding to each data table, and further includes: After negotiating with the server to determine the vertex identifiers for generating the graph structure, the generated vertex identifiers are sent to the server's graph database cluster, wherein the number of vertex identifiers matches the number of data in all entity data tables to be written to the graph database cluster.
[0077] In one possible implementation, the negotiation module is used to determine that the vertex identifier is in hash format, and that each vertex in the graph structure has a different vertex identifier, wherein the number of bits in the vertex identifier is no greater than a preset number of bits.
[0078] In one possible implementation, before the negotiation module sequentially writes each data table to the graph database cluster according to the number of data items corresponding to each data table, it further includes: After negotiating with the server to determine which data tables to be written to the graph database cluster should be cleaned up, duplicate data in the data tables to be written to the graph database cluster should be cleaned up.
[0079] In one possible implementation, before the negotiation module sequentially writes each data table to the graph database cluster according to the number of data items corresponding to each data table, it further includes: For each target attribute information, if the description of the target attribute information includes a data consistency parameter, then after negotiating with the server to determine whether to lock the data, the data in the corresponding data table will be locked.
[0080] Based on the same inventive concept, this application also provides a data writing device for executing the data writing method provided in this application.
[0081] The device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the data writing methods in the above embodiments.
[0082] The following reference Figure 8 This describes the data writing device 130 according to this embodiment of the application. Figure 8 The data writing device 130 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0083] like Figure 8As shown, the data writing device 130 is presented in the form of a general-purpose data writing device. The components of the data writing device 130 may include, but are not limited to: at least one processor 131, at least one memory 132, and a bus 133 connecting different system components (including memory 132 and processor 131).
[0084] The processor 131 is used to read and execute instructions from the memory 132, so that the at least one processor can execute the data writing method provided in the above embodiments.
[0085] Bus 133 represents one or more of several bus structures, including a memory bus or memory controller, peripheral bus, processor, or local bus using any of the various bus structures.
[0086] The memory 132 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 1321 and / or cache memory 1322, and may further include read-only memory (ROM) 1323.
[0087] The memory 132 may also include a program / utility 1325 having a set (at least one) of program modules 1324, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0088] The data writing device 130 can also communicate with one or more external devices 134 (e.g., keyboard, pointing device, etc.), one or more devices that enable a user to interact with the data writing device 130, and / or any device that enables the data writing device 130 to communicate with one or more other data writing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 135. Furthermore, the data writing device 130 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 136. As shown, network adapter 136 communicates with other modules used for the data writing device 130 via bus 133. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the data writing device 130, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0089] In some possible implementations, various aspects of the data writing method provided in this application can also be implemented in the form of a program product, which includes program code that, when the program product is run on a computer device, causes the computer device to perform the steps of a data writing method according to various exemplary embodiments of this application as described above.
[0090] In addition, this application also provides a computer-readable storage medium storing a computer program for causing a computer to perform the method described in any of the above embodiments.
[0091] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0092] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0093] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0094] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A data writing method, characterized in that, Applied to a client, the method includes: Obtain multiple data tables to be written to the graph database cluster. The categories of the data tables include entity data tables that describe entity attributes and relationship data tables that describe the relationship between a pair of entity data table attributes. For each data table, based on the category and table structure of the data table, query the target attribute information that matches from multiple predefined attribute information describing the graph structure; Based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, the number of data written to a single graph database node for each data table is determined. Based on the product of the number of data written to a single graph database node each time, the number of nodes in the graph database cluster, and the maximum concurrency supported by a single graph database node, the number of data sent to the graph database cluster for each data table each time is determined. If the target attribute information includes a data consistency parameter, then after negotiating with the server to determine whether to lock the data, the data in the data table corresponding to the data consistency is locked, and the data locking switch pre-set on the server is in the off state; After negotiating with the server to determine the vertex identifiers for generating the graph structure, the generated vertex identifiers are sent to the server's graph database cluster. The switch for generating vertex identifiers, which is pre-configured on the server, is in the off state. After negotiating with the server to determine the preset processing for the multiple data tables to be written to the graph database cluster, data is sent to the graph database cluster sequentially according to the number of data corresponding to each data table. The switch on the server that is preset to perform the preset processing on the multiple data tables is in the off state.
2. The method according to claim 1, characterized in that, For each data table, based on the table's category and structure, matching attribute information is queried from multiple predefined attribute information describing the graph structure, including: If the data table belongs to the category of entity data table, according to the field describing entity attributes in the entity data table, query the target attribute information that matches the entity data table from multiple predefined attribute information describing the vertex attributes of the graph structure; If the data table belongs to the category of relational data table, then based on the field in the relational data table that describes the relationship between the attributes of a pair of entity data tables, the target attribute information that matches the relational data table is queried from multiple attribute information of predefined description graph structure edge attributes.
3. The method according to claim 1, characterized in that, The number of vertex identifiers matches the number of data entries in all entity data tables to be written to the graph database cluster.
4. The method according to claim 1, characterized in that, The vertex identifier is in hash format, and each vertex in the graph structure has a different vertex identifier. The number of bits in the vertex identifier is no greater than a preset number of bits.
5. The method according to claim 1, characterized in that, Before writing each data table sequentially into the graph database cluster according to the number of data items corresponding to each data table, the process also includes: After negotiating with the server to determine which data tables to be written to the graph database cluster should be cleaned up, duplicate data in the data tables to be written to the graph database cluster should be cleaned up.
6. A data writing device, characterized in that, The device includes: The acquisition module is used to acquire multiple data tables to be written to the graph database cluster. The categories of the data tables include entity data tables that describe entity attributes and relation data tables that describe the relationship between a pair of entity data table attributes. The query module is used to query the target attribute information that matches the data table from multiple predefined attribute information describing the graph structure, based on the category and table structure of the data table. The determination module is used to determine the number of data written to a single graph database node for each data table each time based on the pre-established correspondence between the target attribute information and the number of data written to a single graph database node each time, and to determine the number of data sent to the graph database cluster for each data table each time based on the product of the number of data written to a single graph database node each time, the number of nodes in the graph database cluster, and the maximum concurrency supported by a single graph database node. The writing module is used to, when determining that the target attribute information includes data consistency parameters, negotiate with the server to determine that data locking processing is required, and then lock the data in the data table corresponding to the data consistency, with the data locking switch pre-set on the server in a closed state; after negotiating with the server to determine that vertex identifiers for generating the graph structure are generated, the generated vertex identifiers are sent to the graph database cluster on the server, with the vertex identifier generation switch pre-set on the server in a closed state; after negotiating with the server to determine that preset processing is required for multiple data tables to be written to the graph database cluster, data from each data table is sent to the graph database cluster sequentially according to the number of data tables corresponding to each data table, with the preset processing switch pre-set on the server for the multiple data tables in a closed state.
7. A data writing device, characterized in that, The device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-5.
8. A computer storage medium, characterized in that, The computer storage medium stores a computer program that enables the computer to perform the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Data exporting method and device, computer equipment and storage medium
CN109739928A
Data storage method and device and data retrieval method and device
CN112445889A
Method and a device for storing graph data
CN113656411A