Edge Construction Method and Device for Graph Database

By generating strings and converting them into long integer identifiers, the performance problems caused by unknown identification when there are already edges in the graph database are solved, and more efficient edge update and rewrite performance is achieved.

CN119691234BActive Publication Date: 2025-06-27ZHEJIANG DAHUA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510204576.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

In the graph database, when there is a secondary update or rewrite of an edge, since the edge identification is unknown, the edge needs to be queried and updated or rewrite is extremely affected, which greatly affects the performance of update or rewrite.

Method used

By generating a string based on the attribute values ​​of edges in the graph database, converting characters at preset positions in the string into long integers, obtaining the numerical long integer identifier of the edge, and writing it to the graph database, thereby changing the unknown edge identifier into known.

Benefits of technology

Avoid the need to query the edge before updating or rewriting the edge, which significantly improves the performance of updating or rewriting the edge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691234B_ABST
    Figure CN119691234B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for constructing edges of a graph database. The method for constructing edges of the graph database includes: generating a string based on the attribute values of the edges in the graph database; converting the characters at preset positions in the string into long integers to obtain a numerical long integer identifier of the edge; and writing the numerical long integer identifier as the identifier of the edge into the graph database. The present application can change unknown edge identifiers into known ones through the edge construction method of the present application, and can avoid querying edges before updating or rewriting edges, thereby improving the performance of updating or rewriting edges.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of graph databases, and particularly to a method and apparatus for constructing edges in a graph database. Background Art

[0002] A graph database is a database system for storing and processing graph data. It is mainly used for storing and querying data with relational properties, such as social networks, logistics networks, knowledge graphs, etc. Compared with traditional relational databases, graph databases have better capabilities in processing graph data, and thus have extensive applications in fields such as big data analysis, artificial intelligence, and machine learning.

[0003] In the long-term research and development process, the inventors of this application found that when performing secondary updates or rewrites on existing edges in a graph database, since the edge identifiers are unknown, it is necessary to query the edges before updating or rewriting them, which will extremely affect the performance of the update or rewrite. Summary of the Invention

[0004] This application provides a method and apparatus for constructing edges in a graph database. The unknown edge identifiers can be made known through the edge construction method of this application, which can avoid the action of querying edges before updating or rewriting edges, thereby improving the performance of updating or rewriting edges.

[0005] To achieve the above object, this application provides a method for constructing edges in a graph database. The method includes:

[0006] Generating a string based on the attribute values of edges in the graph database;

[0007] Converting the characters at preset positions in the string into long integers to obtain the numerical long integer identifier of the edge;

[0008] Writing the numerical long integer identifier as the identifier of the edge into the graph database.

[0009] To achieve the above object, this application also provides an electronic device, which includes a processor; the processor is used to execute instructions to implement the steps of the above method.

[0010] To achieve the above object, this application also provides a computer-readable storage medium, which is used to store instructions / program data, and the instructions / program data can be executed to implement the above method.

[0011] The edge construction method of the graph database in this application generates a string from the attribute value of the edge in the graph database; converts the characters at preset positions in the string into long integers to obtain the numerical long integer identifier of the edge; writes the numerical long integer identifier as the identifier of the edge into the graph database. In this way, the unknown edge identifier can be made known through the edge construction method of this application, and the action of querying the edge before updating or rewriting the edge can be avoided, thereby improving the performance of updating or rewriting the edge. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of this application and form a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0013] Figure 1 is a schematic flowchart of an embodiment of the edge construction method of the graph database in this application;

[0014] Figure 2 is a schematic flowchart of an embodiment of the edge construction method of the graph database in this application;

[0015] Figure 3 is a schematic structural diagram of an embodiment of an electronic device in this application;

[0016] Figure 4 is a schematic structural diagram of an embodiment of a computer-readable storage medium in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application. Additionally, unless otherwise specified (e.g., "or alternatively" or "or in an alternative"), the term "or" as used herein refers to a non-exclusive "or" (i.e., "and / or"). And, the various embodiments described herein are not necessarily mutually exclusive, because some embodiments can be combined with one or more other embodiments to form new embodiments.

[0018] As one of the representatives of open-source graph databases, Janusgraph uses a technology called point partitioning to implement the storage and query of graph data. Point partitioning is a commonly used graph database partitioning method. In short, it stores relational data on the two end points respectively. For example, when HBase is used as the storage engine of Janusgraph, for a "communication connection relationship between device A and device B", a "device A" point and a "device B" point will be generated during storage, and the "communication connection" relationship edge will be stored in the cells of the "device A" and "device B" points respectively. Point partitioning can greatly improve the concurrent read and write performance and query efficiency of JanusGraph because it distributes data across multiple nodes and only requires querying or modifying operations on the target nodes instead of the entire graph data.

[0019] For the edge data of a graph database, a common edge is uniquely determined by four elements, which are: the left point identifier (id), the right point id, the edge label (label) name, and the edge id. When updating and rewriting an edge, these four elements are required to update or rewrite the edge. In actual usage scenarios, the edge id of the open-source Janusgraph is automatically generated as a long-type edge id when writing an edge. When performing secondary updates or rewrites on an existing edge in the graph database, since the id of the edge is unknown, it is necessary to query the edge before updating or rewriting, which greatly affects the performance of updating or rewriting.

[0020] Based on this, this application proposes a method for constructing edges in a graph database. This edge construction method generates a string based on the attribute values of the edges in the graph database; converts the characters at preset positions in the string into long integers to obtain the numerical long integer identifier of the edge; writes the numerical long integer identifier as the identifier of the edge into the graph database. In this way, the unknown edge identifier can be made known through the edge construction method of this application, avoiding the action of querying the edge before updating or rewriting the edge, thereby improving the performance of updating or rewriting the edge.

[0021] Specifically, as Figure 1 shown, a specific implementation of the edge construction method for a graph database proposed in this application specifically includes the following steps. It should be noted that the following step numbers are only used for simplified description and are not intended to limit the execution order of the steps. The steps of this implementation can be arbitrarily changed in the execution order without violating the technical idea of this application.

[0022] S101: Generate a string based on the attribute values of the edges in the graph database.

[0023] Optionally, a string can be generated first based on the attribute values of the edges in the graph database, so that the identifier of the edge can be generated based on the string and written into the graph database subsequently.

[0024] Optionally, string calculation can be performed based on at least one attribute value that can uniquely determine the relationship between the two nodes connected by the edge. That is, at least one attribute value (which can be a combination of one or more features) that can uniquely determine the relationship between the two nodes connected by the edge can be used as the input parameter for string calculation. For example, for the identifier of the edge between the object node and the location node to be confirmed in the object leaving event, the attribute values of the time when the object left and the location where it left can be used as the input parameters for string calculation.

[0025] Of course, in other embodiments, string calculation can also be performed based on the general attribute values of the edge. For example, string calculation can be performed according to the type of the edge, that is, the general attribute value of the edge can be used as the input parameter for string calculation.

[0026] In other implementation manners, a string can be generated based on the attribute values of the edge and the attribute values of the two nodes connected by the edge. That is, the attribute values of the edge and the attribute values of the two nodes connected by the edge can be used as the input parameters for string calculation. For example, a string can be calculated based on the identifiers of the two nodes connected by the edge and the general attribute value of the edge. Or a string can be calculated based on the identifiers of the two nodes connected by the edge and the characteristic attribute value of the edge.

[0027] Among them, the generated string can be regarded as the stringEdgeId of the string type. The calculation method of the string is not limited.

[0028] For example, md5 calculation can be performed on the input parameters of string calculation to obtain a string. Exemplarily, for the communication relationship between device nodes, the attribute values of the communication method and the communication time can be used as the input parameters for md5 calculation to obtain a string of 32 randomly hashed characters.

[0029] For another example, hash calculation can be performed on the input parameters of string calculation to obtain a string. Exemplarily, the identifiers of the two nodes connected by the edge and the attribute value of the edge can be used as the input parameters for hash calculation to obtain a string.

[0030] Optionally, step S101 can be generated by the upper business layer.

[0031] S102: Convert the characters at the preset positions in the string into long integers to obtain the numerical long integer identifier of the edge.

[0032] After calculating the string based on step S101, the characters at the preset positions in the string can be converted into long integers to obtain the numerical long integer identifier of the edge.

[0033] Among them, the preset position can refer to all positions in the string, that is, the entire string can be converted into a long integer to obtain the numerical long integer identifier of the edge.

[0034] The preset position can also refer to partial positions in the string.

[0035] For example, the preset position can refer to the first number of characters at the front of the string. In this way, the first number of characters at the front of the string can be converted into a long integer to obtain the numerical long integer identifier of the edge.

[0036] Among them, the first number is set according to the actual situation and is not restricted here. For example, it can be 12 or 24. In a specific example, the first 12 characters in the string can be converted into a long integer to obtain the numerical long integer identifier of the edge. More preferably, the first number is less than or equal to 12 because 36^13 > 2^63 + 1, that is, when it is 12 characters, it is less than the maximum value of the long type, and when it is 13 characters, it will exceed the storage of the long type and overflow. Even more preferably, the first number is equal to 12. Since the custom identifier of the edge is calculated by the business layer based on the attribute value on the edge and the obtained is a randomly hashed string, only taking the first 12 bits, the possibility of conflict is not great in probability statistics; on the other hand, even if the stringIds of two different types of edges conflict, it doesn't matter because the ordinary edge is uniquely determined by the left point id + right point id + edge label + edge id, so it will not be repeated for the edge.

[0037] In this step, as Figure 2 shown, it can be first determined whether the length of the string is greater than the first number. If the length is greater than the first number, the string can be directly truncated to truncate the excess character part, so as to obtain the first number of characters at the front of the string, and then the first number of characters at the front of the string is converted into a long integer; if the length is less than or equal to 12, the string can be directly converted into a long integer.

[0038] Another example is that the preset position can refer to the second number of characters in the middle of the string. In this way, the second number of characters in the middle of the string can be converted into a long integer to obtain the numerical long integer identifier of the edge.

[0039] Another example is that the preset position can refer to the third number of characters at the back of the string. In this way, the third number of characters at the back of the string can be converted into a long integer to obtain the numerical long integer identifier of the edge.

[0040] Another example is that the preset number can refer to all even - numbered characters in the string. In this way, all even - numbered characters in the string can be converted into a long integer to obtain the numerical long integer identifier of the edge.

[0041] Before step S102, it can be determined whether the string conforms to the specification first; if it conforms to the specification, step S102 is executed; if it does not conform to the specification, step S102 may not be performed, that is, the edge identification update fails. Exemplarily, it can be determined whether the string is a character in the base string. Among them, the base string can be set according to the actual situation and is not limited here. For example, it can be "0123456789abcdefghijklmnopqrstuvwxyz".

[0042] Optionally, the edge online access interface of the graph database can be called to concurrently write or update edge data, where the edge data can cover the above-mentioned string; then the edge online access interface of the graph database receives the data and starts to process the data, and then step S102 is executed for each edge data. Moreover, as Figure 2 shown, the rule engine calls the edge online access interface of the graph database to concurrently write edge data to improve the processing efficiency of edge data.

[0043] Among them, there are various methods to convert the character at the preset position in the string into a long integer, which is not limited here.

[0044] For example, a codec for converting characters to long integers can be used to decode the characters at the preset position.

[0045] Among them, each character at the preset position can be traversed to find the position order of the currently traversed character in the base string, where the characters at different positions in the base string are different; based on the position order, the total number of characters in the base string, and the sorting of the currently traversed character among all the characters at the preset position, a first value is calculated; the first value is added to the current numeric long integer identifier to obtain the updated current numeric long integer identifier, where the current numeric long integer identifier starts from a preset value; if the currently traversed character is not the last one among the characters at the preset position, continue to use the next character as the currently traversed character, and return to execute the step of finding the position order of the currently traversed character in the base string and the subsequent steps; if the currently traversed character is the last one among the characters at the preset position, the current numeric long integer identifier is used as the final numeric long integer identifier.

[0046] Among them, the calculation formula for the first value is: ; where s is the first numerical value, m is determined based on the total number of characters. For example, m can be equal to the total number of characters, or m can be equal to the sum of the total number of characters and a specified numerical value. n is the sorting of the currently traversed character from right to left among all the characters at the preset position, and k is the position order. The above-mentioned specified numerical value and preset value can also be set according to the actual situation and are not limited here. For example, they can be 0 or 1, etc.

[0047] In one example, the total number of characters in the above-mentioned base string is 36. In this way, the string is regarded as a number in base 36 (the length of BASE_SYMBOLS), and then converted to decimal to obtain the decimal numerical value, thereby obtaining the numerical long integer identifier. As Figure 2 shown, each character of the characters at the preset position can be traversed, from left to right, regarded as the high-order bit to the low-order bit of base 36; find the position of this character in the base string BASE_SYMBOLS, denoted as pos; the decoded edge value id num = num * 36 + pos; if the currently traversed character is not the last one among the characters at the preset position, continue to traverse the next character; if the currently traversed character is the last one among the characters at the preset position, use id num as the numerical long integer identifier.

[0048] In a specific example, the base string BASE_SYMBOLS = "0123456789abcdefghijklmnopqrstuvwxyz" has 36 characters. Regarding the incoming string from left to right as a base 36 number from high-order bit to low-order bit, the conversion between the string and the long type can be completed. The calculation method is exemplified as follows: when the input string is abc, the numerical long integer identifier is 10 * 36^2 + 11 * 36^1 + 12 * 36^0 = 13368.

[0049] Another example is that a weighted conversion can be performed on the characters at the preset position based on the character positions. Exemplarily, they are weighted according to the positions of the characters in the string, and then combined into a long integer to obtain the numerical long integer identifier.

[0050] Another example is that the characters at the preset position can be converted according to certain attributes of the characters (such as the position in the alphabet, case, etc.) to obtain the numerical long integer identifier.

[0051] S103: Write the numerical long integer identifier as the identifier of the edge into the graph database.

[0052] After calculating the numerical long integer identifier based on step S101, the numerical long integer identifier can be written as the identifier of the edge into the graph database.

[0053] Optionally, a write request can be constructed based on the edge identifier, edge label name, and the identifiers of the two nodes connected by the edge to write edge data into the graph database, so as to load the edge identifier onto the edge, making the constructed edge an edge with a custom representation, thereby realizing the update or rewrite of the edge data.

[0054] In this implementation manner, a string is generated based on the attribute value of the edge in the graph database; the characters at preset positions in the string are converted into long integers to obtain the numerical long integer identifier of the edge; the numerical long integer identifier is written into the graph database as the identifier of the edge. In this way, the unknown edge identifier can be made known through the edge construction method of the present application, and the operation of querying the edge before updating or rewriting the edge can be avoided, thereby improving the performance of updating or rewriting the edge.

[0055] Please refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device 20 includes a processor 22, and the processor 22 is used to execute instructions to implement the above method. For the specific implementation process, please refer to the description of the above embodiments, which will not be repeated here.

[0056] The processor 22 can also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 22 may be an integrated circuit chip with signal processing capabilities. The processor 22 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 22 can also be any conventional processor, etc.

[0057] The electronic device 20 may further include a memory 21 for storing instructions and data required for the operation of the processor 22.

[0058] The processor 22 is used to execute instructions to implement the method provided by any embodiment of the above method of the present application and any non-conflicting combination.

[0059] Among them, the electronic device of the present application can be a confusion control system.

[0060] Please refer to Figure 4 , Figure 4This is a schematic diagram of the structure of the computer-readable storage medium in the embodiments of the present application. The computer-readable storage medium 30 of the embodiments of the present application stores instruction / program data 31, and when the instruction / program data 31 is executed, it implements the methods provided by any embodiment of the edge construction method of the graph database of the present application and any non-conflicting combination. In one embodiment, the instruction / program data 31 may form a program file and be stored in the above storage medium 30 in the form of a software product, so that a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium 30 includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0061] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0062] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0063] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

[0064] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for constructing edges in a graph database, characterized in that: The method comprises: Generate a string based on the attribute value of the edge in the graph database; In response to a character in the string being a character in the basic string, decoding the character at a preset position in the string using a character-to-long integer codec to obtain a numerical long integer identifier of the edge; and writing the numerical long integer identifier as an identifier of the edge into the graph database; The codec for converting characters to long integers is used to decode the characters at the preset positions to obtain the numerical long integer identifier, including: Traversing each character at the preset position to find out the position sequence of the currently traversed character in the basic string, wherein the characters at different positions in the basic string are different; A first value is obtained by calculating based on the position sequence, the total number of characters in the basic character string, and the order of the currently traversed character among all characters at the preset position; the calculation formula of the first value is: , where s is a first value, m is determined based on the total number of characters, n is the order of the currently traversed character from right to left among all characters at the preset position, and k is the position order; Adding the first value and the current numeric long integer identifier to obtain an updated current numeric long integer identifier, wherein the current numeric long integer identifier starts with a preset value; If the currently traversed character is not the last one of the characters at the preset position, continue to use the next character as the currently traversed character, and return to the step of finding the position sequence of the currently traversed character in the basic string; If the currently traversed character is the last one of the characters at the preset position, the current numeric long integer identifier is used as the final numeric long integer identifier.

2. The method according to claim 1, characterized in that The generating a string based on the attribute value of the edge in the graph database includes: The string calculation is performed based on at least one attribute value that can uniquely determine the relationship between two nodes connected by the edge.

3. The method according to claim 2, characterized in that The performing string calculation based on at least one attribute value that can uniquely determine the relationship between two nodes connected by an edge includes: performing md5 calculation on the at least one attribute value that can uniquely determine the relationship between two nodes connected by an edge to obtain the string.

4. The method according to claim 1, characterized in that: The characters at the preset positions are the first first number of characters in the character string.

5. The method according to claim 4, characterized in that The converting the characters at the preset positions in the character string into long integers to obtain the numerical long integer identifier of the edge includes: If the length of the string exceeds the first number, truncating the first number of characters in the string to obtain the character at the preset position; If the length of the character string is less than or equal to the first number, the character string is used as the character at the preset position.

6. An electronic device, characterized in that: The electronic device comprises a processor; the processor is configured to execute instructions to implement the steps of the method according to any one of claims 1-5.

7. A computer-readable storage medium having a program and / or an instruction stored thereon, characterized in that: When the program and / or instruction is executed, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method and a device for storing graph data

    CN113656411A

  • Data duplicate checking method and device, equipment and storage medium

    CN117076509A

  • Database character string comparison method and device, electronic equipment and storage medium

    CN119474476A