Graph data processing method, apparatus and device, and computer readable storage medium

By encoding and storing ByteGraph's attribute index into a key-value database, the problem that ByteGraph's index technology requires two operations is solved, which improves the query efficiency and storage efficiency of graph data, and improves the performance of large-scale graph data processing.

CN120277240APending Publication Date: 2025-07-08TENCENT DIGITAL TIANJIN
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410021738.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

ByteGraph's indexing technology requires two operations to find the relationship between the two points, resulting in extended lookup time and high storage costs, affecting the performance and scalability of large-scale graph data processing.

Method used

By encoding each attribute index, including graph space identification, partition identification, index type, index identification, vertex identification and attribute value, key-value pair information is generated and stored in the key-value database to improve query efficiency.

Benefits of technology

Reduces query time, reduces redundant data storage, improves the overall performance of graph data processing, supports fast query operations and efficient data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277240A_ABST
    Figure CN120277240A_ABST
Patent Text Reader

Abstract

The invention provides a graph data processing method, device and equipment and a computer readable storage medium. The method comprises the following steps: acquiring to-be-processed graph data, and acquiring a plurality of attribute indexes corresponding to the to-be-processed graph data; obtaining index data of each attribute index, and obtaining a vertex identifier and an attribute value corresponding to each attribute index; encoding the index data of each attribute index and the vertex identifier and the attribute value corresponding to each attribute index to obtain key value pair information corresponding to each attribute index; and storing the key value pair information corresponding to each attribute index into a key value database. Through the method and the device, high attribute query efficiency can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and computer-readable storage medium for graph data processing. Background Art

[0002] Graph data is a data structure used to represent and process complex relationships, consisting of nodes and edges, and is widely used in fields such as social network analysis, recommendation systems, knowledge graphs, etc. In graph data, Byte-oriented Graph (ByteGraph) is a storage format with an efficient indexing technology. ByteGraph adopts a byte-based encoding method to compress and store graph data on disk, thereby saving storage space. In addition, ByteGraph also provides an efficient indexing technology by dividing graph data into multiple shards and building indexes on each shard to accelerate query operations. However, there is a problem with the indexing technology of ByteGraph. To find the relationship between two points, two operations are required. First, find the shard where the starting point is located through the index and read the index information on that shard; then, perform another indexing operation on the found shard to find the location of the target point. This dual operation results in an extended search time. Summary of the Invention

[0003] Embodiments of this application provide a method, apparatus, and computer-readable storage medium for graph data processing, which can improve the search efficiency of graph data and save the search time of graph data.

[0004] The technical solution of the embodiments of this application is implemented as follows:

[0005] Embodiments of this application provide a method for graph data processing, the method includes:

[0006] Obtain the graph data to be processed, and obtain multiple attribute indexes corresponding to the graph data to be processed;

[0007] Obtain the index data of each attribute index, and obtain the vertex identifier and attribute value corresponding to each attribute index;

[0008] Perform encoding processing on the index data of each attribute index, and the vertex identifier and attribute value corresponding to each attribute index, to obtain key-value pair information corresponding to each attribute index;

[0009] Store the key-value pair information corresponding to each attribute index into a key-value database.

[0010] Embodiments of this application provide a graph data processing apparatus, including:

[0011] A first acquisition module, configured to acquire graph data to be processed and acquire a plurality of attribute indexes corresponding to the graph data to be processed;

[0012] A second acquisition module, configured to acquire index data of each attribute index and acquire vertex identifiers and attribute values corresponding to each attribute index;

[0013] An encoding module, configured to perform encoding processing on the index data of each attribute index and the vertex identifiers and attribute values corresponding to each attribute index to obtain key-value pair information corresponding to each attribute index;

[0014] A first storage module, configured to store the key-value pair information corresponding to each attribute index into a key-value database.

[0015] An embodiment of the present application provides an electronic device, where the electronic device includes:

[0016] A memory, configured to store computer-executable instructions;

[0017] A processor, configured to implement the method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0018] An embodiment of the present application provides a computer-readable storage medium, storing a computer program or computer-executable instructions, configured to implement the graph data processing method provided by the embodiment of the present application when executed by a processor.

[0019] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, where when the computer program or computer-executable instructions are executed by a processor, the graph data processing method provided by the embodiment of the present application is implemented.

[0020] The embodiment of the present application has the following beneficial effects:

[0021] In the embodiment of the present application, by performing encoding processing on the index data of each attribute index and the vertex identifier and attribute value, key-value pair information corresponding to each attribute index is obtained. In addition, by extracting attribute information such as graph space identifier, partition identifier, index type, and index identifier in the index data and performing encoding processing, the structure and storage method of the index can be optimized. This can further improve the query efficiency, reduce the storage of redundant data, and improve the overall performance. By performing encoding processing on multiple attribute indexes of graph data and storing them in a key-value database, the query operation can be made more efficient. The key-value database has efficient read and write capabilities and fast index query capabilities, which can accelerate the query operation of graph data. Since the key-value pair information corresponding to the attribute index is stored in the key-value database, the nodes or edges that meet the conditions can be quickly found through the key-value pair, and the target data can be located faster, thus saving query time. Description of the Drawings

[0022] Figure 1 is a model diagram of the graph storage topology model with indexes included in ByteGraph;

[0023] Figure 2 is a schematic diagram of the architecture of the graph data processing system provided by an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of the structure of the server provided by an embodiment of the present application;

[0025] Figure 4A is a schematic diagram of an implementation process of the graph data processing method provided by an embodiment of the present application;

[0026] Figure 4B is a schematic diagram of the process for determining key-value pair information corresponding to each attribute index provided by an embodiment of the present application;

[0027] Figure 4C is a schematic diagram of the process for determining key-value pair information corresponding to each attribute index provided by an embodiment of the present application;

[0028] Figure 4D is a schematic diagram of the implementation process for querying based on a target index identifier provided by an embodiment of the present application;

[0029] Figure 4E is a schematic diagram of the implementation process of the data insertion process provided by an embodiment of the present application;

[0030] Figure 4F is a schematic diagram of the implementation process of the data update process provided by an embodiment of the present application;

[0031] Figure 4G is a schematic diagram of the implementation process of the data deletion process provided by an embodiment of the present application;

[0032] Figure 5 is a schematic diagram of the index key encoding provided by an embodiment of the present application;

[0033] Figure 6 is a schematic diagram of the implementation process for point query based on an attribute index provided by an embodiment of the present application;

[0034] Figure 7 is a schematic diagram of the implementation process for point insertion based on an attribute index provided by an embodiment of the present application;

[0035] Figure 8 is a schematic diagram of the implementation process for point update based on an attribute index provided by an embodiment of the present application;

[0036] Figure 9 is a schematic diagram of the implementation process for point deletion based on an attribute index provided by an embodiment of the present application. Detailed implementation manners

[0037] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0038] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0039] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0040] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0041] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0042] 1) Key-Value: A data structure used to store and retrieve data. It consists of a unique key and an associated value. The key is usually used as the unique identifier of the element, while the value represents the actual data stored and read.

[0043] 2) Graph Database: A database that uses a graph structure for semantic queries. It uses nodes, edges, and properties to represent and store data. Nodes represent entities, edges represent the relationships between nodes, and properties add more descriptive information to nodes and edges. Graph databases have efficient semantic query capabilities and can perform complex query operations through the relationships between nodes and edges.

[0044] 3) Index: A data structure in a database system used to improve the retrieval efficiency of data. It is similar to the table of contents of a book. By pre - establishing a sorted data structure, it can quickly locate specific data rows. An index is usually created based on the value of a certain column or multiple columns and can accelerate query operations.

[0045] 4) Composite Index: An index composed of two or more columns. It can sort and retrieve based on the values of multiple columns simultaneously, thus improving query efficiency. A composite index can be used for query operations that meet multiple query conditions.

[0046] 5) B - tree: A balanced multi - way search tree commonly used as an index structure in databases and file systems. It has the property of self - balancing and can efficiently support insert, delete, and search operations. B - trees are often used as the underlying data structure for database indexes and can provide fast retrieval performance.

[0047] In existing graph data processing technologies, ByteGraph uses B - trees to store vertices, edges, and the corresponding indexes of vertices and edges. It supports both global indexes and local indexes. Among them, a local index refers to building an index on the attributes of an edge based on a given starting vertex and edge type, such as an index based on the age of a user; a global index refers to finding the identifiers (IDs) of all vertices with a specific attribute value throughout the graph based on an attribute value. To maintain the consistency between data and indexes, ByteGraph uses distributed transaction capabilities for processing. By abstracting each index as a virtual vertex, the ID of this vertex is the specific attribute value, which can be a hash value for types such as strings and double - precision floating - point numbers, and the type of the vertex is a built - in type, and the type values of each attribute are the same. Each time a vertex is written, ByteGraph will additionally create an edge. The starting point of the edge is the virtual attribute - value vertex, and the ending point of the edge is the vertex written this time. The type of the edge is a custom internal type and can be associated with the type of the virtual index - value vertex. In this way, when querying a vertex using an index, only the target vertex that is one - hop away from the index vertex needs to be found. See Figure 1 , Figure 1 is the model diagram of the graph storage topology model with indexes in ByteGraph.

[0048] However, ByteGraph also has obvious drawbacks. The query efficiency of ByteGraph has obvious drawbacks. For example, when looking for points with city A, it needs to be divided into two steps: first, find the index point with point type 100001 and id A, and then, starting from this index point, find the one-hop target node. This method requires two operations to locate the point with id A, so the query efficiency is relatively low. At the same time, the storage cost of ByteGraph is also relatively high because a virtual vertex and an edge need to be created for each index, and when the data volume is large, this method will occupy a large amount of storage space. These drawbacks will affect the performance and scalability of ByteGraph and are not conducive to its application in large-scale graph data processing scenarios.

[0049] To solve the problems existing in the prior art, embodiments of the present application perform encoding processing on each attribute index, including graph space identifier, partition identifier, index type, index identifier, and vertex identifier and attribute value. In this way, the key-value pair information corresponding to each attribute index can be obtained. During query, by quickly locating the nodes or edges that meet the conditions, the query time can be effectively reduced. In addition, encoding multiple attribute indexes of the graph data and storing them in a key-value database can make the query operation more efficient, accelerate the query operation on the graph data, further improve the query efficiency, reduce the storage of redundant data, and improve the overall performance.

[0050] In embodiments of the present application, data related to user information, user voice data, etc. are involved. When embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0051] Embodiments of the present application provide a graph data processing method, apparatus, device, computer-readable storage medium, and computer program product, which can improve the search efficiency of graph data and save the search time of graph data. The following describes the exemplary applications of the electronic device provided by embodiments of the present application. The device provided by embodiments of the present application can be implemented as various types of user terminals such as laptops, tablets, desktop computers, set-top boxes, mobile devices (such as mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), smartphones, smart speakers, smart watches, smart TVs, in-vehicle terminals, etc., or can also be implemented as a server. The following will describe the exemplary applications when the device is implemented as a server.

[0052] See Figure 2 , Figure 2It is a schematic architecture diagram of the graph data processing system 100 provided by an embodiment of the present application. To support an exemplary application, the terminal 200 is connected to the server 400 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of both. The graph data processing system 100 may further include a key-value database 500 for obtaining the graph data to be processed. The key-value database 500 can be independent of the server 400 or located inside the server 400. In Figure 2 the case where the key-value database 500 is independent of the server 400 is taken as an example for illustration.

[0053] The server 400 receives the graph data processing request sent by the terminal 200. The graph data processing request includes the graph data to be processed. The server 400 obtains multiple attribute indexes corresponding to the graph data to be processed by obtaining the graph data to be processed, and then obtains the index data of each attribute index through the multiple attribute indexes, and obtains the vertex identifier and attribute value corresponding to each attribute index through the graph data. After obtaining the index data, vertex identifier, and attribute value of each attribute index, the server 400 encodes the index data of each attribute index, as well as the vertex identifier and attribute value corresponding to each attribute index, to obtain the key-value pair information corresponding to each attribute index. After determining the key-value pair information corresponding to each attribute index, the server 400 stores the key-value pair information corresponding to each attribute index in the key-value database 500 for subsequent use.

[0054] In some embodiments, the server 400 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 200 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.

[0055] See Figure 3 , Figure 3 It is a schematic structural diagram of the server 400 provided by an embodiment of the present application. Figure 3The server 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 3 all kinds of buses are labeled as the bus system 440.

[0056] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0057] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.

[0058] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located far from the processor 410.

[0059] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0060] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0061] An operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0062] A network communication module 452 for reaching other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.;

[0063] A presentation module 453 for enabling the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (such as a display screen, a speaker, etc.);

[0064] An input processing module 454 for detecting and translating one or more user inputs or interactions from one of one or more input devices 432.

[0065] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 3 Shown is a graph data processing device 455 stored in the memory 450, which can be software in the form of programs and plugins, etc., including the following software modules: a first acquisition module 4551, a second acquisition module 4552, an encoding module 4553, and a first storage module 4554. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.

[0066] In other embodiments, the device provided by the embodiments of the present application can be implemented in hardware. As an example, the device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the graph data processing method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can employ one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Programmable Logic Devices (PLDs), Complex Programmable Logic Devices (CPLDs), Field-Programmable Gate Arrays (FPGAs), or other electronic components.

[0067] Next, the graph data processing method provided by the embodiments of the present application will be described. As mentioned above, the electronic device implementing the graph data processing method of the embodiments of the present application can be a server. Therefore, the execution subject of each step will not be repeated hereinafter. Refer to Figure 4A , Figure 4A which is a schematic diagram of an implementation process of the graph data processing method provided by the embodiments of the present application, and will be described in conjunction with the steps shown in Figure 4A . The subject of the steps is the server. Figure 4A

[0068] In step 101, the graph data to be processed is obtained, and a plurality of attribute indexes corresponding to the graph data to be processed are obtained.

[0069] In some embodiments, the graph data to be processed can be imported from an external data source (such as a database, a file, etc.). For example, in social network analysis, relationship data can be obtained from the API interface of a social media platform, and then the relationship data can be converted into graph data. For example, the entities in the relationship data can be mapped to the nodes in the graph, and the relationships in the relationship data can be mapped to the edges in the graph. In this way, the relationships between entities can be more intuitively displayed in the graph data. The graph data to be processed can also be obtained from other systems or applications. For example, in a recommendation system, the relationships between products and users can be extracted from the user purchase records of a shopping website to construct graph data.

[0070] The graph data to be processed can include node information and relationship information between nodes. Among them, the relationship information between nodes is also edge information. Nodes represent entities, which can be people, objects, concepts, etc.; edges represent the relationships between entities, such as friendship, cooperation, inclusion, etc. The graph data to be processed can also include attribute information of nodes and attribute information of edges. For example, the attribute information of nodes can include age, gender, hobbies, etc. The following examples will all be given with the attribute information of nodes.

[0071] An attribute index is a way to index and mark the attributes in the graph data. The purpose of the attribute index is to improve the access and query efficiency of the data. Generally, the attribute index maps the attribute values to the corresponding graph nodes or edges so that nodes or edges containing specific attribute values can be quickly retrieved.

[0072] The attribute index corresponding to the graph data to be processed contains the attribute information of nodes and edges. The attribute information may include the attribute name, attribute type, attribute value, and the identifier of the corresponding node or edge, etc. Among them, the attribute name refers to the name of the attribute in the graph data, the attribute type refers to the data type of the attribute, the attribute value refers to the specific numerical value or string corresponding to the attribute in the node or edge, and the identifier of the node or edge is the number or ID used to uniquely identify each node or edge. For example, in a social network, attributes such as name, age, and gender can be added to each user node, and attributes such as relationship type and relationship establishment time can be added to the friendship edges. The relationship establishment time is the time when the association relationship is established between two objects, such as the time of becoming friends. The attribute index corresponding to the graph data to be processed can search and filter the graph data based on these attribute values. For example, the user groups within a certain age range can be found through the age attribute index, or the user groups with similar interests can be found through the interest attribute index.

[0073] To obtain multiple attribute indexes corresponding to the graph data to be processed, the server preprocesses the graph data. This may involve traversing the graph structure and extracting the attribute information of each node and edge, and then storing it in the index data structure. The specific form and implementation method of the attribute index vary according to the application scenario, and various data structures and algorithms can be used to implement the index function, such as hash tables, B-trees, inverted indexes, etc.

[0074] In step 102, obtain the index data of each attribute index, and obtain the vertex identifier and attribute value corresponding to each attribute index.

[0075] In some embodiments, the index data includes the graph space identifier, partition identifier, index type, and index identifier. In a graph database, the graph space identifier is used to identify an independent graph space, which can be understood as an independent graph database instance. Each graph space can contain multiple partitions, and the partition identifier is used to identify different data shards or data sets. The index type refers to the type of index used to accelerate graph database queries, such as index based on attribute values, index based on points, index based on edges, etc. The index identifier is used to uniquely identify a specific index. The vertex identifier is used to uniquely identify the vertices in the graph database, and the corresponding vertices can be obtained through it. The attribute value is the value of the attribute stored on the point or edge. The attribute value can be different types of data such as strings, integers, and floating-point numbers. The attribute value of an attribute index contains at least one value. For example, the attribute in the attribute index of a vertex can be the "name" attribute, and the attribute value of the "name" attribute can be "Zhang San". The attributes in the attribute index of a vertex can also be the "name" and "student number" attributes, and the attribute values of the "name" and "student number" attributes can be "Zhang San; 0012".

[0076] In some embodiments, the graph space identifier associated with the attribute index is obtained by querying the graph data of the graph database, such as system tables or views. This graph data is typically stored in specific system tables and contains information about the graph space. For each graph space, the partitioned graph data related to that graph space is queried to obtain the partition identifier. The partitioned graph data may include information such as the name and location of the partition. When obtaining the graph data, the graph space identifier and the partition identifier are automatically generated. Information related to the attribute index, including the index type, index identifier, and the vertex identifier and attribute value corresponding to the attribute index, is obtained from the index data corresponding to the graph data.

[0077] In step 103, the index data of each attribute index, as well as the vertex identifier and attribute value corresponding to each attribute index, are encoded to obtain the key-value pair information corresponding to each attribute index.

[0078] In some embodiments, refer to Figure 4B , Figure 4B which is a schematic flowchart of the process for determining the key-value pair information corresponding to each attribute index provided by the embodiments of this application. Figure 4A The step 103 shown can be implemented by steps 1031 to 1032 as shown in Figure 4B which will be specifically described below in conjunction with Figure 4B .

[0079] Step 1031: Determine the meta-information of the attribute index based on the vertex identifier and attribute value corresponding to the attribute index.

[0080] In the graph database, the meta-information is descriptive information about the data. The meta-information is used to describe the relevant information of specific attributes of the attribute index, such as length, data type, etc. In some embodiments, the meta-information content is length information, indicating the number of bytes occupied by the attribute value. The length of the meta-information can be determined based on the vertex identifier and the attribute value.

[0081] In some embodiments, refer to Figure 4C , Figure 4C which is a schematic flowchart of the process for determining the key-value pair information corresponding to each attribute index provided by the embodiments of this application. Figure 4B The step 1031 shown can be implemented by steps 311 to 313 as shown in Figure 4C which will be specifically described below in conjunction with Figure 4C .

[0082] Step 311: Obtain the first data length of the vertex identifier of the attribute index and obtain the second data length of the attribute value of the attribute index.

[0083] In some embodiments, the vertex identifier and attribute value data of the attribute index are obtained by querying relevant graph data sources. The identifier data stored in the attribute index is read, and the first data length of the vertex identifier is calculated. This length may refer to bytes, bits, or qubits, depending on the data storage method and encoding rules. The attribute value data stored in the attribute index is read, and the second data length of the length attribute value is calculated. Similar to the vertex identifier, the length of the attribute value may also be represented in different units.

[0084] Step 312: Determine the third data length occupied by storing the first data length and the second data length.

[0085] In some embodiments, the number of bytes required to store the first data length and the second data length is calculated and added together to obtain the third data length. The first data length and the second data length can be converted into a specific binary representation form, and the number of bytes they occupy is calculated. Then, based on this length information, it is calculated how many bits or bytes of space are needed to store these length data. For example, assume that the first data length occupies 16 bits and the second data length occupies 8 bits, then the third data length occupies 24 bits. Therefore, the content in the meta-information is [16, 8, 24].

[0086] Step 313: Determine the first data length, the second data length, and the third data length as the meta-information of the attribute index.

[0087] In some embodiments, the first data length, the second data length, and the third data length are stored as the meta-information of the attribute index. These meta-information are usually stored together with the actual attribute index data so that the attribute index can be accurately parsed and used when needed. The calculated length information is saved in a specific data structure, such as a metadata table, an index file header, etc. When other systems or objects need to access the attribute index, they can quickly parse the index data according to the meta-information and correctly process the data lengths of the vertex identifier and the attribute value.

[0088] Step 1032: Encode the index data of the attribute index, as well as the vertex identifier, attribute value, and meta-information corresponding to the attribute index in a preset format to obtain the key-value pair information corresponding to the attribute index.

[0089] In some embodiments, the index data of the attribute index, the corresponding vertex identifier, the attribute value, and the meta-information are encoded to obtain the key-value pair information corresponding to the attribute index. The key information corresponding to the attribute index can be: [graph space identifier, partition identifier, index type, index identifier, attribute value, vertex identifier, meta-information]. The preset encoding format can be composed of the graph space identifier, partition identifier, index type, index identifier, and the vertex identifier, attribute value, and meta-information corresponding to the attribute index arranged in sequence. Among them, the content order in the preset encoding format can be modified, but the content in the encoding cannot be changed. The value information corresponding to the attribute index is a null value. The key-value pair information is composed of the key information and the value information.

[0090] In step 104, the key-value pair information corresponding to each attribute index is stored in the key-value database.

[0091] In some embodiments, the key-value pairs are written into the key-value database. This process involves operations such as data serialization, compression, or encryption to ensure data integrity and security. After the data is successfully written into the data block, a successful response is returned to the terminal. During storage, various methods can be used to manage the key-value database. One common method is to use a hash table, which maps keys to hash buckets to achieve efficient lookup and insertion operations. In addition, data structures such as B-trees and LSM-trees are also often used for the storage and management of key-value databases. Furthermore, techniques such as backup, replication, and sharding can be used to improve data reliability and scalability. Backup can create data copies on different nodes or disks to prevent single-point failures. Replication can synchronize data between multiple nodes, increasing system throughput and fault tolerance. Sharding can divide data into multiple segments and distribute them on different nodes, thereby achieving horizontal expansion and load balancing.

[0092] A key-value database is a non-relational database that uses simple key-value pairs to store and retrieve data. Each key is unique and is associated with a specific value. This type of database is suitable for scenarios that require fast data reading and writing because they can perform these operations with a constant time complexity. The structure of a key-value database is similar to a dictionary or hash table, where the key is a hash value calculated by a hash function, and the value can be any type of data, such as a string, integer, list, or object. This flexibility makes key-value databases very suitable for storing large amounts of semi-structured data, such as configuration information, session state, and log data.

[0093] In some embodiments, after step 104, it is also possible to Figure 4D perform queries based on the target index identifier through steps 201 to 207 as shown to obtain query results. Figure 4DThe following is a schematic diagram of the implementation process for querying based on a target index identifier provided by an embodiment of this application. The following will be specifically described in conjunction with Figure 4D Specifically explain.

[0094] Step 201: Obtain a query request.

[0095] In some embodiments, obtaining a query request can be through network communication. The query request includes the attribute to be queried and the target value of the attribute to be queried. When the terminal sends a query request to the server, the server will receive and parse the query request. Among them, the attribute to be queried refers to a specific attribute of the object or information that is desired to be queried. For example, in a commodity database, the attributes to be queried can be commodity names, prices, etc. The target value of the attribute to be queried can be a specific value, a range, a keyword, etc. For example, for the price attribute, the target value can be less than 100 yuan, more than 500 yuan, or simply equal to 120 yuan. In the query request, the terminal will pass the attribute to be queried and the target value of the attribute to be queried as parameters to the server. The server performs corresponding query operations based on these parameters and returns the query results that meet the conditions to the terminal.

[0096] Step 202: Determine whether there is a corresponding target index identifier for the attribute to be queried.

[0097] In some embodiments, when receiving a query request, the query statement will be analyzed first to determine the attribute to be queried and judge whether there is a corresponding target index identifier for these attributes.

[0098] In some embodiments, when there is a target index identifier corresponding to the attribute to be queried in the key-value database, go to step 203; when there is no target index identifier corresponding to the attribute to be queried in the key-value database, go to step 206.

[0099] Step 203: Determine the target key-value pair information including the target index identifier from the key-value database.

[0100] In some embodiments, the attribute value included in the target key-value pair information satisfies the matching condition with the target value. The attribute value included in the target key-value pair information satisfying the matching condition can be that the attribute value included in the target key-value pair information is exactly the same as the target value, or that the attribute value included in the target key-value pair information is partially the same as the target value. The target key-value pair information usually can also include multiple attributes and their attribute values. After determining the target index identifier, candidate key-value pair information containing the target index identifier is retrieved through the graph database. Then, the target key-value pair information whose attribute value satisfies the matching condition with the target value is filtered out from the candidate key-value pair information.

[0101] Step 204: Determine the target vertex information corresponding to the target vertex identifier included in the target key-value pair information as the query result.

[0102] In some embodiments, after determining the target key-value pair information, obtain the target vertex identifier from the target key-value pair information, and the vertex information corresponding to the vertex identifier is obtained from the graph data. Then, obtain the target vertex information corresponding to the target vertex identifier from the graph data. The target vertex information may include information such as the attributes of the point and the associated edges.

[0103] For example, assume there is a key-value database storing information about products, including product names, prices, and types. If you want to query the information of all products of type electronic products, you can use "type" as the attribute to be queried and locate the target key-value pair containing the target index identifier in the key-value database through its corresponding target index identifier. Then, traverse the located target key-value pairs and check whether their type attribute values match the query condition "electronic products". Only the vertex information corresponding to the vertex identifier included in the target key-value pair information with the type attribute value of "electronic products" will be returned as the query result.

[0104] Step 205: Output the query result.

[0105] In some embodiments, the query result can be displayed and output in various ways. For example, on the display device of the server, the query result can be displayed in the form of a list, or through data visualization, such as in the form of charts, statistical graphs, etc. In addition, the query result can also be sent to the terminal. The specific form of outputting the query result depends on the application scenario and user requirements.

[0106] For example, assume that there is a key-value database storing information about some products. Each product has a unique identifier (ID), as well as attributes such as name and price. A query request is received, asking to find information about all fruit products whose names start with "A". First, check whether there is an identifier indexed by the name. Assume there is an index named "name_index". Through the "name_index" index, find all key-value pair information containing the target index identifier in the key-value database. According to the matching condition (the name starts with "A"), filter out the key-value pair information that meets the condition. Take some key-value pair information as an example, such as {"Apple", 2.99}, {"Avocado", 1.99}, etc. Among the filtered key-value pair information, the value of the name attribute meets the matching condition. Find the corresponding product information according to the target vertex identifier (product ID) in the key-value pair information, such as {"Apple", 2.99}. Finally, take the information whose name (attribute) contains "A", such as {"Apple", 2.99}, {"Avocado", 1.99}, etc., as the query result and output it to the terminal.

[0107] Step 206: Traverse the graph data to determine the target vertex information.

[0108] In some embodiments, when there is no target index identifier corresponding to the attribute to be queried in the key-value database, it is necessary to traverse all the graph data to determine the target vertex information, and the attribute value of the attribute to be queried in the target vertex information is the target value. Traversing the graph data means retrieving and searching for nodes and edges in the graph according to a certain method through the storage structure of the graph database. Algorithms such as breadth-first search (BFS) or depth-first search (DFS) can be used to traverse the graph data. These algorithms can start from a starting node, gradually explore the connected nodes, and further explore their connected nodes, and so on. During the traversal process, it is judged whether the attributes of each vertex meet the query conditions. If the attribute value of the attribute to be queried of a certain vertex is the target value, then the vertex information of this vertex is determined as the target vertex information. Only after traversing all the graph data in the graph database will the found target vertex information be returned to the terminal to complete the query operation.

[0109] Step 207: Determine the target vertex information as the query result and output the query result.

[0110] In some embodiments, before processing graph data, it is necessary to establish an attribute index. The traditional method is to establish an attribute index separately. For example, in a school, the attributes of students can be separately established with a gender attribute index, an age attribute index, a grade index, etc. However, the embodiments of the present application propose a new method that can establish a combined attribute index according to requirements. For example, in a school, the gender and name of students rarely change, so a combined attribute index of gender and name can be established, such as [gender, name]. To encode the attribute index, it is necessary to use a preset format to encode the information related to the combined attribute index. This information includes the graph space identifier, the partition identifier, the index type, the index identifier, and the vertex identifier, attribute value, and meta-information corresponding to the combined attribute index. Through the encoding process, the key-value pair information corresponding to the combined attribute index can be obtained. The key is the information related to the combined attribute index, and the value is a null value. Then, the key-value pair information corresponding to the combined attribute index is stored in the key-value database.

[0111] When performing a query, this index can be used to query the students corresponding to the attributes starting with gender. In addition, the index can also be used to query the students corresponding to the attributes starting with [gender, name]. By establishing a combined attribute index, the storage space required for separately establishing attribute indexes can be reduced, and the query efficiency can be improved.

[0112] In the above steps 201 to 207, by obtaining a query request, including the attribute to be queried and the target value of the attribute to be queried, the target index identifier corresponding to the attribute to be queried is determined in the key-value database. If the target index identifier exists in the key-value database, the attribute value in the target key-value pair information is matched with the target value, and the target vertex information corresponding to the target vertex identifier in the target key-value pair information is determined as the query result and output, so that the query effect is efficient and accurate, and the target vertex information that meets the conditions can be quickly located. However, if the target index identifier corresponding to the attribute to be queried does not exist in the key-value database, it is necessary to traverse the graph data to determine the target vertex information. During the traversal process, for each target vertex information, it is checked whether the attribute value of the attribute to be queried matches the target value. When the matching target vertex information is found, the target vertex information is determined as the query result and output, so that the target vertex information that meets the conditions can be guaranteed to be found in the graph data. Therefore, by obtaining the corresponding target vertex information from the key-value database or the graph data according to the attribute to be queried and the target value and outputting it as the query result, the query request can be effectively implemented.

[0113] In some embodiments, after completing step 104, the data insertion process can also be implemented through Figure 4E the steps 301 to 309 shown, Figure 4E which is a schematic flowchart of the implementation process of the data insertion process provided by the embodiments of the present application. The following combines Figure 4EDetailed description.

[0114] Step 301: Obtain a data insertion request.

[0115] Among them, the data insertion request includes at least one object attribute of the object to be inserted and the attribute value of the object attribute.

[0116] In some embodiments, when a request to insert data is received, the insertion request is parsed and the object to be inserted and its attributes and attribute values are extracted. These attributes and attribute values describe the characteristics and characteristic values of the object to be inserted. The data insertion request includes one or more attributes of the object to be inserted, and each attribute corresponds to an attribute value.

[0117] Step 302: Insert the object to be inserted into the graph data and obtain the vertex identifier of the object to be inserted.

[0118] In some embodiments, the object to be inserted is inserted into the graph database in the form of a vertex, and after successful insertion, the unique identifier of the object to be inserted, that is, the vertex identifier, is obtained for subsequent access and management of the object to be inserted. During this period, the corresponding API or an insertion language is called to represent the object to be inserted as a vertex and perform the insertion operation. After the insertion operation is successful, the vertex identifier of the object to be inserted is obtained from the return result.

[0119] Step 303: Determine whether to establish an attribute index for the object to be inserted.

[0120] Before inserting the key-value pair information of the object to be inserted into the key-value database, it is determined whether an index needs to be established for the attributes of the object. When it is necessary to establish an attribute index for the object to be inserted, go to step 304. When it is not necessary to establish an attribute index for at least one object attribute, go to step 308.

[0121] Step 304: Establish an attribute index for at least one object attribute.

[0122] In some embodiments, the corresponding API or a query language can be called to create an attribute index. The index establishment process associates the attribute value with the corresponding vertex or edge. In addition to the index of at least one attribute, other attributes can be optionally selected to be grouped together to establish an index. For example, an attribute index that includes both the name and gender attributes can be established.

[0123] Step 305: Obtain the index data of the attribute index of at least one object attribute.

[0124] In some embodiments, the index data of the attribute index of the object attribute can be obtained by accessing the metadata or system tables of the graph database system. The index data consists of a graph space identifier, a partition identifier, an index type, and an index identifier. In the metadata, the graph space containing the attribute index to be queried can be found, and the corresponding graph space identifier can be obtained. If the graph space is divided into multiple partitions, the corresponding partition identifier can be obtained according to the relevant information of the partition where the index is located. In addition, the metadata can be further queried to obtain the type information of the attribute to be queried. Finally, the attribute index identifier to be queried can be obtained from the metadata for subsequent query and management operations.

[0125] Step 306: Generate the key-value pair information to be inserted based on the index data of the attribute index of at least one object attribute, the vertex identifier of the object to be inserted, and the attribute value of the point attribute.

[0126] Step 307: Store the key-value pair information to be inserted into the key-value database.

[0127] In some embodiments, according to the graph space identifier, partition identifier, index type, index identifier, the vertex identifier of the object to be inserted, and the attribute value of the point attribute, an appropriate coding algorithm and data structure are used to determine the key in the key-value pair, and the value in the key-value pair is empty. Finally, the key-value pair information to be inserted is stored in the key-value database.

[0128] Step 308: Generate the key-value pair information to be inserted based on the vertex identifier corresponding to the object to be inserted, at least one object attribute of the object to be inserted, and the attribute value of the object attribute.

[0129] Step 309: Store the key-value pair information to be inserted into the key-value database.

[0130] In some embodiments, when it is not necessary to establish an attribute index for the object to be inserted, the attribute information of the object to be inserted needs to be added to the graph database, and the key-value pair information to be inserted is generated based on the graph space identifier, partition identifier, vertex identifier of the object to be inserted, and the attribute value of the point attribute. That is, the index type and index identifier are not included in the key-value pair information to be inserted. Therefore, after being converted into the form of a key-value pair, the content of the index type and index identifier is empty. Finally, the key-value pair information after conversion and without the content of the index type and index identifier is stored in the key-value database.

[0131] In the above steps 301 to 309, through the insertion method of the embodiments of the present application, an efficient data insertion operation can be achieved. First, obtain a data insertion request including at least one object attribute of the object to be inserted and the corresponding attribute value. Then, insert the object to be inserted into the graph data and obtain its corresponding vertex identifier. When it is necessary to establish an attribute index for the object to be inserted, at least one object attribute's attribute index can be created. By obtaining information such as the graph space identifier, partition identifier, index type, and index identifier of the attribute index, as well as the vertex identifier of the object to be inserted and the attribute value of the point attribute, the key-value pair information to be inserted can be generated. The generated key-value pair information can be stored in the key-value database for subsequent query and retrieval operations. If it is not necessary to establish an attribute index, the key-value pair information to be inserted can be generated based on the vertex identifier corresponding to the object to be inserted, at least one object attribute of the object to be inserted, and the attribute value of the object attribute, and stored in the key-value database. Therefore, through the data insertion method of the embodiments of the present application, after inserting the object to be inserted into the graph data, the key-value pair information of the object to be inserted will also be generated, and the key-value pair information corresponding to the object to be inserted will be added to the key-value database, thereby ensuring the consistency between the key-value database and the graph data.

[0132] In some embodiments, after completing step 104, it can also be updated through Figure 4F the steps 401 to 409 shown below, Figure 4F which is a schematic flowchart of the implementation process of the data update process provided by the embodiments of the present application. The following will be specifically described in conjunction with Figure 4F this.

[0133] Step 401: Obtain a data update request.

[0134] In some embodiments, the terminal sends a data update request to the server to update the data in the database. This data update request needs to include at least one object attribute of the object to be updated and the corresponding attribute value of the object attribute. After receiving the data update request, the server will check whether the data update request contains the necessary information. Among them, the information in the data update request should clearly specify the object to be updated, usually distinguished by a unique identifier or other attributes, and clearly specify at least one attribute of the object to be updated, such as name, age, address, etc. It is also necessary to obtain the new attribute value of the attribute to be updated. After determining that the data update request contains the necessary information, the object to be updated is determined according to the information carried in the data update request, and the specified attribute value of the object to be updated is updated to the new attribute value.

[0135] Step 402: Update the graph data based on at least one object attribute of the object to be updated and the attribute value of the object attribute to obtain the updated graph data.

[0136] In some embodiments, according to the object attributes and attribute values provided in the update request, the corresponding data in the graph database is updated, which may involve updating the attributes of points, updating the attributes of edges, or adjusting the graph structure. If it is to update the attributes of points, the points to be updated will be found and the attribute values of the specified object attributes therein will be updated. If it is to update the attributes of edges, the edges to be updated will be found and the attribute values of the specified object attributes therein will be updated. Finally, after all the update operations are completed, the updated graph data is obtained, and these updates may include data updates corresponding to at least one object attribute of the object to be updated and the attribute values of the object attributes.

[0137] Step 403: Determine whether the object to be updated has an attribute index.

[0138] Before updating the key-value pair information of the object to be updated in the key-value database, it is determined whether the object has an attribute index. When the object to be updated has an attribute index, step 404 is entered; when the object to be updated does not have an attribute index, step 407 is entered.

[0139] Step 404: Obtain the vertex identifier of the object to be updated, and obtain the index data of the attribute index of at least one object attribute.

[0140] In some embodiments, when the object to be updated has an attribute index, the vertex identifier of the object to be updated will be obtained, and further the index data of the attribute index of at least one object attribute will be obtained. The index data includes the graph space identifier, partition identifier, index type, and index identifier. By obtaining the relevant information of the attribute index, the index structure of the updated object can be located, so as to perform operations on the attribute index. For example, the query language or API of the graph database will be used to access the database, and the required vertex identifier, attribute index, graph space identifier, partition identifier, index type, and index identifier will be obtained by traversing the index and searching for relevant information.

[0141] Step 405: Generate new key-value pair information of the object to be updated based on the index data of the attribute index of at least one object attribute, the vertex identifier of the object to be inserted, and the attribute values of the point attributes.

[0142] In some embodiments, when the object to be updated has an attribute index, new key-value pair information of the object to be updated can be generated based on the index data of the attribute index of at least one object attribute. The new key-value pair information consists of the updated graph space identifier, partition identifier, index type, index identifier, the vertex identifier of the object to be inserted, and the attribute values of the point attributes.

[0143] Step 406: Update the initial key-value pair information of the object to be updated in the key-value database to the new key-value pair information.

[0144] In some embodiments, in a key-value database, new key-value pair information is used to replace the original key-value pair information. A connection to the key-value database needs to be established for data reading and writing operations. Then, the object to be updated is searched in the key-value database according to specified conditions, and the specified conditions can be using a unique identifier or other query parameters to determine the location of the object. After obtaining the initial key-value pair information, the current key-value pair information of the object is retrieved from the key-value database, which includes the key and the associated value. Then, the original key-value pair information is overwritten with the new key-value pair information, which can be completed by directly replacing the existing key and value, or using the corresponding database commands. Finally, the updated key-value pair information is stored back in the key-value database.

[0145] Step 407: Obtain the vertex identifier of the object to be updated.

[0146] In some embodiments, when the object to be updated has no attribute index, it cannot be directly queried through the index structure. Therefore, the vertex identifier of the object to be updated needs to be obtained. One method is to retrieve the vertex identifier of the object to be updated by using a graph database query language. For example, when using Java as the graph database query language to query in the system, a query program can be written to obtain the vertex identifier of the object to be updated. One method is to directly access the vertex identifier of the object to be updated by using the API provided by the server. For example, specific API functions can be used to obtain the identifier of the object.

[0147] Step 408: Generate new key-value pair information based on the vertex identifier corresponding to the object to be updated, at least one object attribute of the object to be updated, and the attribute value of the object attribute.

[0148] Step 409: Update the initial key-value pair information of the object to be updated in the key-value database to the new key-value pair information.

[0149] In some embodiments, the updated attributes of the object to be updated are entered into the graph database. The relevant information of the object to be updated can be extracted and converted into the form of key-value pairs. The relevant information of the object to be updated includes the graph space identifier, the partition identifier, the vertex identifier of the object to be updated, and the attribute value of the point attribute. Since the object to be updated has no attribute index, the key-value pair does not contain the index type and the index identifier. Therefore, after being converted into the form of key-value pairs, the index type and the index identifier content are empty. Next, the initial key-value pair information of the object to be updated in the key-value database is replaced with the new key-value pair information. The initial key-value pair of the object to be updated is found according to the vertex identifier and updated to the new key-value pair information. Finally, the key-value pair information that has been converted and does not contain the index type and the index identifier content is stored in the key-value database.

[0150] In some embodiments, in a data update request, the object to be updated can be the attribute value of a single point, the attribute value of a single edge, or all the attribute indexes of points or edges. For example, in class A of a school, if it is necessary to update the age attribute value of student B, and the attribute index and its attribute value are "age: 18", then in the update request, the age of student B is 19. After the update, the age attribute index and its attribute value of student B are "age: 19", and then it is written into the key-value database for storage. If it is necessary to update the grade attribute value of all students in class A, and the attribute index and its attribute value are "grade: sophomore", then in the update request, the grade of all students in class A is senior. After the update, the grade attribute indexes and their attribute values of all students are updated to "grade: senior", and then it is encoded and written into the key-value database for storage. Additionally, if it is necessary to update the relationship attribute between student C and student D to deskmate, and the original edge relationship attribute index and its attribute value are "relationship: classmate", then in the update request, the relationship attribute between student C and student D is deskmate. After the update, the edge relationship attribute index and its attribute value between student C and student D are "relationship: deskmate", and then it is encoded and written into the key-value database for storage.

[0151] In the above steps 401 to 409, by obtaining a data update request and updating the graph data based on the attributes and attribute values of the object to be updated, the embodiments of the present application can implement the update operation of objects in the graph database. When the object to be updated has an attribute index, by obtaining information such as vertex identifier, graph space identifier, partition identifier, index type, and index identifier, new key-value pair information of the object to be updated can be generated, and the original key-value pair information in the key-value database is replaced with the new key-value pair information. When the object to be updated has no attribute index, new key-value pair information can also be generated by obtaining the vertex identifier and the attributes and attribute values of the object to be updated, and the original key-value pair information in the key-value database is updated. Therefore, it is convenient to update the objects in the graph database by specifying attributes and attribute values, improving the efficiency and accuracy of data update.

[0152] In some embodiments, after completing step 104, the deletion process can also be carried out through Figure 4G the steps 501 to 505 shown below, Figure 4G which is a schematic flowchart of the implementation process of the data deletion process provided by the embodiments of the present application. The following will be specifically described in conjunction with Figure 4G this.

[0153] Step 501: Obtain a data deletion request.

[0154] In some embodiments, a data deletion request sent by a terminal is received. This request can be transmitted in various ways. For example, the terminal sends the data deletion request to the server via a network. After the server receives the data deletion request, it parses the content of the request. The data deletion request may include: the unique identifier of the object to be deleted, which is used to locate the position of the object to be deleted in the key-value database; the object graph data information of the object to be deleted, such as the name of the object graph where the object to be deleted is located, node information, edge information, etc.; and other data information related to the object to be deleted, such as object attributes, association relationships, etc.

[0155] Step 502: Delete the object to be deleted from the graph data.

[0156] In some embodiments, the unique identifier and object graph data information of the object to be deleted are extracted from the data deletion request. According to the unique identifier of the object to be deleted, the position of the object to be deleted in the key-value database is located. Then, a deletion operation is performed to delete the object to be deleted and its related information from the key-value database.

[0157] Step 503: Determine whether the object to be deleted has an attribute index.

[0158] In some embodiments, before deleting the key-value pair information of the object to be deleted from the key-value database, it is determined whether the object has an attribute index. When the object to be deleted has an attribute index, step 504 is entered; when the object to be deleted does not have an attribute index, 505 is entered.

[0159] Step 504: Delete the key-value pair information of the object to be deleted from the key-value database.

[0160] In some embodiments, since the object to be deleted has an attribute index, the object to be deleted has been encoded into the form of key-value pairs according to the existing attribute index and stored in the key-value database. Now, a data deletion request is received, which includes the object to be deleted and its attribute information. After deleting the object graph data information of the object to be deleted in the graph data, the key-value pair information of the object to be deleted also needs to be deleted in the key-value database.

[0161] Step 505: Delete the object graph data information of the object to be deleted from the graph data.

[0162] In some embodiments, since the object to be deleted does not have an attribute index, when deleting the object to be deleted, only the object graph data information of the object to be deleted in the graph database needs to be deleted.

[0163] In some embodiments, deleting the object to be deleted may also include deleting a specific attribute in the joint attribute index. For example, consider that the joint attribute index of all students in a class includes student number, name and gender, and the joint attribute index is represented as [student number, name, gender]. Now, due to the need to delete the student number attribute for further study, and retain the name and gender attributes, the attribute index after deletion is [name, gender]. Subsequently, the updated attribute index is encoded into a key-value pair and stored in a key-value database.

[0164] In the above steps 501 to 505, a comprehensive deletion operation of the object to be deleted can be implemented by obtaining a data deletion request. The request contains the object graph data information of the object to be deleted, and the object graph data information of the object can be directly deleted from the graph data. If the object to be deleted has an attribute index, the key-value pair information of the object can also be deleted from the key-value database. If the object to be deleted does not have an attribute index, the key-value pair information of the object can also be deleted from the key-value database. Therefore, regardless of whether the object to be deleted has an attribute index, the object and its related data can be completely deleted to ensure data integrity.

[0165] Through the above-mentioned method, the embodiment of the present application can realize fast retrieval and query of each attribute in the graph database by encoding the attribute index into key-value pair information. First, each attribute index is encoded, including information such as graph space identifier, partition identifier, index type, index identifier and attribute value, and then this information is combined with vertex identifier to form a key-value pair and stored in the key-value storage system. Using this encoding method, the query performance can be greatly improved and the search time can be saved. Since all attribute indexes have been pre-encoded and stored in the key-value storage system, only one query operation is required to directly scan the vertex identifier corresponding to the index. Compared with the traditional traversal query method, this encoding method reduces the overhead of multiple searches and greatly speeds up the query speed. In addition, this encoding method also supports multiple matching queries. By flexibly selecting the appropriate index type and index identifier according to different query requirements, multiple query operations such as exact matching, range query, and fuzzy matching can be realized. At the same time, since the index information has been encoded and stored in the key-value storage system, it can be easily updated and maintained, so that the consistency of the data is guaranteed.

[0166] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.

[0167] It is understandable that in the embodiments of the present application, related data such as user information is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0168] The graph data processing method provided by the embodiments of the present application can be applied to scenarios where it is necessary to find specified points through point attributes. For example, in the fields of friend recommendation in social networks and short video recommendation, etc., the graph data processing method can be used to quickly locate specific users, so as to provide recommended content with a higher response speed to users. For example, according to the video content liked by users, the social circle of users, etc., more accurate personalized recommendations can be made. This graph data processing method based on attribute indexing can bring significant advantages to graph database applications, not only improving the query efficiency, but also providing stronger support for personalized recommendation systems and social network analysis. By combining the characteristics of the graph database and this indexing technology, the performance of the system and the user experience can be effectively improved, meeting the high-efficiency requirements for data query and recommendation in complex business scenarios.

[0169] In a graph database, the structure of the entire graph can be stored in the form of key-value pairs, including point elements and edge elements. Through the identification information of the points, the out-degree or neighbor nodes of the points can be quickly queried. However, in some business scenarios, it is necessary to retrieve all points in the entire graph according to the specific attributes of the points. Without indexing technology, it is necessary to scan each point one by one and compare its attributes; this operation efficiency is very low. For example, in a social scenario, it is necessary to find all users located in SZ for targeted information recommendation. At this time, the point attribute indexing technology can be used to speed up the query. The embodiments of the present application propose a graph data processing method based on key-value pairs, which is designed for point attributes. The embodiments of the present application establish an index key encoding method through the graph space identifier, partition identifier, index type, index identifier of the attribute index, and the vertex identifier and attribute value corresponding to each attribute index, so as to quickly filter out the points with specified attribute values. At the same time, during the process of creating and deleting indexes, speed limiting operations are supported to minimize the impact on the read and write operations of the existing graph database. The embodiments of the present application can not only improve the query efficiency, but also optimize the data processing performance, meeting the real-time and high-concurrency query requirements. By establishing appropriate attribute indexes, the point attribute retrieval process can be accelerated, and the performance of graph data analysis and recommendation systems in business scenarios can be improved.

[0170] In the embodiments of the present application, the index encoding is to convert the index into the form of key-value pairs and store it in a key-value storage system. Specifically, the index value encoding includes the graph space identifier (space_id), partition identifier (partition_id), index type (flag), index identifier (index_id), meta information (meta) of the attribute index, and the vertex identifier (vertex_id) and attribute value (index_binary) corresponding to each attribute index, see Figure 5 , Figure 5 which is a schematic diagram of the index key encoding provided by the embodiments of the present application.

[0171] The graph space identifier (space_id) represents the unique identifier of the graph space, which is used to distinguish different graph data stores. When establishing graph data, each graph space will be assigned a specific identifier. This identifier can be used to determine the specific graph space to be operated on, so as to accurately locate and access the required graph data in a multi-graph data storage environment.

[0172] The partition identifier (partition_id) represents the partition identifier to which the data belongs and is used to divide the data into different partitions. When establishing graph data, a specific partition to which each data object belongs has been specified. Through the partition identifier, the specific partition containing the required data can be quickly located, thereby improving the efficiency of querying and retrieving.

[0173] The index type (flag) is used to mark the type of the index to distinguish different types of keys. For example, 'i' can be used to represent the point attribute index type, '1' or '0' can be used to represent the edge attribute index type, and other characters can be used to represent other attribute index types. Such a marking method can be customized according to specific needs to meet the differentiation and identification of index types in different scenarios.

[0174] In the index identifier (index_id), for a specific graph space, each index has a unique identifier for identification. This identifier is pre-assigned when the index is created and is used to uniquely identify the index object. Through this unique identifier, the specific index to be operated on can be accurately specified for relevant data query and retrieval operations.

[0175] The attribute value (index_binary) refers to the attribute information associated with the graph object. Through the index identifier, the corresponding column information, including descriptions of data type, data length, etc., can be obtained in the SchemaManager. The attribute value can be a binary format for storing attribute values. To support variable-length data types, the length of each non-fixed-length field is saved at the end of the attribute value, and each length is represented by an int32 type. At the same time, an attribute contains at least one piece of data. Different storage methods can be adopted according to different data types. For example, string-type attributes can be stored using UTF-8 encoding, and numerical-type attributes can be stored using a binary format.

[0176] The vertex identifier (vertex_id) is an identifier used to uniquely identify vertex objects in a graph database. The vertex identifier can be represented as a vertex identifier in string format, with the length of the vertex identifier saved at the end of the string. This design facilitates quickly locating the end of the attribute value to obtain the string columns length field stored therein. At the same time, the vertex identifier can also be found through the index to determine that the vertex identifier is actively input by the user. When inputting data, the corresponding vertex identifier will be provided, and through this identifier, the corresponding product or data can be directly retrieved.

[0177] Meta information is a professional term that contains metadata information related to graph objects, such as the type of graph object, creation time, modification time, etc. It is calculated based on the lengths of the previously mentioned content (index_binary and vertex_id).

[0178] After encoding, the value corresponding to the key is a null value, and all the information required for the index is saved on the key. By constructing a suitable key-value range as the index identifier, query operations can be effectively performed. To implement range queries, relevant interfaces for prefix scanning of the underlying key-value pair storage structure can be called. Prefix scanning is a query method that scans according to a specified prefix, and it can return all key-value pairs starting with that prefix. By reasonably setting the prefix of the key and the sorting rule, prefix scanning can be used to efficiently obtain and return vertex identifiers that meet the conditions. Therefore, through the encoded key-value pair structure and the reasonably constructed index identifier value range, combined with the relevant interfaces for prefix scanning of the underlying key-value pair storage structure, efficient acquisition and return of vertex identifiers can be achieved, which can not only improve the query performance and retrieval efficiency of the database system, but also effectively support the requirements of data management and query operations.

[0179] See Figure 6 , Figure 6 which is a schematic diagram of the implementation process of point query based on attribute index provided by the embodiments of this application. Next, specific descriptions will be made in combination with the Figure 6 steps shown.

[0180] Step 601: Query the target point according to the attribute.

[0181] During the query of the target point (object) according to the attribute, obtain the attribute conditions in the query request, such as "age is 20 years old", and then determine the set of vertices related to the target point attribute in some way, such as traversing the graph data or using the index. Then, for each target point, check whether its attribute value matches the query condition. The target point is the point containing the point to be queried.

[0182] Step 602: Determine whether an attribute index is established for the target point.

[0183] First, it will check whether there is an index related to the property to be queried in the graph database. These indexes can quickly locate the corresponding vertices according to the property values. If there is a property index, the key-value pair information corresponding to the index will be used to accelerate the query operation. Through the index, the vertices with matching property values can be directly located without traversing the entire graph data. If there is no property index, it will have to traverse the graph data to find the target points that meet the query conditions, which may result in longer query times and higher system overhead.

[0184] Among them, when the property index is established for the target point, it enters step 603; when the property index is not established for the target point, it enters step 604.

[0185] Step 603: When the property index is established for the target point, call the point property index to query the key-value pair information corresponding to the point property and obtain the result.

[0186] If the property index has been established for the target point, it is necessary to select and read the property index of the target point. During this period, a new remote procedure call (RPC) interface provided by the B-tree (Btree) needs to be called. Specify the property column to be scanned in the request and support equality indexing. During this process, a request will be initiated to require the B-tree to perform the scanning and retrieval operations of the property index. Since the number of points obtained may be relatively large, it is necessary to consider using the offset+limit method to obtain the points that meet the conditions in batches. This can process a large amount of data in batches and avoid performance problems caused by obtaining too much data at one time. Since each B-tree is responsible for several partition data and uses the hash sharding technology, when the query layer queries the equality properties that meet a certain condition, it needs to send the request concurrently to all B-trees and aggregate the results returned by each B-tree. This can make full use of the advantages of parallel processing and improve the query efficiency.

[0187] By calling the index query interface and providing the property value to be queried, the key-value pair information is obtained. Since the indexes may be distributed on multiple nodes or partitions, query requests will be sent concurrently to each node. Each node will perform the query operation in its local index and return the key-value pairs that meet the conditions. After receiving the results returned by each node, these results will be aggregated. Finally, a result set containing the key-value pair information that meets the query conditions will be generated.

[0188] When selecting an attribute index, simple prefix matching can be considered, that is, select the index with the longest prefix match with the query condition for querying. This can reduce the amount of data scanned and improve query efficiency. In subsequent iterations, the selection of the index can be further optimized. For example, for a query statement like "prop1>'xxx' and prop2=='xxx'", assuming there are the following indexes: index1(prop1), index2(prop2), index3(prop1,prop2), index4(prop2,prop1). In this case, index4 should be preferred, and the equal-valued prop2 index should be placed in the front for matching. The reason for choosing index4 is that by placing the equal-valued prop2 index in the front, the matching data can be located more quickly during the query. Because the condition for prop2 in the query statement is an equal comparison, while prop1 is a range comparison. Placing the prop2 index in the front can first use the equal condition for exact matching, and then use the range condition of prop1 for further filtering.

[0189] Step 604: Scan all graph data and obtain the results.

[0190] When there is no attribute index established for the target point, it is necessary to traverse all nodes and edges in the graph one by one to find the data that matches the query condition. First, the graph data will be loaded into memory for quick access and processing. Then, a global scan operation is performed, traversing each node and edge to check whether its attributes meet the query condition. This process can be understood as a complete search on the entire graph data. During the search process, some optimization strategies will be used to improve performance. For example, the graph data may be divided into multiple regions and each region is processed in parallel to speed up the search. In addition, a caching mechanism can be utilized to save frequently accessed data in memory to reduce the IO overhead. However, since there is no established attribute index, it is impossible to directly skip the nodes and edges that do not meet the query condition, but it is necessary to traverse the entire graph data. This leads to a problem of low query efficiency, especially when the scale of the graph data is large.

[0191] See Figure 7 , Figure 7 FIG. Figure 7 shows the implementation process diagram of point insertion based on attribute index provided by the embodiment of the present application. Next, specific descriptions will be made in combination with the

[0192] Step 701: Obtain the operation of inserting the target point.

[0193] When performing a write operation on a graph database, the target points to be inserted (corresponding to the objects to be inserted in other embodiments) can be obtained through appropriate methods, which can be achieved by using query languages or APIs. For example, in the Cypher query language, the CREATE statement can be used to create new nodes and specify the labels and properties of the nodes.

[0194] Step 702: Determine whether the target points to be inserted need to establish attribute indexes.

[0195] Among them, if the target points to be inserted need to establish attribute indexes, go to step 703; if the target points to be inserted do not need to establish attribute indexes, go to step 705.

[0196] Step 703: Add attribute indexes for the target points.

[0197] When an attribute index needs to be established, operations will be performed according to the specific attribute information of the target points. By establishing mapping relationships between the values of certain key attributes and corresponding records or data entries, the speed of retrieving and searching for these attributes can be accelerated. For example, when receiving a request to establish an attribute index, the attributes to be indexed and the target points will be determined first. The target points usually refer to records, documents, or other storage units in the database. Then, for each target point, its attribute set will be traversed, and indexes will be created for each specified attribute. During the index creation process, appropriate index algorithms will be selected according to the attribute type and data structure. For example, if the attribute is numeric, a B+ tree or hash table may be used to build the index; if the attribute is text, a full-text search index, etc. can be considered.

[0198] When inserting a new target point data, in addition to writing its attributes and values into the main data table, corresponding index data will be constructed according to the associated attribute column index of the point. For example, if the associated attribute column of a certain point is "age", an index data structure based on "age" will be created additionally, and the attribute value of the point and the primary key information corresponding to the attribute value of the point will be stored in this index data structure. Then, through the batch write (partition batch) operation provided by the distributed key-value storage system (tikv client), data can be written into the key-value database in batches. The batch write operation refers to dividing the data to be written into multiple partitions and performing batch write operations for each partition. For each partition, using the partition batch write operation interface provided by the distributed key-value storage system, the data of this partition is written into the key-value database in batches. In each partition, by submitting multiple write operations at one time, the write efficiency and performance can be significantly improved.

[0199] Step 704: Encode the attribute indexes of the newly added target points into key-value pairs and store them in the key-value database.

[0200] When the attribute index of the target point is created, the attribute index of the target point includes the graph space identifier, partition identifier, index type, index identifier, vertex identifier of the target point to be inserted, and the attribute value of the target point attribute. Encoding is performed based on the graph space identifier, partition identifier, index type, index identifier, vertex identifier of the target point to be inserted, and the attribute value of the target point attribute to obtain the key-value pair and key-value pair information of the newly added target point. Finally, the key-value pair and key-value pair information of the newly added target point are stored in the key-value database.

[0201] Step 705: Point information of the newly added target point.

[0202] When the target point does not need to establish an attribute index, the data is processed according to the normal write operation, and the point information of the newly added target point is stored in the graph database. The normal write operation means writing the data into the database in a normal way, that is, inserting the attributes and values of each point as a record into the database table.

[0203] Step 706: Encode the point information of the newly added target point into a key-value pair and store it in the key-value database.

[0204] The point information of the newly added target point includes the graph space identifier, partition identifier, vertex identifier of the target point to be inserted, and the attribute value of the target point attribute, and does not include the index type and index identifier. Then, encoding is performed on the newly added target point according to the graph space identifier, partition identifier, vertex identifier of the target point to be inserted, and the attribute value of the target point attribute to obtain a key-value pair, but the index type and index identifier parts in the key are null values, and the key-value pair of the newly added target point is stored in the key-value database.

[0205] See Figure 8 , Figure 8 which is the schematic diagram of the implementation process of point update based on attribute index provided by the embodiment of the present application. Next, specific descriptions will be made in combination with the Figure 8 steps shown.

[0206] Step 801: Obtain the operation of updating the target point.

[0207] In some embodiments, similar to step 701, the node to be updated is obtained by using a query language or API, including information such as the attributes and attribute values of the node, and the target point is the point containing the attribute to be updated.

[0208] Step 802: Determine whether the target point to be updated has an attribute index.

[0209] In some embodiments, the metadata of the database system is queried to check the index information of the table to which the target point to be updated belongs. By querying the metadata of the table, the index information of the table can be obtained, including which attributes have indexes established, the type and uniqueness of the indexes, etc. Based on the attribute information of the target point to be updated and the index information in the metadata, it is evaluated whether there is a suitable attribute index available, which includes determining whether the attribute to be updated belongs to the established index columns.

[0210] Among them, if there is an attribute index for the target point to be updated, go to step 803; if there is no attribute index for the target point to be updated, go to step 806.

[0211] Step 803: When there is an attribute index for the target point to be updated, delete the old index of the target point and add the new index of the target point.

[0212] When there is an attribute index for the target point to be updated, for each target point to be updated, its old index is deleted, that is, the old attributes of the target point are removed from the index. Deleting the old index can ensure that the outdated index is no longer used to query the target point. Then, a data update operation is performed on the target point to be updated, which may include changing the value of an attribute, adding or deleting an attribute, etc. After updating the data of the target point, the server regenerates the new index according to the updated data, which may involve inserting the attribute values of the target point into the appropriate index data structure and reconstructing or adjusting the index.

[0213] To ensure data consistency, the operations of deleting the old index and adding the new index are encapsulated in an atomic operation. An atomic operation refers to a single, indivisible operation unit in the database system, either all of which are executed successfully or all of which fail. By placing the operations of deleting the old index and adding the new index in an atomic operation, it can be ensured that data inconsistency does not occur during the update process.

[0214] Step 804: Update the key-value pair information of the target point.

[0215] Step 805: Store the updated key-value pair information of the target point into the key-value database.

[0216] In some embodiments, a data update operation is performed on the target point to be updated, which involves replacing the original key-value pair information with the new key-value pair information. For example, according to the attributes and attribute values specified in the update operation, the corresponding key-value pair information is updated. After updating the key-value pair information of the target point, the updated key-value pair information needs to be stored into the key-value database. For example, insert the key-value pair information into the key-value database, or update the data in the existing key-value database according to the key-value pair information specified in the update operation.

[0217] Step 806: Update the point information of the target point.

[0218] Step 807: Encode the point information of the target point into key-value pair information and store it in the key-value database.

[0219] In some embodiments, when the target point to be updated has no attribute index, a data update operation is performed on the point information of the target point to be updated. Then, the point information of the target point is encoded into key-value pair information. For example, the key in the key-value pair information is encoded from the graph space identifier, partition identifier, index type, index identifier, and the vertex identifier of the target point to be inserted and the attribute value of the target point attribute. Among them, the index type and index identifier are null values, and the value is also a null value. After updating the target point data and encoding it into key-value pair information, the key-value pair information needs to be stored in the key-value database.

[0220] See Figure 9 , Figure 9 which is a schematic diagram of the implementation process of point deletion based on attribute index provided by the embodiments of the present application. Next, it will be specifically described in combination with the steps shown in Figure 9 shown below.

[0221] Step 901: Obtain the operation of deleting the target point.

[0222] In some embodiments, similar to Step 901, the node to be updated is obtained by using a query language or API, including information such as the attributes and attribute values of the node, and the target point is the point to be deleted.

[0223] Step 902: Determine whether the target point to be deleted needs to delete the attribute index.

[0224] In some embodiments, the list of attribute indexes associated with the target point to be deleted can be checked. This list records all the attribute indexes related to the target point. Compare the target point to be deleted with the attribute indexes, such as attribute type, value range, etc. Only when the attributes of the target point exactly match the attribute indexes will the attribute indexes be deleted.

[0225] Among them, if the target point to be deleted needs to delete the attribute index, go to Step 903; if the target point to be deleted does not need to delete the attribute index, go to Step 905.

[0226] Step 903: Delete the attribute index and key-value pair information of the target point to be deleted.

[0227] Step 904: Store the key-value pair information corresponding to the remaining attributes in the key-value database.

[0228] In some embodiments, when there is an attribute index for the target point to be deleted, an operation to delete the attribute index is performed. This includes deleting the index entry corresponding to the target point to be deleted from the attribute index list and updating the index data structure. By deleting the attribute index, it can be ensured that the attribute index of the deleted target point is no longer used in subsequent query operations, avoiding invalid query results and performance losses. Then, according to the identifier or other unique identifier of the target point, the key-value pair information corresponding to the target point to be deleted is located and deleted from the key-value database, which can clear the data related to the target point to be deleted, ensure the consistency and accuracy of the data, and guarantee atomicity. After deleting the key-value pair information of the target point to be deleted, the key-value pair information corresponding to the remaining attributes needs to be stored in the key-value database, including encoding the remaining attribute values of the target point into key-value pair information and storing it in the key-value database.

[0229] Step 905: Delete the point information of the target point to be deleted.

[0230] In some embodiments, when the target point to be deleted does not need to have its attribute index deleted, it means that the target point to be deleted has no attribute index. According to the identifier or other unique identifier of the target point, the point information corresponding to the target point to be deleted is located and deleted from the graph database, which includes the edges connected to the target point and the vertex of the target point itself.

[0231] Step 906: Encode the remaining point information of the target point to be deleted into key-value pair information and store it in the key-value database.

[0232] After deleting the point information of the target point to be deleted, the remaining point information needs to be encoded into key-value pair information and stored in the key-value database. The encoding process includes generating the key and the value. The key is obtained by encoding the graph space identifier, partition identifier, index type, index identifier, vertex identifier of the target point to be inserted, and the attribute value of the target point attribute. The index type and index identifier are null values, and the value is also a null value, without storing specific data content. Finally, the encoded key-value pair information is stored in the key-value database. By storing the remaining point information of the target point to be deleted in the form of key-value pairs, the information of other vertices and edges connected to the target point to be deleted can be retained, while ensuring the consistency and integrity of the data.

[0233] In summary, by encoding the attribute index into key-value pair information in the embodiments of the present application, fast retrieval and query of each attribute in the graph database can be achieved. First, each attribute index is encoded, including information such as graph space identifier, partition identifier, index type, index identifier, and attribute value. Then, these information are combined with the vertex identifier to form key-value pairs and stored in the key-value storage system. Using this encoding method, the query performance can be greatly improved and the search time can be saved. In addition, the encoding method also has atomicity guarantee. When performing operations such as inserting, updating, and deleting attribute indexes, since the attribute indexes have been pre-encoded and stored in the key-value storage system, the atomicity of the operations can be guaranteed, that is, the operations are either completely successful or completely failed, and there will be no situation of partial success.

[0234] Next, the exemplary structure of the graph data processing device 455 provided in the embodiments of the present application implemented as software modules will be further described. In some embodiments, as Figure 3 shown, the software modules in the graph data processing device 455 stored in the memory 440 may include:

[0235] A first acquisition module 4551, configured to acquire the graph data to be processed and acquire a plurality of attribute indexes corresponding to the graph data to be processed;

[0236] A second acquisition module 4552, configured to acquire the index data of each attribute index and acquire the vertex identifier and attribute value corresponding to each attribute index;

[0237] An encoding module 4553, configured to perform encoding processing on the index data of each attribute index, and the vertex identifier and attribute value corresponding to each attribute index, to obtain key-value pair information corresponding to each attribute index;

[0238] A first storage module 4554, configured to store the key-value pair information corresponding to each attribute index into the key-value database.

[0239] In some embodiments, the encoding module 4553 is further configured to determine the meta information of the attribute index based on the vertex identifier and attribute value corresponding to the attribute index; perform encoding processing on the index data of the attribute index, and the vertex identifier, attribute value, and meta information corresponding to the attribute index in a preset format, to obtain the key-value pair information corresponding to the attribute index.

[0240] In some embodiments, the encoding module 4553 is further configured to acquire a first data length of the vertex identifier of the attribute index and acquire a second data length of the attribute value of the attribute index; determine a third data length occupied by storing the first data length and the second data length; and determine the first data length, the second data length, and the third data length as the meta information of the attribute index.

[0241] In some embodiments, the graph data processing apparatus further includes: a third acquisition module configured to acquire a query request, where the query request includes a to-be-query attribute and a target value of the to-be-query attribute; a first determination module configured to determine a target index identifier corresponding to the to-be-query attribute; a second determination module configured to, when the target index identifier exists in the key-value database, determine, from the key-value database, target key-value pair information including the target index identifier, where an attribute value included in the target key-value pair information satisfies a matching condition with the target value; a third determination module configured to determine target vertex information corresponding to the target vertex identifier included in the target key-value pair information as a query result; and a first output module configured to output the query result.

[0242] In some embodiments, the graph data processing apparatus further includes: a fourth determination module configured to, when the target index identifier corresponding to the to-be-query attribute does not exist in the key-value database, traverse the graph data to determine target vertex information, where an attribute value of the to-be-query attribute in the target vertex information is the target value; and a second output module configured to determine the target vertex information as a query result and output the query result.

[0243] In some embodiments, the graph data processing apparatus further includes: a fourth acquisition module configured to acquire a data insertion request, where the data insertion request includes at least one object attribute of an object to be inserted and an attribute value of the object attribute; an insertion module configured to insert the object to be inserted into the graph data and acquire a vertex identifier of the object to be inserted; a establishment module configured to establish an attribute index of the at least one object attribute when it is necessary to establish an attribute index of the object to be inserted; a fifth acquisition module configured to acquire index data of the attribute index of the at least one object attribute; a first generation module configured to generate key-value pair information to be inserted based on the index data of the attribute index of the at least one object attribute, the vertex identifier of the object to be inserted, and the attribute value of the point attribute; and a second storage module configured to store the key-value pair information to be inserted into the key-value database.

[0244] In some embodiments, the graph data processing apparatus further includes: a second generation module configured to, when it is not necessary to establish an attribute index of the at least one object attribute, generate key-value pair information to be inserted based on the vertex identifier corresponding to the object to be inserted, the at least one object attribute of the object to be inserted, and the attribute value of the object attribute; and a third storage module configured to store the key-value pair information to be inserted into the key-value database.

[0245] In some embodiments, the graph data processing apparatus further includes: a sixth acquisition module, configured to acquire a data update request, where the data update request includes at least one object attribute of an object to be updated and an attribute value of the object attribute; a first update module, configured to update the graph data based on the at least one object attribute of the object to be updated and the attribute value of the object attribute, so as to obtain updated graph data; a seventh acquisition module, configured to, when the object to be updated has an attribute index, acquire a vertex identifier of the object to be updated, and acquire index data of the attribute index of the at least one object attribute; a third generation module, configured to generate new key-value pair information of the object to be updated based on the index data of the attribute index of the at least one object attribute, the vertex identifier of the object to be inserted, and the attribute value of the point attribute; a second update module, configured to update the initial key-value pair information of the object to be updated in the key-value database to the new key-value pair information.

[0246] In some embodiments, the graph data processing apparatus further includes: an eighth acquisition module, configured to, when the object to be updated has no attribute index, acquire a vertex identifier of the object to be updated; a fourth generation module, configured to generate new key-value pair information based on the vertex identifier corresponding to the object to be updated, at least one object attribute of the object to be updated, and the attribute value of the object attribute; a third update module, configured to update the initial key-value pair information of the object to be updated in the key-value database to the new key-value pair information.

[0247] In some embodiments, the graph data processing apparatus further includes: a ninth acquisition module, configured to acquire a data deletion request, where the data deletion request includes object graph data information of an object to be deleted; a first deletion module, configured to delete the object to be deleted from the graph data; a second deletion module, configured to, when the object to be deleted has an attribute index, delete the key-value pair information of the object to be deleted from the key-value database.

[0248] An embodiment of the present application provides a computer program product, which includes a computer program or computer executable instructions, and the computer program or computer executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer-readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the graph data processing method in the embodiments of the present application described above.

[0249] An embodiment of the present application provides a computer-readable storage medium storing computer executable instructions, where computer executable instructions or a computer program are stored therein. When the computer executable instructions or the computer program are executed by a processor, the processor will be caused to execute the graph data processing method provided in the embodiments of the present application. For example, as Figures 4A to 4G shown in the graph data processing method.

[0250] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; it may also be various devices including one or any combination of the above memories.

[0251] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0252] As an example, the computer-executable instructions may or may not correspond to a file in the file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code).

[0253] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0254] In summary, the embodiments of the present application encode the attribute index into key-value pair information and store it in the key-value storage system, which can achieve efficient attribute query and retrieval. This encoding method provides fast query performance, saved lookup time, and the flexibility to support multiple query operations, while ensuring the atomicity of operations.

[0255] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A method for processing graph data, characterized in that, The method includes: Obtaining the graph data to be processed, and obtaining a plurality of attribute indexes corresponding to the graph data to be processed; Obtaining the index data of each attribute index, and obtaining the vertex identifier and attribute value corresponding to each attribute index; Performing encoding processing on the index data of each attribute index, and the vertex identifier and attribute value corresponding to each attribute index, to obtain key-value pair information corresponding to each attribute index; Storing the key-value pair information corresponding to each attribute index into a key-value database.

2. The method according to claim 1, wherein The performing encoding processing on the index data of each attribute index, and the vertex identifier and attribute value corresponding to each attribute index, to obtain key-value pair information corresponding to each attribute index, includes: Executing the following process for each of the attribute indexes: Determining the meta information of the attribute index based on the vertex identifier and attribute value corresponding to the attribute index; Performing encoding processing on the index data of the attribute index, and the vertex identifier, attribute value and meta information corresponding to the attribute index according to a preset format, to obtain the key-value pair information corresponding to the attribute index.

3. The method according to claim 2, wherein The determining the meta information of the attribute index based on the vertex identifier and attribute value corresponding to the attribute index, includes: Obtaining a first data length of the vertex identifier of the attribute index, and obtaining a second data length of the attribute value of the attribute index; Determining a third data length occupied by storing the first data length and the second data length; Determining the first data length, the second data length and the third data length as the meta information of the attribute index.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtaining a query request, where the query request includes a to-be-query attribute and a target value of the to-be-query attribute; Determining a target index identifier corresponding to the to-be-query attribute; When the target index identifier exists in the key-value database, determining, from the key-value database, target key-value pair information including the target index identifier, where the attribute value included in the target key-value pair information satisfies a matching condition with the target value; Determining target vertex information corresponding to the target vertex identifier included in the target key-value pair information as a query result; Outputting the query result.

5. The method according to claim 4, wherein The method further includes: When the target index identifier corresponding to the to-be-query attribute does not exist in the key-value database, traversing the graph data to determine target vertex information, where the attribute value of the to-be-query attribute in the target vertex information is the target value; Determining the target vertex information as a query result, and outputting the query result.

6. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtaining a data insertion request, where the data insertion request includes at least one object attribute of an object to be inserted and an attribute value of the object attribute; Inserting the object to be inserted into the graph data, and obtaining a vertex identifier of the object to be inserted; When it is necessary to establish an attribute index of the object to be inserted, establishing an attribute index of the at least one object attribute; Obtaining index data of the attribute index of the at least one object attribute; Generating key-value pair information to be inserted based on the index data of the attribute index of the at least one object attribute, the vertex identifier of the object to be inserted, and the attribute value of the point attribute; Store the key-value pair information to be inserted into the key-value database.

7. The method according to claim 6, characterized in that, The method further includes: When it is not necessary to establish an attribute index for the at least one object attribute, generate key-value pair information to be inserted based on the vertex identifier corresponding to the object to be inserted, the at least one object attribute of the object to be inserted, and the attribute value of the object attribute; Store the key-value pair information to be inserted into the key-value database.

8. The method according to any one of claims 1 to 3, characterized in that, The further includes: Obtain a data update request, where the data update request includes at least one object attribute of the object to be updated and the attribute value of the object attribute; Update the graph data based on the at least one object attribute of the object to be updated and the attribute value of the object attribute to obtain updated graph data; When the object to be updated has an attribute index, obtain the vertex identifier of the object to be updated, and obtain the index data of the attribute index of the at least one object attribute; Generate new key-value pair information for the object to be updated based on the index data of the attribute index of the at least one object attribute, the vertex identifier of the object to be inserted, and the attribute value of the point attribute; Update the initial key-value pair information of the object to be updated in the key-value database to the new key-value pair information.

9. The method according to claim 8, characterized in that, The method further includes: When the object to be updated does not have an attribute index, obtain the vertex identifier of the object to be updated; Generate new key-value pair information based on the vertex identifier corresponding to the object to be updated, the at least one object attribute of the object to be updated, and the attribute value of the object attribute; Update the initial key-value pair information of the object to be updated in the key-value database to the new key-value pair information.

10. The method according to any one of claims 1 to 3, characterized in that The method further includes: Obtain a data deletion request, where the data deletion request includes the object graph data information of the object to be deleted; Delete the object to be deleted from the graph data; When the object to be deleted has an attribute index, delete the key-value pair information of the object to be deleted from the key-value database.

11. A graph data processing device, characterized in that, The apparatus includes: A first acquisition module, configured to acquire graph data to be processed and acquire a plurality of attribute indexes corresponding to the graph data to be processed; A second acquisition module, configured to acquire index data of each attribute index, and acquire the vertex identifier and attribute value corresponding to each attribute index; An encoding module, configured to perform encoding processing on the index data of each attribute index, and the vertex identifier and attribute value corresponding to each attribute index, to obtain key-value pair information corresponding to each attribute index; A first storage module, configured to store the key-value pair information corresponding to each attribute index into a key-value database.

12. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the method according to any one of claims 1 to 10 when executing the computer-executable instructions stored in the memory.

13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or the computer program, when executed by the processor, implement the method according to any one of claims 1 to 10.

14. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or the computer program, when executed by the processor, implement the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Graph database decimal index coding and decoding method based on key value storage

    CN120849671A

  • Key-value storage based graph database decimal index encoding and decoding method

    CN120849671B