Data fragment query method in distributed multi-mode database

By adopting a unified sharding decision-making mechanism based on Hash and a centralized unified sharding management architecture in multimodal databases, the problems of complex data sharding management and low cross-model query efficiency in multimodal databases are solved, efficient data storage and access are achieved, and user experience is improved.

CN120179729APending Publication Date: 2025-06-20上海沄熹科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510223657.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In multimodal databases, due to the existence of multiple data models, data sharding management is complex, sharding strategies are not targeted, data consistency is difficult to guarantee, cross-model query efficiency is low, affecting user experience.

Method used

The unified sharding decision-making mechanism based on Hash and a centralized unified sharding management architecture are used to calculate the Hash value through the same Key format (Table/table_id/HashPoint) and SHA-256 function to achieve unified and accurate sharding of data of different models.

Benefits of technology

It improves data storage and access efficiency, gives full play to the performance advantages of multi-mode databases, meets diversified business needs, reduces the complexity of user operations, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179729A_ABST
    Figure CN120179729A_ABST
Patent Text Reader

Abstract

The invention discloses a data fragment query method in a distributed multi-mode database, and relates to the technical field of distributed multi-mode databases. Comprising the steps of 1, using the same Key format for data in different modes, calculating a Hash value according to a data type, and performing modulo operation on the maximum value of the HashPoint according to the Hash value to obtain the HashPoint according to the Hash value, and 2, when a user writes or queries data, calculating the Hash value according to the type of the data written or queried by the user, performing modulo operation on the maximum value of the HashPoint according to the Hash value to obtain the HashPoint, and calculating the Hash value according to the type of the data written or queried by the user to obtain the HashPoint according to the Hash value. The method comprises the following steps: step 1, determining a data fragment where HashPoint is located according to HashPoint to obtain distribution information of the data fragment, and step 2, sending a related request to a corresponding node according to the distribution information of the data fragment, and receiving a result returned by the corresponding node after executing related operation according to the request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a method for data sharding query in a distributed multi-modal database, which relates to the technical field of distributed multi-modal databases. Background Art

[0002] A multi-modal database is a database that can support the processing of a mixture of multiple data models. This type of database provides a better solution for heterogeneous data, enabling the database to process structured, semi-structured, and unstructured data simultaneously. Data sharding refers to splitting data blocks into multiple smaller, independent fragments or subsets to improve system performance, scalability, and manageability. By dispersing data across multiple servers or nodes, this technology can effectively distribute the load and enhance the overall performance of the system.

[0003] In a multi-modal database, due to the existence of multiple data models, if each data type corresponds to a data management method, the management of multiple types of data will be very cumbersome. Especially in data sharding management, the sharding strategy lacks pertinence, data consistency is difficult to guarantee, and the cross-model query efficiency is low, reducing the user experience. Summary of the Invention

[0004] In view of the problems of the prior art, the present invention provides a method for data sharding query in a distributed multi-modal database, which realizes the efficient collaborative management of multi-modal data, improves the performance, scalability, and maintainability of the database, and meets diverse business requirements.

[0005] The specific solution proposed by the present invention is as follows:

[0006] The present invention provides a method for data sharding query in a distributed multi-modal database, including:

[0007] Step 1: For data of different modes, use the same Key format. The Key format is Table / table_id / HashPoint. Table / is the prefix, table_id is the ID of the table where the data shard is located. Calculate the Hash value according to the data type, and obtain HashPoint by taking the modulus of the Hash value with respect to the maximum value of HashPoint.

[0008] Step 2: When a user writes or queries data, calculate the Hash value according to the type of data written or queried by the user, obtain HashPoint by taking the modulus of the Hash value with respect to the maximum value of HashPoint, determine the data shard where HashPoint is located, and obtain the distribution information of the data shard.

[0009] Step 3: Send a relevant request to the corresponding node according to the distribution information of the data shard, and receive the result returned after the corresponding node executes relevant operations according to the request.

[0010] Further, in step 1 of the method for data shard query in a distributed multi-modal database, calculating the Hash value according to the data type and obtaining the HashPoint by taking the modulus of the maximum value of HashPoint with the Hash value includes:

[0011] For relational data, select the primary key field as the Hash key; for document data, select the document ID as the Hash key; for graph data, select the unique identifier of the node where the graph data is located as the Hash key.

[0012] Calculate the Hash value of the Hash key using the SHA-256 function, and obtain the HashPoint by taking the modulus of the Hash value with the maximum value of HashPoint.

[0013] Further, in step 1 of the method for data shard query in a distributed multi-modal database, preset the maximum value of HashPoint, and evenly distribute the range of HashPoint where each data shard is located according to the maximum value of HashPoint when creating the data table.

[0014] Further, in step 2 of the method for data shard query in a distributed multi-modal database, also manage the metadata of the data shards. The metadata of the data shards includes the replica location, data type, shard size, data volume, access permission of each shard, and the range of HashPoint corresponding to each shard.

[0015] Update the metadata in real time. When there are new shards created, data migrated, or shards merged, adjust the metadata records in a timely manner to ensure that the database kernel grasps the changes in the entire shard layout.

[0016] Further, in step 3 of the method for data shard query in a distributed multi-modal database, send a write data request to the corresponding node according to the distribution information of the data shards, and receive the result returned by the corresponding node after storing the data in the specified shard of this node according to the write data request;

[0017] Send a query data request to the corresponding node according to the distribution information of the data shards, and receive the query result returned by the corresponding node after querying the specified shard of this node according to the query data request.

[0018] The present invention also provides a data shard query device in a distributed multi-modal database, including a shard decision module, a data routing module, and a metadata management module.

[0019] The sharding decision module uses the same Key format for different types of data. The Key format is Table / table_id / HashPoint. Table / is the prefix, and table_id is the ID of the table where the data shard is located. Calculate the Hash value according to the data type, and obtain HashPoint by taking the modulo of the Hash value with the maximum value of HashPoint.

[0020] When the user writes or queries data, the sharding decision module calculates the Hash value according to the type of data written or queried by the user, and obtains HashPoint by taking the modulo of the Hash value with the maximum value of HashPoint. The metadata management module returns the distribution information of the data shard according to the data shard where HashPoint is located.

[0021] The data routing module sends relevant requests to the corresponding nodes according to the distribution information of the data shards, and receives the results returned after the corresponding nodes perform relevant operations according to the requests.

[0022] Furthermore, the sharding decision module of the data shard query device in a distributed multi-mode database calculates the Hash value according to the data type, and obtains HashPoint by taking the modulo of the Hash value with the maximum value of HashPoint, including:

[0023] For relational data, select the primary key field as the Hash key; for document data, select the document ID as the Hash key; for graph data, select the unique identifier of the node where the graph data is located as the Hash key.

[0024] Use the SHA-256 function to calculate the Hash value of the Hash key, and obtain HashPoint by taking the modulo of the Hash value with the maximum value of HashPoint.

[0025] Furthermore, the sharding decision module of the data shard query device in a distributed multi-mode database presets the maximum value of HashPoint, and evenly distributes the range of HashPoint where each data shard is located according to the maximum value of HashPoint when creating a data table.

[0026] Furthermore, the metadata management module of the data shard query device in a distributed multi-mode database also manages the metadata of the data shards. The metadata of the data shards includes the replica location of each shard, the data type, the size of the shard, the data volume, the access permission of the shard, and the range of HashPoint corresponding to each shard.

[0027] Update the metadata in real time. When there are new shards created, data migrated, or shards merged, adjust the metadata records in time to ensure that the database kernel grasps the changes in the entire shard layout.

[0028] Furthermore, the data routing module of the data sharding query device in the distributed multi-mode database sends a write data request to the corresponding node according to the distribution information of the data shards, and receives the result returned by the corresponding node after storing the data in the specified shard of this node according to the write data request.

[0029] Sends a query data request to the corresponding node according to the distribution information of the data shards, and receives the query result returned by the corresponding node after querying the specified shard of this node according to the query data request.

[0030] The beneficial effects of the present invention are as follows:

[0031] Through the unified sharding decision mechanism based on Hash and the centralized unified sharding management architecture, the unified and accurate sharding of different model data is realized, the data storage and access efficiency are improved, the performance advantages of the multi-mode database are fully exerted, and the diverse business requirements are met. For users, for databases or tables of different models, users only need to execute the same syntax to control the data sharding of different engines. This reduces the complexity of user operations and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of the application process of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited are not intended to limit the present invention.

[0034] Embodiment 1

[0035] The present invention provides a data sharding query method in a distributed multi-mode database, including:

[0036] Step 1: For data of different modes, use the same Key format, where the Key format is Table / table_id / HashPoint, Table / is the prefix, table_id is the ID of the table where the data shard is located, calculate the Hash value according to the data type, and obtain HashPoint by taking the modulus of the Hash value with respect to the maximum value of HashPoint.

[0037] Among them, calculating the Hash value according to the data type and obtaining HashPoint by taking the modulus of the Hash value with respect to the maximum value of HashPoint may include:

[0038] For relational data, select the primary key field as the Hash key; for document data, select the document ID as the Hash key; for graph data, select the unique identifier of the node where the graph data is located as the Hash key.

[0039] Use the SHA-256 function to calculate the Hash value of the Hash key, and obtain the HashPoint by taking the modulus of the Hash value with respect to the maximum value of HashPoint.

[0040] Preset the maximum value of HashPoint. When building a data table, evenly distribute the range of HashPoint where each data shard is located according to the maximum value of HashPoint, and the number of HashPoints contained in each shard is the same. The initial number of shards can also be set by the user himself.

[0041] Step 2: When the user writes or queries data, calculate the Hash value according to the type of data written or queried by the user, obtain the HashPoint by taking the modulus of the Hash value with respect to the maximum value of HashPoint, determine the data shard where the HashPoint is located, and obtain the distribution information of the data shard.

[0042] Among them, in Step 2, the metadata of the data shard can also be managed. The metadata of the data shard includes the replica location, data type, size of the shard, data volume, access permission of the shard, and the range of HashPoint corresponding to each shard.

[0043] Update the metadata in real time. When there are new shards created, data migrated, or shards merged, adjust the metadata records in time to ensure that the database kernel grasps the changes in the entire shard layout.

[0044] Step 3: Send relevant requests to the corresponding nodes according to the distribution information of the data shards, and receive the results returned after the corresponding nodes execute relevant operations according to the requests.

[0045] Among them, in Step 3, send a write data request to the corresponding node according to the distribution information of the data shard, and receive the result returned after the corresponding node stores the data in the specified shard of this node according to the write data request;

[0046] Send a query data request to the corresponding node according to the distribution information of the data shard, and receive the query result returned after the corresponding node queries the specified shard of this node according to the query data request.

[0047] Embodiment 2

[0048] The present invention also provides a data shard query device in a distributed multi-mode database, including a shard decision module, a data routing module, and a metadata management module.

[0049] For data of different modes, the sharding decision module uses the same Key format, which is Table / table_id / HashPoint. Table / is the prefix, table_id is the ID of the table where the data shard is located. Calculate the Hash value according to the data type, and obtain HashPoint by taking the modulus of the Hash value with respect to the maximum value of HashPoint.

[0050] When the user writes or queries data, the sharding decision module calculates the Hash value according to the type of data written or queried by the user, and obtains HashPoint by taking the modulus of the Hash value with respect to the maximum value of HashPoint. HashPoint determines the shard to which the data should belong. The maximum value of HashPoint is a fixed value and can be set by the user. This method ensures that data of different models are sharded based on the same Hash rule, guaranteeing the uniformity and consistency of data distribution.

[0051] The metadata management module returns the distribution information of the data shard according to the data shard where HashPoint is located. The metadata management module also manages the shard metadata of the multi-mode database, including information such as the replica location, data model type, shard size, data volume, and access permissions of each shard. The metadata management module will update the metadata in real time. When there are operations such as creating a new shard, data migration, or shard merging, it will adjust the metadata records in a timely manner to ensure that the database kernel can accurately grasp the changes in the entire shard layout. At the same time, the range of HashPoint corresponding to each shard is also recorded in the metadata to facilitate quickly locating the shard where the data is located.

[0052] The data routing module sends relevant requests to the corresponding nodes according to the distribution information of the data shards, and receives the results returned after the corresponding nodes execute relevant operations according to the requests.

[0053] Among them, the data routing module is responsible for accurately routing requests to the corresponding shards based on the Hash calculation results during data read and write operations. When writing data, it receives the data from the application layer and the shard number calculated by the sharding decision module based on Hash, and sends the data to the specified shard of the corresponding storage node. When reading data, it recalculates the Hash value according to the Hash key in the query condition, combines the global index information, determines the involved shards, and distributes the query requests to these shards. For example, when receiving a cross-model query request involving a user table of relational data and user comments of document-type data, the data routing module parses the query statement, extracts the Hash key, calculates the Hash value, and then combines the physical location mapping relationship of the unified sharding of different model data recorded in the global index to route the query requests to the corresponding relational data shard and document-type data shard respectively.

[0054] For the information interaction and execution process among the modules in the above device, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.

[0055] Similarly, the device of the present invention realizes unified and accurate sharding of different model data through a unified sharding decision mechanism based on Hash and a centralized unified sharding management architecture, improves data storage and access efficiency, gives full play to the performance advantages of the multi-model database, and meets diverse business requirements. For users, for libraries or tables of different models, users only need to execute the same syntax to control the data sharding of different engines. This reduces the complexity of user operations and improves the user experience.

[0056] It should be noted that not all steps and modules in the above processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structure described in the above embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities separately, or some components in multiple independent devices may be jointly implemented.

[0057] The above embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. A data sharding query method in a distributed multi-mode database, characterized by: include: Step 1: For data in different modes, use the same Key format, which is Table / table_id / HashPoint. Table / is the prefix, and table_id is the ID of the table where the data shard is located. Calculate the Hash value based on the data type, and obtain the HashPoint by taking the modulus of the maximum HashPoint value based on the Hash value. Step 2: When the user writes or queries data, the Hash value is calculated according to the type of data written or queried by the user, and the HashPoint is obtained by taking the modulus of the maximum value of the HashPoint according to the Hash value, and the data shard where the HashPoint is located is determined to obtain the distribution information of the data shard. Step 3: Send relevant requests to the corresponding nodes according to the distribution information of the data shards, and receive the results returned by the corresponding nodes after performing relevant operations according to the requests.

2. According to the method for querying data fragments in a distributed multi-mode database according to claim 1, it is characterized by: In step 1, the Hash value is calculated according to the data type, and the HashPoint is obtained by taking the modulus of the maximum value of the HashPoint according to the Hash value, including: For relational data, select the primary key field as the hash key; for document data, select the document ID as the hash key; for graph data, select the unique identifier of the node where the graph data is located as the hash key. The SHA-256 function is used to calculate the Hash value of the Hash key, and the HashPoint is obtained by taking the modulus of the maximum value of the HashPoint according to the Hash value.

3. According to the method for querying data fragments in a distributed multi-mode database according to claim 1, it is characterized by: In step 1, the maximum HashPoint value is preset. When building the data table, the range of HashPoint for each data shard is evenly distributed according to the maximum HashPoint value.

4. The method for querying data fragments in a distributed multi-mode database according to claim 1, characterized in that Step 2 also manages the metadata of the data shards, which includes the replica location of each shard, data type, size of the shard, amount of data, access rights to the shard, and the range of the HashPoint corresponding to each shard. Update metadata in real time. When new shards are created, data is migrated, or shards are merged, adjust metadata records in a timely manner to ensure that the database kernel is aware of changes in the entire shard layout.

5. According to the method for querying data fragments in a distributed multi-mode database according to claim 1, The feature is that in step 3, a write data request is sent to a corresponding node according to the distribution information of the data shard, and the corresponding node is received to store the data in the designated shard of the node according to the write data request and return the result; Send a query data request to the corresponding node according to the distribution information of the data shard, and receive the query result returned by the corresponding node after querying the specified shard of the node according to the query data request.

6. A data fragmentation query device in a distributed multi-mode database, characterized in that Including sharding decision module, data routing module and metadata management module, The sharding decision module uses the same Key format for data in different modes. The Key format is Table / table_id / HashPoint, where Table / is the prefix and table_id is the ID of the table where the data shard is located. The Hash value is calculated based on the data type, and the HashPoint is obtained by taking the modulus of the maximum HashPoint value based on the Hash value. When a user writes or queries data, the sharding decision module calculates the Hash value according to the type of data written or queried by the user, and obtains the HashPoint modulo the maximum value of the HashPoint according to the Hash value. The metadata management module returns the distribution information of the data shard according to the data shard where the HashPoint is located. The data routing module sends relevant requests to the corresponding nodes according to the distribution information of the data shards, and receives the results returned by the corresponding nodes after performing relevant operations according to the requests.

7. The data sharding query device in a distributed multi-mode database according to claim 6 is characterized in that the sharding The decision module calculates the Hash value according to the data type, and obtains the HashPoint by taking the modulus of the maximum HashPoint value according to the Hash value, including: For relational data, select the primary key field as the hash key; for document data, select the document ID as the hash key; for graph data, select the unique identifier of the node where the graph data is located as the hash key. The SHA-256 function is used to calculate the Hash value of the Hash key, and the HashPoint is obtained by taking the modulus of the maximum value of the HashPoint according to the Hash value.

8. The data sharding query device in a distributed multi-mode database according to claim 6 is characterized in that the sharding The decision module presets the maximum value of HashPoint. When building a data table, the range of HashPoint where each data shard is located is evenly distributed according to the maximum value of HashPoint.

9. The data fragmentation query device in a distributed multi-mode database according to claim 6 is characterized in that The metadata management module also manages the metadata of data shards, which includes the replica location, data type, size, data volume, access rights of each shard, and the range of HashPoint corresponding to each shard. Update metadata in real time. When new shards are created, data is migrated, or shards are merged, adjust metadata records in a timely manner to ensure that the database kernel is aware of changes in the entire shard layout.

10. The data fragmentation query device in a distributed multi-mode database according to claim 6, characterized in that The data routing module sends a write data request to the corresponding node according to the distribution information of the data shards, and receives the result returned by the corresponding node after storing the data in the specified shard of the node according to the write data request; Send a query data request to the corresponding node according to the distribution information of the data shard, and receive the query result returned by the corresponding node after querying the specified shard of the node according to the query data request.