Data object storage method and device, storage medium and electronic equipment
Selecting data storage nodes through hash processing and weighted hash algorithms solves the problem of low storage efficiency in the existing technology, realizes more efficient and stable data storage, reduces load imbalance between nodes, and improves system performance and scalability.
Patent Information
- Application Number
- CN202510553159.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, the storage efficiency of data objects is low and the effect is poor, especially when the system scale changes, the single point of failure problems caused by high data migration costs, unstable load balancing and centralized metadata management affect system availability and scalability.
By hashing the identification of the data group to be stored, combined with the distributed hierarchical hash algorithm and the new weighted hash algorithm, the data storage node with the largest weight is selected, which reduces the computational complexity of node traversal, avoids load imbalance, and optimizes the performance and stability of the storage system.
It improves the overall performance and stability of the storage system, reduces load imbalance between nodes, reduces computing complexity and bandwidth usage, and improves storage efficiency and system scalability.
Smart Images

Figure CN120491893A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for storing data objects, a storage medium, and an electronic device. Background Art
[0002] In the current information age, decentralized object storage technology is rapidly developing, becoming a crucial tool for addressing the needs of massive data storage and dynamic expansion. This technology encompasses two main areas: consistent hashing algorithms, widely used in distributed storage systems to achieve balanced data distribution and dynamic expansion; and decentralized storage systems, which, through distributed architecture design, aim to reduce reliance on centralized management and improve system scalability and availability.
[0003] However, this approach firstly leads to high data migration costs. When the system scale changes, some data needs to be remapped and migrated, which consumes a large amount of resources and affects system performance. Secondly, load balancing is unstable. Especially when storage node performance varies significantly, the consistent hashing algorithm struggles to ensure even data distribution, causing some nodes to be overloaded and other nodes to waste resources. Finally, the centralized metadata management component introduces a single point of failure, affecting the availability and scalability of the entire system. In other words, the existing technology suffers from low data object storage efficiency, resulting in poor storage performance.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The embodiments of the present application provide a method and apparatus for storing data objects, a storage medium, and an electronic device, so as to at least solve the technical problems of low storage efficiency and poor effect of data objects in the prior art.
[0006] According to one aspect of an embodiment of the present application, a method for storing data objects is provided, comprising: obtaining at least one data group to be stored, and performing hash processing on an identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored; determining a target node group that matches the data group to be stored in a node group set based on a hash processing result of the data group to be stored, wherein each of the multiple node groups in the node group set includes at least one data storage node; determining a target data storage node that matches the data object to be stored in the target node group based on a weight selection condition, wherein the weight selection condition is determined based on a weight value of the data storage node in the node group; and storing the data object to be stored in the target data storage node.
[0007] According to another aspect of an embodiment of the present application, a storage device for data objects is also provided, including: an acquisition unit, which acquires at least one data group to be stored and performs hash processing on an identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored; a first determination unit, which determines, in a node group set, a target node group that matches the data group to be stored based on a hash processing result of the data group to be stored, wherein each of the multiple node groups in the node group set includes at least one data storage node; a second determination unit, which determines, in the target node group, a target data storage node that matches the data object to be stored based on a weight selection condition, wherein the weight selection condition is determined based on a weight value of the data storage node matching in the node group; and a storage unit, which stores the data object to be stored in the target data storage node.
[0008] Optionally, the above-mentioned first determination unit includes a third determination module, which is used to determine at least one of the data storage nodes in the target layer of the deployment architecture according to the placement rules in the initialization information, wherein information synchronization is achieved between the various layers in the deployment architecture through a target detection process, and the placement rules are used to indicate the storage method of the data object; the data storage nodes that pass the node status detection are stored in the data storage node set.
[0009] Optionally, the above-mentioned third determination module includes a grouping module, which is used to group the data storage node set according to the target grouping number to obtain the node group set; perform modulo processing on the hash value of the data to be stored according to the target grouping number to obtain a modulo result value; determine the target node group according to the modulo result value, wherein the serial number value of the target node group is consistent with the modulo result value.
[0010] Optionally, the above-mentioned grouping module is also used to determine the sum of the weight values based on the weight values of each data storage node in the data storage node set; determine the target overall weight value corresponding to each of the multiple node groups based on the target number of groups and the sum of the weight values; group the data storage nodes in the data storage node set according to the weight balancing rule to obtain the node group set, wherein the weight balancing rule is used to make the overall weight value of each node group and the target overall weight value meet the target difference condition.
[0011] Optionally, the above-mentioned second determination unit is also used to determine the first parameter based on the identifier of the data group to be stored, the disturbance factor parameter value and the node serial number of the data storage node; determine the selection node weight value of the data storage node based on the product of the first parameter and the weight value of the data storage node; and determine the data storage node with the largest selection node weight value as the target data storage node.
[0012] Optionally, the above-mentioned storage unit includes a detection module for determining the current data storage node information of the current deployment architecture when it is detected that the number of data storage nodes in the deployment architecture meets the data object migration trigger condition; updating the node group set and the data storage nodes in each node group in the node group set according to the current data storage node information; and migrating the reference data according to the overall weight value of each node group in the node group set after the update until the weight balance condition is met.
[0013] Optionally, the above-mentioned detection module includes a migration module, which is used to determine that the data object migration trigger condition is met and obtain the current data storage node information of the current deployment architecture when the number of data storage nodes whose node load is greater than the target load threshold is greater than a first number; determine that the data object migration trigger condition is met and obtain the current data storage node information of the current deployment architecture when the difference between the number of nodes of the current data storage nodes and the number of nodes of the original data storage nodes in the deployment architecture is greater than a first threshold; determine that the data object migration trigger condition is met and obtain the current data storage node information of the current deployment architecture when the difference between the overall weight values of any two node groups in the deployment architecture is greater than a second threshold.
[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned method for storing data objects when running.
[0015] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-mentioned data object storage method through the computer program.
[0016] In an embodiment of the present application, first, the identifier of the data group to be stored is processed by hashing, wherein the data object to be stored is the data object in the data group to be stored. Then, based on the hashing result of the data group to be stored, the target node group is accurately screened from the node group set. By utilizing a hierarchical architecture and pre-built node groups, the computational complexity of traversing all network nodes is reduced, ensuring that the data is stored on the most appropriate set of nodes. The target node is further selected based on the weight value of the data storage node, avoiding load imbalance between nodes, reducing hot spots, and improving the overall performance and stability of the storage system. Finally, the data object to be stored is stored in the target data storage node. This solves the technical problems of low storage efficiency and poor effect of data objects in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 is a schematic diagram of an application environment of an optional data object storage method according to an embodiment of the present application;
[0019] Figure 2 is a flow chart of an optional method for storing data objects according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of an optional data object storage method according to an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of another optional method for storing data objects according to an embodiment of the present application;
[0022] Figure 5 is a schematic diagram of an optional deployment architecture according to an embodiment of the present application;
[0023] Figure 6 This is an optional process communication diagram according to an embodiment of the present application;
[0024] Figure 7 is a schematic diagram of another optional method for storing data objects according to an embodiment of the present application;
[0025] Figure 8 is a schematic diagram of another optional method for storing data objects according to an embodiment of the present application;
[0026] Figure 9 is a schematic structural diagram of an optional data object storage device according to an embodiment of the present application;
[0027] Figure 10 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] According to one aspect of the embodiment of the present application, a method for storing a data object is provided. Optionally, as an optional implementation, the method for storing a data object can be applied to, but is not limited to, Figure 1 in the environment shown.
[0031] The terminal device 102 includes a display 108, a processor 106, and a memory 104; the server 112 includes a database 114 and a processing engine 116; and the environment also includes a network 110;
[0032] The server 112 executes S102-S108, S102, obtains at least one data group to be stored, and performs hash processing on the identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored;
[0033] S104, determining a target node group that matches the data group to be stored in the node group set according to the hash processing result of the data group to be stored, wherein each of the multiple node groups in the node group set includes at least one data storage node;
[0034] S106, determining a target data storage node that matches the data object to be stored in the target node group according to a weight selection condition, wherein the weight selection condition is determined according to a weight value of the data storage node in the node group;
[0035] S108: Store the data object to be stored in the target data storage node.
[0036] Optionally, in this embodiment, the above-mentioned terminal device can be a terminal device configured with a target client, which can include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client can be a video client, an instant messaging client, a browser client, an education client, etc. The above-mentioned network can include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The above-mentioned server can be a single server, or it can be a server cluster composed of multiple servers, or a cloud server. The above is only an example, and this embodiment does not impose any limitation on this.
[0037] Alternatively, as an optional implementation, Figure 2 As shown, the storage method of the above data object includes:
[0038] S202, obtaining at least one data group to be stored, and performing hash processing on an identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored;
[0039] S204, determining a target node group that matches the data group to be stored in the node group set according to the hash processing result of the data group to be stored, wherein each of the multiple node groups in the node group set includes at least one data storage node;
[0040] S206, determining a target data storage node that matches the data object to be stored in the target node group according to a weight selection condition, wherein the weight selection condition is determined according to a weight value of the data storage node in the node group;
[0041] S208: Store the data object to be stored in the target data storage node.
[0042] In the above step S202, the above data objects can be any type of data, such as files, pictures, videos, etc. The data group aggregates multiple data objects together for unified management, maps them to physical storage nodes, and simplifies the data copying and recovery process. For example, video data objects correspond to video-type data groups. The data in the data group is distributed and redundantly stored according to the strategy set in the storage bucket (such as the number of copies or the erasure code scheme). Each data group is assigned a unique identifier and mapped to the leaf-layer storage node using the distributed hierarchical hashing (DHH) algorithm. When a storage node fails, is overloaded, goes online, or goes offline, the mapping information of the data group can be used to quickly locate the data that needs to be rehashed and backfilled to ensure data integrity.
[0043] It should be noted that the mapping relationship between storage buckets and data groups is one-to-many, while the mapping relationship between data groups and storage nodes is many-to-many. The number of data groups has a significant impact on object storage performance. If the number is fewer than the number of storage nodes, data hashing will be uneven, significantly affecting storage efficiency and preventing the full utilization of distributed features. If the number is too large, the load on DHH's hash distribution calculations will increase significantly, significantly reducing computational efficiency. The recommended number of data groups is approximately 20 to 100 times the number of storage nodes. When system administrators preset the data group size for a storage bucket, if the data scale is expected to expand by 2 to 5 times the current estimated storage capacity, the data group multiplier can be preset to 50 to 100 to reduce the number of subsequent expansion adjustments. The result calculated by DHH is the many-to-many mapping relationship between data groups and storage nodes.
[0044] In the above step S204, the above node group set is a set of all node groups, each node group includes one or more data storage nodes, is located at the lowest layer of the distributed system, and is responsible for actual data storage tasks.
[0045] As an optional implementation, the target node group is determined by a modulo operation of a consistent hash, for example, Figure 3 As shown in the figure, the target group is selected as Group2 by calculating the hash value hash(x,r)mod N, where x is the unique identifier of the data group to be stored, and r is the disturbance factor used to adjust the distribution of data objects among storage nodes, so that the same x is hashed to different locations under different r values, and the value of r will be adjusted according to the currently selected node or the number of algorithm iterations to ensure uniform distribution of data and load balancing.
[0046] Further in steps S206-S208, the weight selection condition may be, for example, selecting the data storage node with the largest weight value as the target data storage node for storing the data object to be stored, thereby completing the data storage operation.
[0047] Alternatively, for example Figure 3 As shown in the target group by the formula max i (w i hash(x,r,i)) calculates the selection weight of each node and determines the node with the largest calculation result as the target data storage node, where w i is the weight of node i, i is the node number, that is, the first step is to select the group, and the second step is to select the node. Give each element a lottery method, combined with its own weight, the node with a larger weight has a larger probability of being drawn. A hierarchical node selection diagram is shown as follows Figure 4 shown.
[0048] Through the above-mentioned implementation of the present application, an algorithm that comprehensively considers performance and rebalancing efficiency is developed by combining the distributed hierarchical hash algorithm (DHH) and the new weighted hash algorithm. By introducing the concept of grouping, while ensuring the hashing of data, the computational efficiency is optimized while reducing the scale of data migration when nodes change. Reducing the node traversal range significantly improves storage efficiency, reduces computational complexity, improves storage stability, and reduces bandwidth usage and consumption.
[0049] In an optional embodiment, before determining a target node group matching the data group to be stored in the node group set based on the hash processing result of the data group to be stored, the following steps are included:
[0050] S1, determining at least one data storage node in a target layer of a deployment architecture according to a placement rule in the initialization information, wherein information synchronization is achieved between the layers in the deployment architecture through a target detection process, and the placement rule is used to indicate how the data object is stored;
[0051] S2, storing the data storage nodes that pass the node status detection in the data storage node set.
[0052] It's important to note that the object's storage location, replica information, backup methods, and data bucketing all require an intermediate abstraction layer to handle. Therefore, the concept of buckets is introduced to isolate and manage storage objects based on business needs, data characteristics, or security policies. Each bucket can define specific storage policies, such as the number of replicas, erasure coding parameters, and performance levels, to meet the data reliability, availability, and performance requirements of different business scenarios. Data in different buckets is independent of each other, supporting multi-tenant environments. Data from different tenants is stored in different logical partitions to ensure data privacy and security.
[0053] In step S1, the initialization information may include, but is not limited to, a unique identifier x, a selection number n, a level t, level nodes nodes, and a placement rule pr. A unique identifier refers to an identifier that uniquely identifies an object within a bucket. The selection number refers to the number of elements to be selected, the level t refers to the level at which the selected element is located, and the level nodes refers to an ordered set of all NRDs in the level, which together form a hierarchical mapping table. The placement rule refers to a pre-set data placement method under a bucket, such as the type of replica, the number of replicas, and the fault domain division.
[0054] Alternatively, as Figure 5The deployment architecture shown in the figure includes the row layer, the rack layer, the host layer, and the storage layer. At each layer, a unified node runtime daemon (NRD), namely the target detection process mentioned above, manages the nodes, including the topology data. Figure 6 As shown in the figure, NRDs at each level maintain data synchronization through a consistency protocol, exchanging the status of all nodes in their layer and their underlying topology information in real time to ensure high availability and topological consistency for the entire system. The leaf nodes at the bottom of the tree structure, known as storage nodes, are responsible for recording the topological information of nodes at the same level. They also manage stored data, planning designated storage media space, monitoring media health, adjusting storage space, defragmenting storage, storing data, recording metadata, and directly providing data read and write services. Typically, this node is directly responsible for a single storage medium, such as a single disk.
[0055] By replacing centralized management with NRD (Node Runtime Daemon), each tier of the storage system can independently maintain its own node-level topology information, avoiding the bottleneck problem of global management components. NRD's main functions are to maintain topology information and ensure high availability within the tier. Its specific functions are as follows:
[0056] Maintaining topological information for nodes and subnodes at the same level: Each NRD holds information about all NRDs within its level and the topological data for each node's subnodes (e.g., a rack NRD records the status of its hosts, and a host NRD records the information of its storage nodes). Topological information is regularly synchronized between NRDs at the same level using a heartbeat mechanism and an incremental update protocol to ensure that all nodes have the latest status.
[0057] Achieve high availability within the layer: A consistent communication network is established between NRDs on the same layer to achieve fault detection, state backup, and consistency synchronization. When an NRD fails, the other NRDs quickly detect it and mark the NRD node as invalid in their topology information, eliminating it from the node list mapped in the layer topology. An alarm is also triggered to alert upper-layer services for troubleshooting. Because multiple NRDs at the same level back up each other, availability is not affected, but the cluster health status is marked as yellow, requiring resolution before cluster upgrades and node adjustments can be performed.
[0058] No data distribution computing responsibilities: NRD's functions are limited to the collection, maintenance, and dissemination of topology information, and it does not directly participate in the calculation of object data distribution. This reduces the NRD's computational burden and removes its bottleneck as a node in traditional centralized object storage systems.
[0059] NRD description attributes include: position identifier is used to clarify the specific position of the node in the hierarchical structure; level information indicates the level of the node in the architecture (such as row, rack, host, storage); connection information connect ensures communication between nodes; siblings set of peer nodes and children set of subordinate nodes maintain the relationship and communication between levels; status reflects the current operating status of the node; weight represents the weight of the node, which affects data distribution decisions; and a series of extensible fields such as version number, geographic location information, storage capacity and business node status.
[0060] As an optional implementation, Figure 7 As shown, first use the take method to add the ordered set under the layer to the ordered set calculated at this layer. For example, in the first calculation, the root node NRD ordered set is output.
[0061] The top rule described in the select method is used to describe the placement rule. When selection fails and selection is retried, the replica method and the erasure code method have different replacement rules. The top selection method applies to replicas, that is, the elements at the back of the ordered set queue are mapped forward in sequence to replace the positions left vacant due to failures, that is, the nodes at the front of the ordered set are given priority. The erasure code method has different requirements for replacement. Fixed positions cannot be replaced subsequently. The rule requires replacement of elements at position r+n, where r is the starting position in the selection and n is the number of nodes that need to be selected in the current level. If the storage policy requires 3 copies of the data, then the value of n will be 3.
[0062] The select method also includes a weighted hashing algorithm (alg), which is used to perform weighted calculations on the nodes in the current level, ensuring that the probability of each node being selected is proportional to its weight value. When calculating r' (i.e., the new index), if the original weight is used directly, the probability of all nodes being selected will be uneven, and nodes with higher weights will usually be selected more frequently. Therefore, in order to further optimize, the weights can be adjusted through the alg method to make the node selection more intelligent and balanced, such as Figure 7 The select algorithm outputs the selected nodes in each level.
[0063] Optionally, in the above step S2, the data storage nodes that pass the node status detection include but are not limited to nodes that meet health status standards, meet storage resource occupancy conditions, have normal network connectivity, etc.
[0064] In an optional embodiment, after determining a target node group matching the data group to be stored in the node group set according to a hash processing result of the data group to be stored, the following steps are included:
[0065] S1, grouping the data storage node set according to the target grouping number to obtain a node group set;
[0066] S2, performing a modulo process on the hash value of the data to be stored according to the target number of groups to obtain a modulo result value;
[0067] S3, determining the target node group according to the modulo result value, wherein the sequence number value of the target node group is consistent with the modulo result value.
[0068] In the above step S1, for example, the above data storage node set can be grouped according to the preset target grouping number in an evenly distributed manner to obtain the above node group set. The number of nodes in the above node groups can be the same, or the sum of the node weights in different node groups can be the same. This is just an example.
[0069] In the above steps S2-S3, as an optional implementation, the hash value is divided by the target number of groups, that is, hash(x,r) mod N, and the remainder is obtained. The remainder is the above modulo result value, and then the above data group to be stored is matched with the node group pointed to by the modulo result value, for example, Figure 8 As shown, data group 1 is assigned disks 1, 3, and 5 (storage nodes store) of node group 1 through a hash algorithm, and data group 2 is assigned disks 2, 3, and 4 of node group 2 through a hash algorithm. If the data object (obj) to be stored belongs to data group 1, it means that the data can be stored in a storage node selected in node group 1.
[0070] In an optional implementation, grouping the data storage node set according to the target number of groups to obtain a node group set includes:
[0071] S1, determining the sum of weight values according to the weight values of each data storage node in the data storage node set;
[0072] S2, determining the target overall weight value corresponding to each of the multiple node groups according to the sum of the target group number and the weight value;
[0073] S3, grouping the data storage nodes in the data storage node set according to a weight balancing rule to obtain a node group set, wherein the weight balancing rule is used to ensure that the overall weight value of each node group and the target overall weight value meet a target difference condition.
[0074] In the above step S1, the above weight value of the above data storage node is used to indicate the availability and storage capacity of the node, and can be comprehensively evaluated based on multiple factors such as the node's storage capacity, processing power, network bandwidth, etc.
[0075] Further in step S2, the target overall weight value corresponding to each of the above multiple node groups can be the ratio of the sum of the weight values and the target number of groups, that is, the standard weight of each group w b =Total weight / N.
[0076] In the above step S3, the storage nodes are grouped according to the weight balancing rule to ensure that the overall weight value of each group is consistent with the target overall weight value w b As close as possible, satisfying the target difference condition, which may be that the difference in the overall weight value of any node group is 0, or the absolute value of the difference is less than a preset threshold.
[0077] As an optional implementation, the number of groups is determined to be N. The sum of the node weights in each group is evenly distributed, ensuring subsequent computing efficiency and load balancing. Calculate the standard weight w for each group b = total weight / N, where total weight is the sum of all node weights and N is the number of groups. Nodes are grouped gradually according to their weights. During the grouping process, the sum of the weights of each group should be close to w. b To avoid some groups with too large or too small weights, which will affect the subsequent load balancing. While ensuring that each group has at least one node, try to keep the weights of each group balanced. If the weight of a group is lower than the standard weight w b , then the node will be added to the group first until all nodes are assigned. The node grouping steps are as follows:
[0078] The first step is to calculate the total weight / number of groups to get the group standard weight w b ; Step 2, for group i to be filled, node j, if the current group (w g (i) <w b &left_nodes(number of remaining nodes)>left_groups(number of remaining groups)) or left_nodes=left_groups, then include the node in the group and record w g (i)=W(i)+W(j), which is the total weight w of the current group i g (i) Less than the target weight w b , and there are still remaining nodes to be allocated (left_nodes>left_groups); or, when the number of remaining nodes is equal to the number of remaining groups (left_nodes=left_groups), even if the total weight of the current group exceeds the target value, it is necessary to continue filling to ensure that each group has at least one node; loop through the second step until all node groupings are completed.
[0079] In an optional embodiment, determining a target data storage node that matches the data object to be stored in the target node group according to a weight selection condition includes:
[0080] S1, determining a first parameter according to an identifier of a data group to be stored, a disturbance factor parameter value, and a node sequence number of a data storage node;
[0081] S2, determining a selection node weight value of the data storage node according to the product of the first parameter and the weight value of the data storage node;
[0082] S3, determine and select the data storage node with the largest node weight value as the target data storage node.
[0083] In the above step S1, the above first parameter is, for example, the calculation result of the formula hash(x,r,i), where the data group identifier x, the disturbance factor r and the node number i, each node has its own NRD, and the position (service location attribute identifier) of the NRD is the node number;
[0084] Then, in step S2, the selection node weight value of the data storage node is determined according to the product of the first parameter and the weight value of the data storage node. The selection node weight value can be w i hash(x,r,i),w i The weighted hash value calculation method ensures that nodes with higher weights (i.e., nodes with richer resources or lower loads) have a higher probability of being selected during data distribution.
[0085] In the above step S3, the data storage node with the largest node weight value is determined to be the target data storage node, and the calculation logic is max i (w i hash(x,r,i)), that is, the product of the weight and the hash value, reflects the node's priority when storing a data group. Nodes with higher weights have greater selection weights, making them more likely to be selected as target data storage nodes. It should be noted that storage nodes within a data group can also be determined based on the storage policy of the bucket to which they belong. If the bucket is in replica mode, multiple copies of the same data group will be stored on disk. If erasure coding is used, the erasure code will be stored. Writes within the group prioritize nodes with higher weights. Alternatively, nodes with higher weights will be selected for backfilling during disaster recovery, without specific restrictions.
[0086] In an optional implementation, after storing the data object to be stored in the target data storage node, the following steps are included:
[0087] S1, determining current data storage node information of the current deployment architecture when detecting that the number of data storage nodes in the deployment architecture meets the data object migration triggering condition;
[0088] S2, updating the node group set and the data storage nodes in each node group in the node group set according to the current data storage node information;
[0089] S3: Migrate the reference data according to the overall weight value of each node group in the updated node group set until the weight balance condition is met.
[0090] In the above step S1, the above data object migration triggering condition can be based on the increase or decrease in the number of nodes, the change in node load, or the deterioration of the node health status. For example, if the number of newly added nodes reaches a preset threshold, or the failure of a node causes an uneven data distribution, the data migration process is triggered.
[0091] The current data storage node information of the current deployment architecture can be recalculated by recalculating the node set corresponding to the data group and determining the storage point information corresponding to the data objects in the data group. This includes but is not limited to recalculating the node list for each group and adjusting the weight distribution of the groups to ensure that the node status of each group is up to date.
[0092] Then, in step S2, the existing set of node groups can be updated based on the latest node status and weights. For example, if a new node is added, it will be added to the appropriate group to maintain the best possible weight balance between groups. If a node fails or is removed, the node list of its group will be updated, and the system will re-evaluate the weights of other nodes to fill the gap left by the missing node, ensuring that data is correctly redistributed.
[0093] In the above step S3, the above parameter data can be the data that needs to be adjusted after comparing the original mapping relationship with the updated mapping relationship. The migration process includes but is not limited to calculating the difference between the overall weight value and the ideal weight value (usually based on the average weight) of each node group to quantify the degree of imbalance between the groups, and then migrate from the high-load node group to the low-load node group according to the real-time status and weight value of the data storage node until the overall weight value of all node groups meets the pre-set weight balance condition; the above weight balance condition may mean that the overall weight values of all node groups are within a specific range, such as not exceeding ±10% of the average value.
[0094] This application uses a pseudo-random algorithm to evenly distribute data to each storage node to ensure load balancing and high availability. When a storage node fails, the storage content of the node needs to be redistributed to other healthy nodes. According to the DHH algorithm, the data migration process will be achieved by dynamically adjusting the node weights and recalculating the hash values to ensure that as little data as possible needs to be moved, reducing the impact on system performance. When the node is overloaded, the nodes with lower loads will take priority to take on more data. After the data hash is recalculated, the data will be migrated to the nodes with lower loads. This process will be re-evaluated based on the node weights to ensure that the data is evenly distributed and avoid being concentrated on a few nodes as much as possible. For the addition and deletion of nodes, data migration and replica redistribution will be dynamically adjusted according to the new node mapping. When a node is added, the new node will participate in the hash calculation according to the weight and serve as the new target node for data storage; when a node is deleted, the system will redistribute the data according to the replica rules to ensure data consistency and availability, while minimizing data migration caused by node deletion.
[0095] In an optional embodiment, when it is detected that the number of data storage nodes in the deployment architecture meets the data object migration triggering condition, determining the current data storage node information of the current deployment architecture includes at least one of the following:
[0096] Method 1: When the number of data storage nodes whose node load is greater than the target load threshold is greater than a first number, determining that the data object migration trigger condition is met, and obtaining current data storage node information of the current deployment architecture;
[0097] Optionally, the load of all storage nodes is monitored in real time, including indicators such as CPU usage, disk I / O, and memory usage; if the load of more than a preset first number of nodes is higher than the target load threshold, the data migration mechanism is immediately started; detailed information on the current data storage nodes in the current deployment architecture is collected, including their load conditions, storage capacity, and grouping, etc. Based on the node load information, part of the data is migrated from the high-load node to the low-load node to rebalance the load of each node.
[0098] Method 2: When the difference between the number of nodes of the current data storage node and the number of nodes of the original data storage node in the deployment architecture is greater than a first threshold, determining that the data object migration trigger condition is met, and obtaining the current data storage node information of the current deployment architecture;
[0099] Optionally, the number of nodes of the current data storage node and the number of nodes in the original deployment architecture are regularly checked, and the difference between the two is calculated. If the difference in the number of nodes exceeds a preset first threshold, it is determined that the current node configuration is not suitable for the existing data distribution and needs to be readjusted through data migration. The latest status information of all data storage nodes is collected, including details of the newly added nodes and information of offline nodes, to provide the necessary context for the data migration process; according to the new node layout, the data distribution is reconstructed by migrating data to the newly added nodes or transferring data from offline nodes.
[0100] Method three: when the difference between the overall weight values of any two node groups in the deployment architecture is greater than the second threshold, it is determined that the data object migration trigger condition is met, and the current data storage node information of the current deployment architecture is obtained.
[0101] Optionally, the overall weight value of each node group in the node group set is analyzed regularly. The overall weight value is the accumulation of the weight values of all nodes in the group. If the weight difference between any two node groups exceeds the preset second threshold, the system determines that the current data distribution strategy is no longer applicable, and then needs to collect information on all data storage nodes under the current deployment architecture, including the node group to which the node belongs, the current weight value and health status. According to the overall weight difference of the node group, choose to migrate data from the high-weight group to the low-weight group until the weight difference of all node groups is within the preset second threshold, thereby achieving load balancing between node groups.
[0102] Before data migration, the recalculated hierarchical mapping is compared with the old hierarchical mapping, and the different migration data at the data group placement location is compared. The data is first made redundant. Only after the migration is complete is the old data redundancy deleted to ensure that data is not lost. During data migration, the DHH algorithm minimizes the scope of data movement and only remaps the affected nodes and their subnodes, avoiding performance bottlenecks caused by global data redistribution. By adjusting the hierarchical mapping nodes and weights, the system can adaptively handle node changes, ensuring that data can always be accessed quickly and reliably, while effectively reducing system load and improving system scalability and fault tolerance.
[0103] The following describes this application in a complete implementation manner:
[0104] Part 1, system deployment. The distributed object storage system of the present invention adopts a hierarchical tree structure, and the system levels include: row (Row), rack (Rack), host (Host), and storage node (Store). In the system design, the management of each level is the responsibility of NRD (Node Running Daemon). The NRD node is responsible for maintaining hierarchical topology information, ensuring high availability of nodes within the level, synchronizing status information in real time, and ensuring data consistency. Each node (such as rack NRD, host NRD, storage node NRD) synchronizes data through a consistency protocol to ensure that the topology information between nodes is always up to date.
[0105] Specifically, the row layer represents a row of cabinets in a computer room, with eight rows in a room and eight racks in each row. The rack layer represents racks within a data center, with each rack containing eight hosts. The host layer represents each server and is responsible for managing the eight underlying storage nodes. Storage nodes, located at the bottom layer, are responsible for data storage, monitoring media health, adjusting storage space, and providing data read and write services. Leaf nodes (storage node NRDs) are also responsible for managing storage media, including disk space planning and storage defragmentation. Based on this approach, a distributed, decentralized, and layered object storage system is deployed.
[0106] Each NRD node not only maintains topology information, but also performs fault detection and repair. Once a node (such as a storage node, host, rack, etc.) is found to have a fault, NRD will immediately issue an alarm and start the automatic recovery mechanism. During the recovery process, NRD will redistribute tasks and data to ensure the high availability of data in the system. After the system is started, NRD will regularly communicate with other NRD nodes at the same level to ensure that the topology information of all nodes is consistent. When a topology change occurs at a certain level (such as a row or rack), the NRD node will trigger the data synchronization mechanism of the relevant level. In order to avoid uneven load between nodes, the system has designed a dynamic load balancing mechanism. NRD will monitor the load of each storage node in real time and adjust the data storage location according to the load situation to avoid excessive load on some storage nodes.
[0107] The second part involves writing data. Before storing data, create a storage bucket within the system. This bucket will be used to store business data. Configure the number of replicas and the initial number of node groups. Specifically, create a new storage bucket for video data (video), configure three replicas, divide the fault domain into different rows (ideally, simultaneous loss of connection, disconnection, and low probability of failure are preferred), and assign each row a different power supply to minimize the probability of loss of connection.
[0108] Based on the storage bucket, number of replicas, and fault domain division, the following client configuration placement rules are generated. For example, "the data belongs to the video storage bucket, select three different row-level nodes, and finally select a storage node under each selected row node to place a data replica, implementing a three-copy redundant storage strategy while ensuring that data is evenly distributed across different levels and locations."
[0109] Then the client can address according to the placement rules, hierarchical mapping and local algorithm library. If the object to be placed is identified as x, the calculation should be placed in the node group x', and calculate which three locations x' will be mapped to. First, starting from the root hierarchical node, after obtaining the hierarchical mapping information, use the DHH algorithm to calculate the storage location of the data locally, and then continue to communicate with the lower nodes to continue to obtain the hierarchical mapping of the lower layer. Until the leaf NRD node store layer node, communicate with the store layer node and send the data to the selected NRD ordered set. The final choice is that the node group x' should be placed on 1 store node in each of the 3 rows to store 3 copies of data. The possible locations are NRD1: row = 1, rack = 2, host = 5, store = 3; NRD2: row = 3, rack = 4, host = 1, store = 1; NRD3: row = 5, rack = 6, host = 7, store = 8;
[0110] Data is distributed to different storage nodes. Each storage node writes data locally and records metadata. Write operations are synchronized through NRD nodes to ensure data consistency across different storage nodes.
[0111] The third part involves data reading. Addressing is done through placement rules, hierarchical mapping, and the local algorithm library. If the object identifier is x and the calculated placement group is x', the placement rules, hierarchical mapping, and local algorithm library remain unchanged. The same three addresses can be requested using the same data storage method to obtain the resulting ordered set. Accessing any NRD allows interactive data reading.
[0112] The fourth part is mapping changes and data migration. During the operation of the system, nodes such as racks or hosts may change (such as addition, deletion, failure, etc.). NRD nodes will re-evaluate the data distribution in groups based on weights and adjust the data storage location as needed. When it is determined that data migration needs to be initiated and the data storage location will change, NRD initiates the data migration operation to migrate the data from the source storage node to the new storage node. During the migration process, the system will ensure data consistency by using the hierarchical mapping of old and new snapshots and appropriate migration redundancy, and delete the old data after the migration is confirmed, so that data loss or access interruption will not occur. The system will start the rebalancing mechanism regularly or according to load changes to automatically adjust the data distribution and reduce the uneven load between nodes.
[0113] It's also important to note that the client module is a key component tightly integrated with the distributed storage system, responsible for the efficient execution of data access operations. This module relies on a pre-defined algorithm library and an ordered list of top-level NRDs (node runtime daemons). Through precise calculation and allocation strategies, it ensures even distribution and efficient storage of data within the storage system. The client module's workflow primarily involves processing data storage requests, obtaining hierarchical mappings, and calculating data storage paths.
[0114] Data placement request and placement rule configuration: When a client initiates a data placement request, the storage task is first configured according to the defined placement rules. Placement rules typically include, but are not limited to, parameters such as the replica type (e.g., replicas or erasure codes), the number of replicas, and the size of the storage system. The client module uses these rules to determine the data allocation strategy, ensuring that data replicas or erasure codes are distributed appropriately across the storage tiers.
[0115] Hierarchical Mapping Acquisition and Node Selection: Based on the configured placement rules, the client module acquires the hierarchical mapping of each layer in a top-down manner, starting from the top-level NRD. The hierarchical mapping of each layer contains information such as the health, storage capacity, weight, and number of nodes of each node in the current layer. This data provides basic data support for the calculation of the data storage path. The client selects nodes by executing the Distributed Hierarchical Hash Algorithm (DHH) and gradually locates the final storage node. The execution of the DHH algorithm ensures that data is stored on healthy and available storage nodes, avoiding data loss or performance bottlenecks caused by node failure or overload.
[0116] Node Health and Fault Tolerance: During actual data storage, the client module also considers node health. When requesting a hierarchical mapping, if the request fails due to network issues or node failure, the client module uses an adaptive retry mechanism to select another node in the same hierarchy to retrieve the hierarchical mapping. This process ensures the efficient operation of the decentralized hierarchical structure and prevents single points of failure from impacting the stability of the overall storage system. The client module implements a NRD backup mechanism between hierarchical nodes to achieve high availability and fault tolerance for the storage system, ensuring that data is always stored on healthy and appropriately loaded nodes.
[0117] Storage Path Calculation and Data Distribution: The client module determines the final data storage path based on hierarchical mapping information through multiple hash calculations and node selection. During data storage, the system manages the health, load, and storage capacity of each storage node and adjusts data storage based on these assessments, achieving load balancing and preventing certain nodes from becoming performance bottlenecks due to overuse. Furthermore, a flexible replica management mechanism ensures redundant data storage across multiple fault domains, enhancing the storage system's fault tolerance.
[0118] Decentralization and scalability: The design of the client module fully reflects the advantages of decentralized distributed storage. By obtaining the hierarchical mapping layer by layer and executing the DHH algorithm, the client can complete the calculation of the data storage path without relying on a single central node. This decentralized architecture effectively avoids the performance bottlenecks and single point failure problems caused by centralized metadata management in traditional storage systems, ensuring that the system can still operate efficiently in the face of node expansion, fault recovery or load fluctuations. When the system expands, the client module can dynamically adapt to the addition of new nodes and automatically update the hierarchical mapping through interaction with other NRD nodes to ensure that new nodes can be quickly integrated into the storage system, reducing the overhead of data migration and maintaining the uniformity of data storage. This dynamic adaptability of the client module makes the entire system more flexible and scalable, and is particularly suitable for large-scale distributed storage environments.
[0119] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0120] According to another aspect of the embodiment of the present application, a data object storage device for implementing the above-mentioned data object storage method is also provided. Figure 9 As shown, the device includes:
[0121] An acquiring unit 902 acquires at least one data group to be stored and performs hash processing on an identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored;
[0122] A first determining unit 904 determines, based on a hash processing result of the data group to be stored, a target node group that matches the data group to be stored in the node group set, wherein each of the plurality of node groups in the node group set includes at least one data storage node;
[0123] A second determining unit 906 determines a target data storage node that matches the data object to be stored in the target node group according to a weight selection condition, wherein the weight selection condition is determined according to a weight value of the data storage node in the node group;
[0124] The storage unit 908 stores the data object to be stored in the target data storage node.
[0125] Optionally, the above-mentioned first determination unit includes a third determination module, which is used to determine at least one of the data storage nodes in the target layer of the deployment architecture according to the placement rules in the initialization information, wherein information synchronization is achieved between the various layers in the deployment architecture through a target detection process, and the placement rules are used to indicate the storage method of the data object; the data storage nodes that pass the node status detection are stored in the data storage node set.
[0126] Optionally, the above-mentioned third determination module includes a grouping module, which is used to group the data storage node set according to the target grouping number to obtain the node group set; perform modulo processing on the hash value of the data to be stored according to the target grouping number to obtain a modulo result value; determine the target node group according to the modulo result value, wherein the serial number value of the target node group is consistent with the modulo result value.
[0127] Optionally, the above-mentioned grouping module is also used to determine the sum of the weight values based on the weight values of each data storage node in the data storage node set; determine the target overall weight value corresponding to each of the multiple node groups based on the target number of groups and the sum of the weight values; group the data storage nodes in the data storage node set according to the weight balancing rule to obtain the node group set, wherein the weight balancing rule is used to make the overall weight value of each node group and the target overall weight value meet the target difference condition.
[0128] Optionally, the above-mentioned second determination unit is also used to determine the first parameter based on the identifier of the data group to be stored, the disturbance factor parameter value and the node serial number of the data storage node; determine the selection node weight value of the data storage node based on the product of the first parameter and the weight value of the data storage node; and determine the data storage node with the largest selection node weight value as the target data storage node.
[0129] Optionally, the above-mentioned storage unit includes a detection module for determining the current data storage node information of the current deployment architecture when it is detected that the number of data storage nodes in the deployment architecture meets the data object migration trigger condition; updating the node group set and the data storage nodes in each node group in the node group set according to the current data storage node information; and migrating the reference data according to the overall weight value of each node group in the node group set after the update until the weight balance condition is met.
[0130] Optionally, the above-mentioned detection module includes a migration module, which is used to determine that the data object migration trigger condition is met and obtain the current data storage node information of the current deployment architecture when the number of data storage nodes whose node load is greater than the target load threshold is greater than a first number; determine that the data object migration trigger condition is met and obtain the current data storage node information of the current deployment architecture when the difference between the number of nodes of the current data storage nodes and the number of nodes of the original data storage nodes in the deployment architecture is greater than a first threshold; determine that the data object migration trigger condition is met and obtain the current data storage node information of the current deployment architecture when the difference between the overall weight values of any two node groups in the deployment architecture is greater than a second threshold.
[0131] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned method for storing data objects when running.
[0132] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned data object storage method is also provided. The electronic device may be Figure 1 The terminal device or server shown in the figure. This embodiment is described by taking the electronic device as a mobile phone or a computer as an example. Figure 10 As shown, the electronic device includes a memory 1002 and a processor 1004. The memory 1002 stores a computer program, and the processor 1004 is configured to execute the steps in any of the above method embodiments through the computer program.
[0133] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0134] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0135] S1, obtaining at least one data group to be stored, and performing hash processing on an identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored;
[0136] S2, determining a target node group that matches the data group to be stored in a node group set according to a hash processing result of the data group to be stored, wherein each of the plurality of node groups in the node group set includes at least one data storage node;
[0137] S3, determining a target data storage node that matches the data object to be stored in the target node group according to a weight selection condition, wherein the weight selection condition is determined according to a weight value of the data storage node in the node group;
[0138] S4: Store the data object to be stored in the target data storage node.
[0139] Alternatively, those skilled in the art will appreciate that Figure 10 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 10 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 10 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 10 Different configurations shown.
[0140] Among them, the memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the storage method and device of the data object in the embodiment of the present application. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, realizes the above-mentioned storage method of the data object. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include a memory remotely located relative to the processor 1004, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, corporate intranet, local area network, mobile communication network and combinations thereof. As an example, Figure 10As shown, the memory 1002 may include, but is not limited to, the acquisition unit 902, the first determination unit 904, the second determination unit 906, and the storage unit 908 in the storage device of the data object. In addition, it may also include, but is not limited to, other module units in the storage device of the data object, which will not be repeated in this example.
[0141] Optionally, the transmission device 1006 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1006 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1006 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0142] In addition, the electronic device further includes: a display 1008 and a connection bus 1010 for connecting various module components in the electronic device.
[0143] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a point-to-point network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the point-to-point network.
[0144] According to one aspect of the present application, a computer-readable storage medium is provided. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the methods provided in the various optional implementations described above.
[0145] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0146] S1, obtaining at least one data group to be stored, and performing hash processing on an identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored;
[0147] S2, determining a target node group that matches the data group to be stored in a node group set according to a hash processing result of the data group to be stored, wherein each of the plurality of node groups in the node group set includes at least one data storage node;
[0148] S3, determining a target data storage node that matches the data object to be stored in the target node group according to a weight selection condition, wherein the weight selection condition is determined according to a weight value of the data storage node in the node group;
[0149] S4: Store the data object to be stored in the target data storage node.
[0150] Alternatively, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0151] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0152] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.
[0153] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0154] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0155] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0156] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0157] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for storing a data object, characterized in that: include: Acquire at least one data group to be stored, and perform hash processing on an identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored; Determining, according to a hash processing result of the data group to be stored, a target node group that matches the data group to be stored in a node group set, wherein each of the plurality of node groups in the node group set includes at least one data storage node; Determining a target data storage node that matches the data object to be stored in the target node group according to a weight selection condition, wherein the weight selection condition is determined according to a weight value of the data storage node in the node group; The data object to be stored is stored in the target data storage node.
2. The method according to claim 1, characterized in that Before determining a target node group matching the to-be-stored data group in a node group set based on a hash processing result of the to-be-stored data group, the method includes: determining at least one data storage node in a target level of a deployment architecture according to a placement rule in the initialization information, wherein information synchronization is achieved between levels of the deployment architecture via a target detection process, the placement rule being used to indicate a storage method for a data object; The data storage nodes that pass the node status detection are stored in a data storage node set.
3. The method according to claim 2, characterized in that After determining, in a node group set, a target node group matching the data group to be stored based on the hash processing result of the data group to be stored, the method further comprises: Grouping the data storage node set according to the target grouping number to obtain the node group set; Performing a modulo process on the hash value of the data to be stored according to the target number of groups to obtain a modulo result value; The target node group is determined according to the modulo result value, wherein the sequence number value of the target node group is consistent with the modulo result value.
4. The method according to claim 3, characterized in that The grouping the data storage node set according to the target grouping number to obtain the node group set includes: Determine the sum of weight values according to the weight values of the respective data storage nodes in the data storage node set; Determine the target overall weight value corresponding to each of the plurality of node groups according to the sum of the target number of groups and the weight value; The data storage nodes in the data storage node set are grouped according to a weight balancing rule to obtain the node group set, wherein the weight balancing rule is used to ensure that the overall weight value of each node group and the target overall weight value meet a target difference condition.
5. The method according to claim 3, characterized in that Determining a target data storage node that matches the data object to be stored in the target node group according to a weight selection condition includes: Determine a first parameter according to the identifier of the data group to be stored, a disturbance factor parameter value, and a node sequence number of the data storage node; determining a selection node weight value of the data storage node according to a product of the first parameter and the weight value of the data storage node; The data storage node with the largest selected node weight value is determined as the target data storage node.
6. The method according to claim 1, characterized in that After storing the data object to be stored in the target data storage node, the method includes: When it is detected that the number of the data storage nodes in the deployment architecture meets the data object migration triggering condition, determining the current data storage node information of the current deployment architecture; Update the node group set and the data storage node in each node group in the node group set according to the current data storage node information; The reference data is migrated according to the updated overall weight value of each node group in the node group set until a weight balance condition is met.
7. The method according to claim 6, characterized in that The determining of the current data storage node information of the current deployment architecture when it is detected that the number of the data storage nodes in the deployment architecture meets the data object migration triggering condition includes at least one of the following: When the number of the data storage nodes whose node load is greater than the target load threshold is greater than a first number, determining that the data object migration trigger condition is satisfied, and obtaining the current data storage node information of the current deployment architecture; When the difference between the number of nodes of the current data storage node and the number of nodes of the original data storage node in the deployment architecture is greater than a first threshold, determining that the data object migration trigger condition is satisfied, and obtaining the current data storage node information of the current deployment architecture; When the difference between the overall weight values of any two node groups in the deployment architecture is greater than a second threshold, it is determined that the data object migration trigger condition is met, and the current data storage node information of the current deployment architecture is obtained.
8. A storage device for a data object, characterized in that: include: an acquiring unit, acquiring at least one data group to be stored, and performing hash processing on an identifier of the data group to be stored, wherein the data object to be stored is a data object in the data group to be stored; a first determining unit, configured to determine, in a node group set, a target node group that matches the data group to be stored based on a hash processing result of the data group to be stored, wherein each of the plurality of node groups in the node group set includes at least one data storage node; a second determining unit, configured to determine a target data storage node in the target node group that matches the data object to be stored according to a weight selection condition, wherein the weight selection condition is determined according to a weight value of the data storage node in the node group; The storage unit stores the data object to be stored in the target data storage node.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program is executed by a processor to perform the method according to any one of claims 1 to 7.
10. An electronic device comprising: A memory and a processor, wherein the memory is used to store a computer program, and wherein the processor is used to execute the steps of the method according to any one of claims 1 to 7 when calling the computer program.
Citation Information
Patent Citations
Object storage data distribution mechanism based on two-sage Hash
CN103905540A
Data storage method and system and data query method and system
CN108334551A
A method and apparatus for storing data
CN109542352A
Data storage method and device based on blockchain network, storage medium and equipment
CN111049902A