Resource allocation and management system for data fragments
By using heartbeat leases and bridging leases, combined with restart cooldown control, the system solves the problems of service interruption caused by system node changes and sharding migration during application restarts, achieving high availability and stability of sharding and optimizing system performance and resource utilization.
Patent Information
- Application Number
- CN202510787123.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies can cause service interruptions and performance degradation when system nodes change (such as expansion, reduction, or failure), as well as system fluctuations and performance degradation caused by unnecessary shard migration during application restarts.
By detecting node status through heartbeat leases, creating bridging leases to trigger shard migration, and prioritizing the restoration of the original shard allocation when the application restarts, combined with the limitation of shard migration during the restart cooldown period, a distributed coordination service is used to manage lease lifecycles and shard change events to achieve high availability of shards.
It improves system stability and availability, reduces service downtime, optimizes shard location query latency, reduces computing resource consumption, enhances resource elasticity and scalability, and reduces the need for manual intervention.
Smart Images

Figure CN120892174A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data sharding, and particularly relates to a resource allocation and management system for data sharding. BACKGROUND
[0002] In a distributed system, data sharding is a commonly used horizontal expansion strategy, which divides a large data set into multiple smaller parts and distributes them on different server nodes to improve the throughput and availability of the system. Existing sharding management techniques can dynamically adjust the sharding allocation according to the application load, but when the system nodes change (such as expansion, contraction or failure), the sharding redistribution process may cause service interruption or performance degradation, and when the application restarts, it often triggers a large number of unnecessary sharding migration, causing system fluctuations and performance degradation. SUMMARY
[0003] The present application provides a resource allocation and management system for data sharding to solve the defects that traditional data sharding may cause service interruption or performance degradation when the system is abnormal, and the poor stability of application restart.
[0004] The present application provides a resource allocation and management system for data sharding, comprising:
[0005] A plurality of worker nodes, wherein each worker node carries the actual processing logic of the shard and periodically reports state information to the shard manager through a heartbeat lease;
[0006] A shard manager for dynamically detecting node failure, expansion or contraction events based on the heartbeat state of each worker node; when the node state changes, triggering shard migration by creating a bridge lease independent of the worker node heartbeat lease; and when the application restarts, preferentially restoring the original shard allocation and limiting shard migration through a restart cooling period;
[0007] A distributed coordination service for managing the life cycle of the heartbeat lease and the bridge lease and publishing shard change events.
[0008] According to the resource allocation and management system for data sharding provided by the present application, the shard manager triggers shard migration by creating a bridge lease independent of the worker node heartbeat lease, comprising:
[0009] Creating a temporary bridge lease with a first validity period;
[0010] Notifying the source node to migrate the shard to be migrated to a bridge state;
[0011] Issuing a new daemon lease for the target node with a second validity period, which is less than the first validity period;
[0012] The shard allocation information is distributed to the target node through an HTTP request, and after the target node confirms to take over, the data ownership of the shard of the source node is released.
[0013] According to the data shard resource allocation and management system provided by the application, the distributed coordination service manages the life cycle of the heartbeat lease, including:
[0014] When the worker node starts, the heartbeat lease is applied to the distributed coordination service, and the heartbeat lease contains a unique identifier and a validity period;
[0015] The worker node prolongs the validity period of the heartbeat lease by sending heartbeat information regularly.
[0016] If the worker node fails to renew, the heartbeat lease will automatically expire, triggering shard reallocation.
[0017] According to the data shard resource allocation and management system provided by the application, the lease validity period is 10-30 seconds, and is dynamically adjusted according to system size and network conditions.
[0018] According to the data shard resource allocation and management system provided by the application, each worker node is also used to maintain a local cache of shard location, and the local cache is updated by listening to the shard change event notification published by the distributed coordination service; the shard location maintained by the worker node is bound to the heartbeat lease, and when the heartbeat lease expires, the shard location is automatically invalidated.
[0019] According to the data shard resource allocation and management system provided by the application, the shard manager is also used to query the shard location, including:
[0020] Merge multiple shard location query requests in a short period of time;
[0021] And, all shard information of the same application is obtained at a time by using the prefix query mechanism.
[0022] According to the data shard resource allocation and management system provided by the application, the shard manager limits shard migration through a restart cooling period, including:
[0023] Set a restart cooling period, and limit shard reallocation within the restart cooling period;
[0024] After the restart cooling period, the number of shard migrations within a unit time is limited, and the migration frequency of the same shard is limited.
[0025] The data sharding resource allocation and management system provided by the application, the shard manager is further used for calculating the priority of the shard based on multiple dimensions, determining the shard migration sequence according to the priority, and the multiple dimensions include service importance, real-time load pressure, data consistency requirement and client dependency amount.
[0026] The data sharding resource allocation and management system provided by the application, the shard manager adopts a leader election mechanism, the leader election mechanism includes presetting a specific key as a leader lock, all shard manager nodes attempt to create a key-value pair with a lease on the specific key, only the first node that successfully writes can obtain the leadership, and other nodes will receive a response of creation failure.
[0027] The data sharding resource allocation and management system provided by the application further includes a tenant isolation module, the tenant isolation module is used for allocating a unique key pair for each application; requiring digital signature verification in a shard operation request; and rejecting a shard migration instruction that fails the key check.
[0028] The data sharding resource allocation and management system provided by the application includes multiple worker nodes, each of which carries shard actual processing logic and periodically reports state information to the shard manager through a heartbeat lease; the shard manager is used for dynamically detecting node failure, expansion or contraction events based on the heartbeat state of each worker node; when the node state changes, a bridge lease independent of the heartbeat lease of the worker node is created to trigger shard migration; and when an application restarts, the original shard allocation is preferentially restored, and shard migration is limited through a restart cooling period; a distributed coordination service is used for managing the life cycle of the heartbeat lease and the bridge lease and publishing a shard change event, realizing shard high availability, improving the ability of the distributed coordination service through shard dynamic scheduling, and being particularly suitable for scenarios such as delay queues, service orchestration worker nodes and other scenarios that require accurate shard dynamic allocation. BRIEF DESCRIPTION OF DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0030] Figure 1 is a functional structure diagram of the data sharding resource allocation and management system provided by the application;
[0031] Figure 2 is a service end architecture diagram provided by the application;
[0032] Figure 3 is a schematic diagram of a client architecture provided by the present application. DETAILED DESCRIPTION
[0033] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in detail with reference to the drawings of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0034] Figure 1 A functional structure diagram of a data sharding resource allocation and management system provided by an embodiment of the present application is shown in FIG. 1, and the data sharding resource allocation and management system provided by the embodiment of the present application includes: Figure 1
[0035] a plurality of worker nodes 101, wherein each worker node carries a sharding actual processing logic and periodically reports state information to a sharding manager through a heartbeat lease;
[0036] In the embodiment of the present application, the heartbeat lease (Heartbeat Lease) confirms the node survival state through a periodic renewal manner, and the sharding state management mechanism based on the heartbeat lease realizes the fast detection of node failure through the lease automatic expiration feature, greatly shortens the node failure detection time, accelerates the start of sharding migration, reduces the service interruption time, and improves the overall availability of the system.
[0037] a sharding manager 102, configured to dynamically detect node failure, capacity expansion or capacity reduction events based on the heartbeat state of each worker node, trigger sharding migration through a bridge lease independent of the heartbeat lease of the worker node when the node state changes, and preferentially restore the original sharding allocation when an application is restarted and limit sharding migration through a restart cooling period;
[0038] a distributed coordination service 103, configured to manage the life cycle of the heartbeat lease and the bridge lease and publish a sharding change event.
[0039] In the embodiment of the present application, in the sharding resource allocation and management method, initial sharding allocation is performed first, specifically including: when the sharding manager is started, checking the existing sharding allocation state; for unallocated sharding, performing initial allocation according to the load condition of the current active worker node; writing the allocation result into the distributed coordination service and notifying the related worker node.
[0040] After the initialization allocation is completed, the sharding state monitoring is entered, specifically including: the sharding manager listens to the sharding state change in the distributed coordination service through the Watch mechanism; the worker node regularly reports the sharding load information (such as CPU, memory usage, request volume, etc.); the sharding manager analyzes the load data and identifies potential hot sharding.
[0041] When the trigger condition is met, dynamic sharding scheduling is performed, the trigger condition includes: node failure (lease expiration), load imbalance, system expansion and contraction; the scheduling strategy includes: comprehensively considering load balancing, migration cost and system stability; the dynamic sharding scheduling execution process includes: calculating the optimal sharding allocation scheme, and safely migrating the sharding through the bridge lease (Bridge Lease) mechanism.
[0042] The traditional sharding management technology can dynamically adjust the sharding allocation according to the application load, but when the system node changes (such as expansion, contraction or failure), the sharding reallocation process may cause service interruption or performance degradation, and when the application restarts, a large number of unnecessary sharding migrations are often triggered, causing system fluctuations and performance degradation.
[0043] The resource allocation and management system for data sharding provided by the embodiment of the application comprises a plurality of worker nodes, wherein each worker node carries actual processing logic of a shard and periodically reports state information to a sharding manager through a heartbeat lease; the sharding manager is used for dynamically detecting node failure, expansion or contraction events based on the heartbeat state of each worker node; when the node state changes, a bridge lease independent of the heartbeat lease of the worker node is created to trigger sharding migration; and when the application restarts, the original sharding allocation is preferentially restored, and sharding migration is limited through a restart cooling period; a distributed coordination service is used for managing the life cycle of the heartbeat lease and the bridge lease and publishing a sharding change event, realizing high availability of the shard, improving the capability of the distributed coordination service through dynamic scheduling of the shard, and being particularly suitable for scenarios such as a delay queue, service orchestration worker nodes and the like that require accurate dynamic allocation of sharding.
[0044] Based on any of the above embodiments, the distributed coordination service manages the life cycle of the heartbeat lease, including:
[0045] The worker node applies for a heartbeat lease from the distributed coordination service when starting, and the heartbeat lease contains a unique identifier and a validity period;
[0046] The worker node renews the heartbeat lease by regularly sending heartbeat information, thereby prolonging the validity period of the heartbeat lease;
[0047] If the worker node fails to renew, the heartbeat lease automatically expires, triggering sharding reallocation.
[0048] In the embodiment of the present application, the lease validity period is 10-30 seconds, and is dynamically adjusted according to the system scale and network condition.
[0049] In the embodiment of the present application, short-term lease (10-30 seconds) is set to quickly detect node failure while ensuring system stability, and the lease duration is dynamically adjusted according to the system scale and network condition to balance availability and performance overhead.
[0050] Based on any of the above embodiments, the shard manager triggers shard migration by creating a bridge lease independent of the worker node heartbeat lease, comprising:
[0051] A temporary bridge lease is created, and a first validity period is set;
[0052] The source node is notified to migrate the shard to be migrated to a bridge state;
[0053] A new guard lease is issued for the target node, and a second validity period is set, which is less than the first validity period;
[0054] The shard allocation information is issued to the target node through an HTTP request, and after the target node confirms to take over, the data ownership of the source node shard is released.
[0055] In the embodiment of the present application, the first validity period is usually 30 seconds; after waiting for the source node to confirm that the migration is completed or the original lease expires, a new guard lease (Guard Lease) is issued for the target node, and the second validity period is usually 15 seconds; a reaction time (about 5 seconds) is reserved for the worker node to process the new lease event; the shard manager issues the shard allocation information to the target node through an HTTP request.
[0056] In the embodiment of the present application, the shard manager is further configured to calculate the priority of the shard based on multiple dimensions, determine the shard migration sequence according to the priority, and the multiple dimensions include business importance, real-time load pressure, data consistency requirement, and client dependency amount.
[0057] In the embodiment of the present application, the shard is safely migrated through the bridge lease (Bridge Lease), and service continuity is ensured during the shard migration process.
[0058] Based on any of the above embodiments, each worker node is further configured to maintain a local cache of shard locations, and the local cache is updated by listening to shard change event notifications published by a distributed coordination service; the shard locations maintained by the worker node are bound to the heartbeat lease, and the shard locations are automatically invalidated when the heartbeat lease expires.
[0059] Based on any of the above embodiments, the shard manager is further configured to query the shard location, comprising:
[0060] merge multiple shard location query requests in a short time;
[0061] and acquire all shard information of the same application at one time by using a prefix query mechanism.
[0062] In the embodiment of the present application, the worker node maintains a local cache of shard locations, reducing access to the distributed coordination service; the cache item is bound to a lease, and the cache item is automatically invalidated when the lease expires; the cache is updated periodically (300 milliseconds) to ensure data consistency.
[0063] The embodiment of the present application can also implement batch query optimization, merge multiple shard location query requests in a short time; acquire all shard information of the same application at one time by using a prefix query mechanism. The local cache of shard location query and the batch optimization strategy significantly improve the system response speed, significantly reduce the shard location query delay, reduce the load of the distributed coordination service, improve the overall throughput of the system, and reduce the consumption of computing resources.
[0064] In the embodiment of the present application, the transaction mechanism of the distributed coordination service is used to ensure the atomicity of the shard allocation operation; the two-phase commit mode is used for shard state change to prevent shard loss or repeated allocation; the worker node locally stores (such as BoltDB) the shard information to support fault recovery.
[0065] Based on any of the above embodiments, the shard manager limits shard migration by a restart cooling period, comprising:
[0066] setting a restart cooling period, and limiting shard reallocation during the period of the restart cooling period;
[0067] after the restart cooling period, limiting the number of shard migrations per unit time, and setting a migration frequency limit for the same shard.
[0068] In the embodiment of the present application, when the worker node is restarted, the original shard allocation is restored first; the shard manager identifies the restart event to avoid triggering unnecessary shard migration; a restart cooling period is set to suppress shard reallocation during the period. Furthermore, a gradual migration control is adopted, specifically including: limiting the number of shard migrations per unit time; setting a migration frequency limit (such as a maximum of one migration in 24 hours) for the same shard; and sorting the priority of shard migration to ensure the stability of critical shard services.
[0069] The embodiment of the present application adopts a gradual shard migration control strategy to balance the system stability and load balancing requirements.
[0070] The application embodiment applies a restart stability guarantee mechanism to avoid unnecessary shard migration and system fluctuation; automatic shard scheduling reduces the need for manual intervention, improves service stability during application restart, shortens system expansion and contraction operation time, and enhances resource elasticity and scalability.
[0071] Based on any of the above embodiments, the shard manager adopts a leader election mechanism, the leader election mechanism includes presetting a specific key as a leader lock, all shard manager nodes attempt to create a key-value pair with a lease on the specific key, only the first successful writing node can obtain leadership, and other nodes will receive a failed creation response.
[0072] The application embodiment adopts a leader (Leader) election mechanism to ensure that only one active shard manager instance exists at any time; the Leader sends a heartbeat to the distributed coordination service regularly to prove its active state; the non-Leader instance listens to the Leader state and automatically takes over when the Leader fails.
[0073] Data consistency storage can be ensured by using a leader node to manage shard allocation.
[0074] Based on any of the above embodiments, the resource allocation and management system of the data shard further includes a tenant isolation module, the tenant isolation module is used to allocate a unique key pair for each application; a digital signature verification is required in a shard operation request; and a shard migration instruction that does not pass the key check is rejected.
[0075] In the application embodiment, the working node locally stores (such as BoltDB) persistent shard information to support fault recovery. In addition, a key (Secret) needs to be applied when the application accesses the shard manager; the shard registration and operation need to carry the key verification to prevent mutual interference between tenants. The application embodiment realizes high data isolation between tenants, improves malicious operation protection capability, and supports large-scale multi-tenant parallel deployment.
[0076] In some embodiments of the application, the resource allocation and management system of the data shard further includes a clock synchronization module for controlling the clocks of the nodes of the system to be synchronized within a second-level precision; the client lease expires 2 seconds in advance to prevent inconsistency caused by slow server clock.
[0077] The data shard resource allocation and management system provided by the embodiment of the application realizes rapid detection of node failure through the lease automatic expiration feature; ensures service continuity during shard migration through the bridge lease; optimizes system performance and resource utilization, significantly reduces shard location query delay, reduces the load of the distributed coordination service, improves the overall throughput of the system, and reduces the consumption of computing resources; automated shard scheduling reduces the need for manual intervention, improves service stability during application restart, shortens system expansion and contraction operation time, and enhances resource elasticity and scalability; the shard management architecture is easy to horizontally expand, supports heterogeneous environment deployment, can be smoothly integrated with existing systems, and reduces the cost of technology migration.
[0078] The embodiment of the application also provides a distributed system architecture including a client and a server, as shown in Figure 2 The server is provided with a shard manager, a worker node and a distributed coordination service, the shard manager is a core control component of the system, is responsible for global shard allocation decision and node monitoring, and needs to be implemented on the server to ensure centralized management and global view. The worker node is a server node that carries shard processing logic and belongs to the part of the server infrastructure. The distributed coordination service is built based on an extended distributed consistency (Etcd) and needs to be deployed and maintained on the server to provide coordination services for the whole system. The client can access these information through API, but the service itself is implemented on the server.
[0079] The server is divided into a network server and a server shard node, the network server includes a service framework and processes requests through a management interface and a shard interface; the server shard node provides a shard protocol interface and includes an add / delete shard manager function.
[0080] The shard node has a leader election mechanism and a shard manager, the shard manager communicates with multiple shard managers, and the implementation principle of the shard manager includes: storing shard configuration information in a strongly consistent storage and storing shard state through a mapping memory data structure.
[0081] The server protocol includes: selecting a suitable network server, and the shard interface provides a delete / add shard manager operation.
[0082] The leader node using the strongly consistent storage manages shard allocation, realizes leader competition to ensure a single master node, and the system presets a specific key (such as " / shard-manager / leader") as a leader lock in Etcd, all shard manager nodes try to create a key-value pair with a lease on this key, only the first node that successfully writes can obtain the leadership, and other nodes will receive a response of failed creation.
[0083] The server processes shard migration events through a clock cycle and performs health checks on shards by detecting heartbeat information from strongly consistent storage. Overall, Etcd coordinates shard status to ensure high availability and data consistency.
[0084] The server and client each implement a portion of the heartbeat lease functionality. The server implements lease management, timeout judgment, and failure handling logic; the client implements periodic sending of heartbeat signals and renewal requests. As a health check mechanism for the distributed system, the two parties need to collaborate to complete the entire process.
[0085] Client architecture such as Figure 3 As shown, it includes an instruction receiving layer, used to receive instructions from the shard management server: acting as a network service receiver, it processes shard addition and deletion instructions from the server, forming the client control entry point. In the container program layer, the container process implements the sharding interface, providing a standardized shard manager framework interface and responsible for the lifecycle management of client shards.
[0086] The clock loop mechanism includes: a heartbeat module that writes heartbeat information to storage at fixed intervals, and uploads load and local shards: periodically collecting and reporting shard running status and resource utilization.
[0087] The sharding interface implements the specific logic of sharding operations according to the framework protocol, handling the actual addition and deletion of shards.
[0088] The sharding process includes:
[0089] Fragmentation protocol encapsulation: Abstracts the underlying fragmentation operations into a unified protocol, and executes add / delete logic based on the fragmentation status;
[0090] Asynchronous queue processing: By buffering sharded operation requests through a queue, the order and reliability of processing are ensured.
[0091] The write mechanism is used to persist the results of sharding operations to local storage, achieving data consistency between the asynchronous queue and local storage. Local storage persistently saves sharding status and configuration information. External storage interaction includes: central hop data nodes in storage (recording the active status of containers) and lease data nodes in storage (maintaining the sharding allocation relationship in external storage).
[0092] The monitoring and synchronization mechanism includes: an observer module that monitors changes to lease nodes in storage and transmits the change information to an asynchronous queue; a lease mechanism implementation that processes messages in the asynchronous queue, verifies the validity of shard leases, and updates the lease status in local storage; and a result synchronizer that tracks changes in local storage and writes shard configuration update information to the asynchronous queue.
[0093] The client architecture forms a complete shard management closed loop through multi-level asynchronous queues, monitoring mechanisms and local storage, ensures shard state consistency and high availability, and realizes reliable shard scheduling and management in a distributed environment.
[0094] The present application details the implementation process of resource allocation and management of data shards through the following examples, specifically including:
[0095] (1) Construct a shard log processing system, create a LogShardManager structure that implements the ShardPrimitives interface, and manage multiple LogProcessor instances;
[0096] (2) Create a LogProcessor for each log shard, including ID, application name and file path parameters, implement Start and Stop methods to control the processing flow;
[0097] LogProcessor corresponds to a worker node in the architecture. When LogProcessor starts, it applies for a unique lease from Etcd through the SM client, with an effective period of 10-30 seconds. In the Start method of LogProcessor, a heartbeat coroutine is started to send a heartbeat lease renewal request to Etcd regularly. Heartbeat renewal is implemented through Etcd's LeaseKeepAliveAPI, which renews the lease every 1 / 3 of the lease duration. The above process implements lease application and renewal.
[0098] (3) Add shard processors through the Add method of LogShardManager, get the application name and file path information from ShardSpec, create or update the processor and start it;
[0099] (4) Remove shard processors through the Drop method of LogShardManager, stop the corresponding processor and delete it from the manager;
[0100] (5) Initialize the LogShardManager instance in the main program, connect to Etcd through the SM client, and provide an HTTP interface for external management of shards;
[0101] The SM client connects to the distributed coordination service in the architecture corresponding to Etcd;
[0102] LogShardManager corresponds to the shard manager in the architecture, and the LogShardManager starts a Watch coroutine when being initialized, and listens to shard state and lease changes in the Etcd; the Watch API of the Etcd is used to listen to key value change events under a specific prefix, and the lease state is acquired in real time; the return information of the lease health state is added in the GetStats method, monitoring data is provided, and the above process realizes shard state monitoring.
[0103] (6) The state monitoring API is added, the current state of all processors is acquired through the GetStats method, and the current state is returned in the JSON format;
[0104] (7) The command line tool is realized, is used for adding and deleting log shards, and interacts with the main service through an HTTP request.
[0105] In the embodiment of the application, when the node where the LogProcessor is located fails and cannot renew, the corresponding lease expires automatically, the LogShardManager detects the lease expiration event through the Watch mechanism, identifies that the node has failed, triggers the Drop method to remove the shard processor of the failed node, and executes a shard redistribution process, creates a bridge lease (30 seconds valid period), ensures that the service does not interrupt during the shard migration process, selects a healthy node as a target node, and assigns a new shard processor through the Add method.
[0106] Each LogShardManager maintains a local cache of shard locations, reduces the query dependence on the Etcd, binds the cache to the lease, automatically cleans up the corresponding cache entry when the lease expires, periodically (300 milliseconds) synchronously updates the cache, ensures data consistency, realizes a batch query mechanism, and optimizes the efficiency of acquiring shard location information.
[0107] The LogShardManager checks the historical shard allocation record of the node when being started, preferentially recovers the original shard allocation relationship, avoids unnecessary shard migration, sets a restart cooling period, suppresses shard redistribution operations during the period, realizes gradual migration control, and limits the number of shard migrations in a unit of time.
[0108] In the embodiment of the application, the system can realize rapid fault detection and safe dynamic scheduling of shards based on the heartbeat lease, and guarantee the high availability and stability of log processing.
[0109] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data sharding resource allocation and management system, characterized in that, include: Multiple worker nodes, each of which carries the actual processing logic of the sharding and periodically reports status information to the sharding manager through heartbeat leases; The shard manager is used to dynamically detect node failures, scaling up or down events based on the heartbeat status of each worker node. When a node's state changes, shard migration is triggered by creating a bridging lease that is independent of the worker node's heartbeat lease; and when the application restarts, the original shard allocation is restored first, and shard migration is restricted through a restart cooldown period. A distributed coordination service is used to manage the lifecycle of the heartbeat leases and bridging leases, and to publish shard change events.
2. The data sharding resource allocation and management system according to claim 1, characterized in that, The shard manager triggers shard migration by creating bridging leases that are independent of worker node heartbeat leases, including: Create a temporary bridging lease and set its first validity period; The source node will be notified to migrate the shard to be migrated to the bridged state; A new guardian lease is issued to the target node, with a second validity period that is shorter than the first validity period. The shard allocation information is sent to the target node via an HTTP request. Once the target node confirms the takeover, it releases the data ownership of the source node's shard.
3. The data sharding resource allocation and management system according to claim 1, characterized in that, The distributed coordination service manages the lifecycle of the heartbeat lease, including: When a worker node starts up, it requests a heartbeat lease from the distributed coordination service. The heartbeat lease includes a unique identifier and a validity period. Work nodes extend the validity period of the heartbeat lease by periodically sending heartbeat messages to renew the lease. If a worker node fails and cannot be renewed, the heartbeat lease will automatically expire, triggering a shard reallocation.
4. The data sharding resource allocation and management system according to claim 3, characterized in that, The lease term is 10 to 30 seconds and is dynamically adjusted according to system size and network conditions.
5. The data sharding resource allocation and management system according to claim 1, characterized in that, Each worker node is also used to maintain a local cache of shard locations, which is updated by listening to shard change event notifications published by the distributed coordination service; the shard locations maintained by the worker node are bound to heartbeat leases, and the shard locations automatically become invalid when the heartbeat leases expire.
6. The data sharding resource allocation and management system according to claim 1, characterized in that, The shard manager is also used to query shard locations, including: Merge multiple shard location query requests within a short period of time; Additionally, a prefix query mechanism is used to retrieve all shard information for the same application at once.
7. The data sharding resource allocation and management system according to claim 1, characterized in that, The fragment manager restricts fragment migration through a restart cooldown period, including: Set a restart cooldown period, during which fragment reallocation is restricted; After the restart cooling-off period, the number of fragment migrations per unit time is limited, and a migration frequency limit is set for the same fragment.
8. The data sharding resource allocation and management system according to claim 1, characterized in that, The shard manager is also used to calculate the priority of shards based on multiple dimensions, and determine the shard migration order according to the priority. The multiple dimensions include business importance, real-time load pressure, data consistency requirements, and client dependency.
9. The data sharding resource allocation and management system according to claim 1, characterized in that, The shard manager employs a leader election mechanism, which includes a preset specific key as a leader lock. All shard manager nodes attempt to create key-value pairs with leases on the specified key. Only the first node to successfully write to the key gains leadership, while other nodes receive a failure response.
10. The data sharding resource allocation and management system according to claim 1, characterized in that, It also includes a tenant isolation module, which is used to assign a unique key pair to each application; require digital signature verification in sharding operation requests; and reject sharding migration instructions that fail key verification.