A resource pool scheduling method based on artificial intelligence
By leveraging artificial intelligence and blockchain technology, the decentralization and immutability of resource pool scheduling are achieved, solving the problems of single point of failure and easy tampering of scheduling records in traditional resource pool scheduling systems, thereby improving the reliability of resource pools and the efficiency of resource allocation.
Patent Information
- Application Number
- CN202511323867.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-09-17
Smart Images

Figure CN120821578B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of resource pool scheduling, in particular to a resource pool scheduling method based on artificial intelligence. BACKGROUND
[0002] In today's era of rapid digital development, with the widespread application of cloud computing, big data, artificial intelligence and other technologies, resource pool scheduling has become a key link to ensure the efficient and stable operation of various applications. Resource pool scheduling aims to reasonably allocate and manage computing, storage, network and other resources to meet the needs of different businesses and tasks, improve resource utilization and reduce operating costs.
[0003] Traditional resource pool scheduling usually adopts a centralized architecture, which has a single point of failure risk. Once the central node fails, the entire scheduling system may be paralyzed, causing resource allocation to be interrupted and seriously affecting business continuity. For example, in a large data center, if the centralized scheduling server encounters hardware failure or network attack, a large number of computing tasks will not be able to obtain the required resources in time, causing business to stagnate and causing huge losses to the enterprise. SUMMARY
[0004] The application provides a resource pool scheduling method based on artificial intelligence, which realizes decentralized scheduling decision-making through resource proof strategy, eliminates single point of failure, ensures the non-tamperability of scheduling records through storage chain, uses cross-chain resource loan mechanism to further optimize resource allocation, improve resource utilization, and ensure the security and non-tamperability of scheduling decisions.
[0005] The application provides a resource pool scheduling method based on artificial intelligence, which comprises:
[0006] S101, splitting or merging resource pools according to geographical areas and resource types, using an algorithm to build a double-chain architecture on a blockchain platform, wherein the resource pool is a sub-chain;
[0007] S102, verifying the use state of node resources through a resource proof strategy, writing the verified state into a scheduling chain, monitoring resources in real time, and setting a threshold based on the obtained real-time resources to allocate resources within the same resource pool;
[0008] S103, detecting the resource conditions of different resource pools, using a scheduling chain to realize cross-chain resource loan, and starting a backup resource pool if more than half of the resource pool nodes fail.
[0009] Preferably, the divided sub-chains are named according to the area + resource type + sub-pool identifier, and the sub-pool identifier is used to distinguish different uses of the same type of resources within the same area.
[0010] Preferably, the method of sub-chain merging is: finding adjacent sub-chains with the same resource type in the same area, constructing an adjacency graph according to resource attributes, and merging into a new sub-chain when detecting that the real-time resource utilization rate of adjacent sub-chains continuously falls below the utilization rate threshold.
[0011] Preferably, the double-chain architecture includes a scheduling chain and a storage chain.
[0012] Preferably, the resource proof strategy is a distributed consensus algorithm based on actual resource contribution of nodes, verifying the storage resource usage state of nodes.
[0013] Preferably,
[0014] S201, in the process of splitting or merging of the resource pool, the node is allocated to the resource pool after splitting or merging, the Pod migration rule is set, and the Pod that needs to be synchronously migrated is identified according to the migration rule;
[0015] S202, before splitting the resource pool, a snapshot is generated by capturing cross-chain transactions and resource state information, the generated snapshot is stored, after splitting the resource pool, the cross-chain transactions and resource state information are monitored in real time, if it is found that the cross-chain transactions and resource state information are greater than the preset failure threshold, that is, the cross-chain transactions and resource state have been interrupted, otherwise, no interruption has occurred, and the interrupted cross-chain transactions and resource state are recovered based on the generated snapshot.
[0016] Preferably, the Pod is a container group that needs to be synchronously migrated and has a dependency relationship with each other when the system is split or merged.
[0017] Preferably, the migration order of the container group with a dependency relationship is to migrate the dependent party first, and then migrate the dependent party.
[0018] Preferably, the snapshot is a structured capture and storage strategy for system state at a specified time point, and the trigger condition for generating the snapshot is that the system obtains a pre-splitting notification of a sub-chain, identifies that the sub-chain performs a splitting operation after a period of time, and simultaneously detects that the activity of the cross-chain transaction is less than a preset threshold.
[0019] Preferably,
[0020] S301, the resource lending chain and the resource borrowing chain are matched according to the service quality and the resource specification, the matching result is recorded, and the service quality index of the loaning task is monitored in real time, wherein the resource pool includes the resource lending chain and the resource borrowing chain when loaning;
[0021] S302, conflict detection is performed on the multi-sub-chain nodes, and the multiple tainted states are merged through three-level arbitration.
[0022] S303, the multi-sub-chain is merged while the cross-chain mobilization is performed.
[0023] One or more technical solutions provided in the application have at least the following technical effects or advantages: the scheduling decision is decentralized by the resource certification strategy, single point failure is eliminated, the scheduling record is ensured to be non-tamperable by the storage chain, the cross-chain resource loan mechanism is used to further optimize resource allocation, improve resource utilization, and ensure the security and non-tamperability of scheduling decisions;
[0024] By generating snapshots, the problem of transaction interruption and resource island caused by resource pool splitting is solved, the probability of transaction interruption due to splitting is extremely low, the system stability is greatly improved, the reliability and state continuity of the resource pool in the splitting process are improved, the misplacement problem in the resource loan process is solved, and the resource can be accurately and efficiently allocated between different sub-chains;
[0025] By linking sub-chain merging and loaning, conflict monitoring and three-level arbitration, efficient and stable collaboration of cross-chain resource loaning and sub-chain merging is achieved, and by dynamic QoS guarantee and fusion of tainted state, resource allocation conflicts and state inconsistency problems are solved. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 A flowchart of the resource pool scheduling method based on artificial intelligence of the application;
[0027] Figure 2 A snapshot generation flowchart of the application;
[0028] Figure 3 A flowchart of the application for merging multiple sub-chains while cross-chain mobilization. DETAILED DESCRIPTION
[0029] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings; the preferred embodiments of the present application are shown in the drawings, but the present application can be realized in many different forms and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0030] It should be noted that the terms "vertical", "horizontal", "up", "down", "left", "right" and similar expressions used herein are only for illustrative purposes and do not represent the only embodiment.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application; the use herein of the terms "and / or" includes a combination of one or more of the associated listed items.
[0032] Embodiment one: Figure 1 is a flowchart of an artificial intelligence-based resource pool scheduling method, which comprises:
[0033] S101, according to the geographical area and the resource type, the resource pool is split, and an algorithm is used to build a double-chain architecture on a blockchain platform;
[0034] Further, the resource pool is divided into North China, East China, South China and West China according to the geographical area, and is divided into CPU pool, GPU pool, memory pool and storage pool according to the hardware attribute, each resource pool is taken as an independent sub-chain, according to the above splitting rule, the divided sub-chain is named according to the area + resource type + sub-pool identifier, for example: East China-CPU-general computing pool, the sub-pool identifier is used to distinguish different uses (such as training / inference) or performance levels of the same type of resources in the same area, the resource pool (sub-chain) is divided by setting a splitting threshold, when each node is started, it will send registration information to the sub-chain management service, the number of sub-chain nodes is determined by extracting the information in the sub-chain management service, when the number of nodes in the resource pool exceeds the preset splitting threshold, it is split into two sub-chains, if the node distribution crosses multiple cities (such as Beijing + Tianjin), it is split into North China-Beijing-GPU pool and North China-Tianjin-GPU pool; if the node GPU model is mixed (such as A100 + V100), it is split into North China-GPU-A100 pool and North China-GPU-V100 pool, the sub-chain node list and the total amount of resources after splitting are updated, the resource utilization is monitored, when it is identified that the sub-chain resource utilization (CPU / GPU / memory occupancy rate) is continuously below 40% for a period of time, the sub-chain is merged, the method of sub-chain merging is: finding adjacent sub-chains with the same resource type in the same area, for example: East China-CPU-pool A and East China-CPU-pool B, constructing an adjacency relationship graph according to the resource attribute, when it is detected that the real-time resource utilization of adjacent sub-chains is continuously below 40% (CPU / GPU / memory weighted average value is calculated by sliding window), it is merged into a new sub-chain, for example: East China-CPU-merged pool, after the merging of the new sub-chain is completed, the original sub-chain management service will release the redundant sub-chain management service resources, and the node metadata is migrated to the new sub-chain.
[0035] The Solana blockchain platform is selected, and the blockchain platform is architected, the HotStuff, PBFT pluggable consensus algorithm and parallel channel are set on the blockchain platform, the global state is horizontally sharded according to resource types (CPU / GPU / storage), the double-chain architecture includes a scheduling chain and a storage chain, for the construction of the scheduling chain, the scheduling chain is built on the Solana blockchain platform, the performance parameters of the scheduling chain are set, the capacity of the block is expanded, the time interval of the block generation is reduced, the delay is reduced, the verification nodes are reduced, the communication rounds are reduced, for the construction of the storage chain, the scheduling log is segmented to form a data block, the data block is stored in the IPFS network after encryption, the SHA-256 hash value of each data block is calculated, the Merkle tree is constructed in chronological order, the Merkle tree root hash is written into the blockchain every day, forming a double verification mechanism on and off the chain, users can obtain the original data through the IPFS network, and then verify the data integrity through the on-chain root hash, complete the compression of the data, and reduce the compression cost; after the scheduling chain generates an instruction, the SHA-3 hash value of the instruction content is calculated, the hash value, timestamp and scheduling chain transaction ID are encapsulated as an anchor point, and written into the storage chain, the node executes the scheduling instruction, and after the execution is completed, the result is fed back to the storage chain, preventing the scheduling record from being tampered with.
[0036] S102, verify the use state of the node resource through the resource proof strategy, write the verified state into the scheduling chain, monitor the resource in real time, and set a threshold based on the obtained real-time resource to allocate the resources in the same resource pool;
[0037] Specifically, the resource proof policy is a distributed consensus algorithm based on actual resource contribution of nodes. The authenticity of the resource usage state of each node is verified, and the resource list of each node is output every certain period of time. The resource list includes static resources and a tainted state. The static resources are CPU model, core number, GPU model and quantity, and memory capacity. The tainted state is dynamic, and the node has been limited in use. The current limitation state and remaining time length need to be indicated. The dynamic tainted state is a temporary state identifier of the resource permission of the node. The system dynamically assigns or releases it according to the resource usage (such as overload or idle). It controls whether the node can receive a specific type of task (such as a GPU dedicated task or a general CPU task) by adding a limitation condition and a time limit. The dynamic tainted state prevents the node from receiving tasks in violation of the rules during the limitation period. The system randomly selects a number of verification nodes from all online nodes according to the VRF algorithm to form a temporary verification group. The verification node obtains the hardware unique identifier of the target node through the TPM (trusted platform module), compares the hardware unique identifier with the historical fingerprint data in the registration library, confirms that the node has not tampered with the hardware, and queries the real-time allocation rate of the resource pool where the target node is located. If the allocation rate is less than 75%, the node is allowed to receive non-dedicated tasks (such as general CPU computing tasks). If the allocation rate is greater than 85%, the node is re-labeled as a limited state and is only allowed to receive dedicated tasks (such as GPU training tasks). The node resource is verified based on the above-mentioned hardware fingerprint and allocation rate dual verification method.
[0038] The resource usage of each sub-chain is monitored in real time, the resource allocation rate is calculated according to the resource usage, and the threshold range of the resource allocation rate is set. The threshold range includes a minimum allocation rate threshold and a maximum allocation rate threshold. When the calculated sub-chain allocation rate is less than the minimum allocation rate threshold, the scheduling sends a de-taint instruction to the target node, i.e. allowing the node to receive non-dedicated tasks. When the calculated sub-chain allocation rate is greater than the maximum allocation rate threshold, the scheduling sends a tainted recovery instruction to the target node, i.e. only allowing to accept dedicated tasks.
[0039] S103, the resource situation of different sub-chains is detected, and cross-chain resource borrowing is realized by using a scheduling chain. If more than half of the sub-chain nodes fail, the backup chain is automatically started.
[0040] Further, the resource usage of different sub-chains is monitored in real time. If the resource pool allocation rate of the first sub-chain is greater than 90% for a period of time, it is determined that the demand of the first sub-chain cannot be met through task splitting or priority adjustment. The scheduling node of the first sub-chain generates cross-chain borrowing information, the cross-chain borrowing information includes the type of requested resources, the borrowing duration and the preference of the target sub-chain, and sends the request to the scheduling decision chain through the cross-chain communication protocol (IBC). The scheduling decision chain is an independent relay chain responsible for global resource coordination. The scheduling decision chain queries the global resource map, filters the sub-chains that meet the conditions, and sends the request to the entry node of the second sub-chain through the cross-chain protocol. The second sub-chain queries the local resource pool and calculates the current allocation rate. If the current allocation rate is less than the preset lending threshold, the second sub-chain meets the lending condition and generates a resource lending voucher. The scheduler of the first sub-chain sends a task data packet to the second sub-chain through the cross-chain channel according to the lending voucher. If more than half of the nodes are found to be unresponsive through real-time monitoring, a fault alarm is triggered and a backup sub-chain is activated. The faulty node is rebuilt in the backup sub-chain. After the rebuilding is completed, the backup sub-chain is upgraded to the main chain, taking over all tasks and resources of the original sub-chain. The scheduling decision chain updates the global routing table to direct new requests to the rebuilt sub-chain.
[0041] The technical solutions in the embodiments of the application have at least the following technical effects or advantages: the scheduling decision is decentralized through the resource proof strategy, single point failure is eliminated, the scheduling record is ensured to be tamper-proof through the storage chain, the cross-chain resource borrowing mechanism is used to further optimize resource allocation, improve resource utilization, and ensure the security and tamper resistance of the scheduling decision.
[0042] Embodiment two: When the sub-chain is split in the above-mentioned embodiment one, the state of the ongoing scheduling transaction is easily lost. This embodiment solves the state drift problem when the sub-chain is split through snapshot anchoring, as shown in Figure 2 .
[0043] S201, during the splitting or merging of the resource pool, the nodes are allocated to the split or merged resource pool, the Pod migration rule is set, and the Pod that needs to be synchronized is identified according to the migration rule;
[0044] Further, the Pod is a container group that needs to be synchronously migrated and has a close functional dependency relationship with each other when the system is split (i.e. the nodes in the existing sub-chain are re-assigned to the new sub-chain), the system receives a split trigger signal through the monitoring module, the split trigger signal contains the split type and the number of target sub-chains, the system identifies the constraint conditions of the split according to the physical layout of the resource pool, the node capacity of a single area and the continuity requirements of the business, the method for identifying the constraint conditions of the split is: in terms of physical topology characteristics, first, the physical layout of the facility needs to be investigated, including the distribution of racks, such as the specific location of the rack to which the node belongs, the power supply method, the attribution of the network switch, etc.; At the same time, the regional division also needs to be analyzed, and the geographical area range of the data center and the boundary definition of the availability zone are clarified; In addition, various resource capacity indicators are evaluated, including the upper limit of the number of nodes that a single rack or area can carry, the allocation of network bandwidth, and the performance of the storage system. Based on these physical characteristic data, different levels of constraint conditions are derived: at the rack level, if it is detected that there are significant differences in inter-rack network delay, a constraint needs to be set that the nodes of each sub-chain are distributed within no more than 3 racks; If the rack adopts independent power design, a rule that the sub-chain must span at least 2 racks should be formulated to avoid service interruption caused by single power failure. At the regional level, for a system deployed across regions, each sub-chain contains at least 2 nodes in an availability zone to prevent the impact of regional disaster events, in terms of business continuity, for stateful services (such as database systems), the constraint condition considers data consistency and availability, typical requirements include that all nodes within a sub-chain must mount the same storage volume; For stateless services (such as web front-end clusters), relatively loose constraints are set, such as the number of nodes contained in a sub-chain is not less than 2 to support horizontal expansion, but at the same time, the number of nodes in a single sub-chain is limited to no more than 30% of the total number of nodes in the system, using a hash algorithm to evenly distribute the nodes to the new sub-chain, mapping the nodes and sub-chains to a hash ring, and attributing them in a clockwise direction, combined with virtual nodes, the uniformity can be further improved, finally verify whether the distribution of nodes in each sub-chain is uniform, determine the business dependency relationship between the container groups (Pods) that need to be synchronously migrated, for example: a web service depends on a database, set the migration order of the container groups, migrate the dependent party first, then migrate the dependent party, input the container groups that need to be synchronously migrated and the migration order of the container groups into the container orchestration tool Kubernetes, Kubernetes evicts the container groups from the original sub-chain nodes and schedules them to the target nodes of the new sub-chain, after the node allocation and container group migration are completed, update the routing rules of the system, and use gradual switching to update, gradually increase the traffic of the new sub-chain.
[0045] S202, before the resource pool splitting, a snapshot is generated by capturing cross-chain transactions and resource state information, the generated snapshot is stored, after the resource pool splitting, the cross-chain transactions and resource state information are monitored in real time, if it is found that the cross-chain transactions and resource state information are greater than a preset failure threshold, that is, the cross-chain transactions and resource state interrupt, otherwise, no interrupt occurs, the interrupted cross-chain transactions and resource state are recovered based on the generated snapshot;
[0046] Specifically, the snapshot is a structured capture and storage mechanism of system state at a specific time point, and the trigger condition for generating the snapshot is that the system obtains the pre-splitting notification of the sub-chain, identifies that the sub-chain performs the splitting operation after a period of time, and at the same time, the system detects that the activity of the cross-chain transaction is less than a preset threshold; the real-time state of the cross-chain transaction is captured and recorded, the cross-chain transaction coordinator generates a globally unique identifier, reads the timeout timestamp and the current timestamp from the metadata of the transaction, calculates the remaining time according to the timeout timestamp and the current timestamp, the remaining time = timeout timestamp - current timestamp, obtains a sub-chain node list to be processed in the next stage from the to-be-processed transaction, the sub-chain node list includes the IP address and port number of each node, the system identifies the external resources relied on by the transaction during execution by scanning the execution context of the transaction, captures the transaction snapshot field by obtaining the globally unique identifier, the remaining time, the sub-chain node list to be processed in the next stage of the transaction and the dependency state, captures the resource state information, that is, records the resource state scheduled across the sub-chain, allocates a unique identifier to each borrowed resource, reads the current sub-chain name to which the resource belongs from the resource scheduling system, extracts the access path of the resource from the network, and captures the resource snapshot field by obtaining the unique identifier, the current sub-chain name to which the resource belongs and the resource access path. The captured transaction snapshot field and resource snapshot field are replaced by binary encoding instead of text, a dictionary table is established, repeated strings are mapped to short indexes, the indexes are stored in the snapshot, and compression of repeated strings is realized.
[0047] The transaction snapshot field is formed into multiple copies, which are respectively stored in the storage chain, the adjacent sub-chain and the local backup, the resource snapshot field is formed into multiple copies, which are respectively stored in the global registry, the sub-chain to which the resource belongs and the standby sub-chain, the storage location, the hash value and the storage timestamp of each copy are recorded, and a storage log is formed.
[0048] After the resource pool splitting is completed, the cross-chain transaction and resource state information after splitting are monitored in real time, a transaction execution failure threshold and a call failure threshold are set, the cross-chain transaction and resource state information obtained by real-time monitoring are compared with the execution failure threshold and the call failure threshold respectively, if the number of execution failures of the cross-chain transaction is greater than the preset execution failure threshold, it indicates that the resource pool splitting causes the transaction interruption problem, if the number of scheduling failures of the resource state is greater than the preset call failure threshold, it indicates that the resource pool splitting causes the resource island problem, the method for recovering the cross-chain transaction is: reading the transaction snapshot from the local database of the sub-chain node or the adjacent sub-chain, synchronizing all nodes with the master clock source through the NTP service, reading the original timeout timestamp of the transaction from the snapshot when the transaction is recovered, calculating the remaining time after recovery in combination with the current time, the remaining time after recovery = original timeout timestamp - (current time + time deviation compensation), according to the target node list in the snapshot, the transaction is redistributed to the corresponding node through the remote procedure call framework (gRPC); the method for recovering the resource state is: querying the resource loan snapshot from the global registry, if the registry is unavailable, obtaining it from the local copy of the sub-chain to which the resource currently belongs, updating the route by calling the cloud provider API according to the route information in the snapshot to realize the recovery of the resource state.
[0049] The technical solutions in the embodiments of the application have at least the following technical effects or advantages: by generating a snapshot, the problems of transaction interruption and resource island caused by resource pool splitting are solved, the probability of transaction interruption due to splitting is extremely low, the system stability is greatly improved, the reliability and state continuity of the resource pool in the splitting process are improved, the misplacement problem in the resource loan process is solved, and the resource can be accurately and efficiently allocated between different sub-chains.
[0050] Embodiment three: based on the above embodiment one and embodiment two, when the resource pool is loaned, delay fluctuation is prone to occur, in this embodiment, QoS negotiation and taint state fusion are combined with cross-module collaborative process to realize the operation of cross-chain resource loan and sub-chain merging, as shown in Figure 3 .
[0051] S301, the resource lending chain and the resource borrowing chain are matched according to the quality of service and the resource specification, and the matching result is recorded, and the service quality index of the loaning task is monitored in real time, wherein the resource pool includes a resource lending chain and a resource borrowing chain when loaning;
[0052] Further, for the resource provider, i.e. the resource lending chain, the minimum requirement of the quality of service (QoS) that the resource lending chain can accept is that the CPU delay is less than a time threshold, indicating that the lending chain matched with the borrowing chain can complete the processing within 10 milliseconds when the borrowing chain uses its resources to process tasks, ensuring timely response of the task, the network bandwidth is greater than the bandwidth threshold, ensuring that data can be transmitted at a faster speed in the network, avoiding data congestion and transmission delay caused by insufficient bandwidth, the availability is greater than the availability threshold, i.e. the probability of the resource being in a normal use state within a specified time reaches the availability threshold, reducing the business interruption time caused by resource failure or unavailability, and the resource borrowing chain meets the above three requirements and can be matched with the resource lending chain; for the resource receiver, i.e. the resource borrowing chain, the resource borrowing chain allocates an exclusive cache area for the resource lending chain, i.e. reserves 10% of the memory of the borrowing chain as an exclusive cache area, which can avoid the competition of other tasks for the cache area and improve the speed of data reading and processing, the resource borrowing chain configures a priority network channel for the resource lending chain, i.e. configures a priority network channel for the lending chain, and sets the VIP bandwidth guarantee ratio to 120% of the demand of the lending chain, i.e. provides at least 1.2Gbps of network bandwidth, the lending chain verifies the exclusive cache area and the priority network channel provided by the borrowing chain, records the matching result, so that the matching result has the characteristics of non-tamperability and traceability.
[0053] By monitoring the quality indicators of the loaned task in real time, when the quality indicators of the loaned task monitored in real time exceed the set threshold, an alarm is triggered, and the resource is compensated.
[0054] S302, conflict detection is performed on the multiple sub-chain nodes, and the multiple stain states are merged through three-level arbitration;
[0055] Specifically, for each of the multiple sub-chains participating in the merger, the proportion of tainted states of their nodes is statistically analyzed. The tainted states are divided into two types: untainted and restricted. The number of nodes in the untainted and restricted states on each sub-chain is collected, and the proportion of each to the total number of nodes is calculated. A threshold range is set, including a maximum threshold and a minimum threshold. Conflict types are classified based on the proportion of tainted states. When the proportion of nodes in the untainted state in the first sub-chain is greater than the maximum threshold, and the proportion of nodes in the restricted state in the second sub-chain is also greater than the maximum threshold, the conflict type is determined to be untainted vs. restricted. In this type, the two sub-chains have a clear opposition in the tainted state of their nodes. If the proportion of nodes in the untainted state in both sub-chains is between the minimum and maximum thresholds, the conflict type is determined to be partially untainted, indicating that the node states of the two sub-chains have some similarity. When the proportion of nodes in the untainted state in both sub-chains is less than the minimum threshold, the conflict type is determined to be a dispersed conflict. Three levels of arbitration include automatic arbitration, weighted arbitration, and manual arbitration. For automatic arbitration, the Jaccard coefficient is used to calculate the similarity of the node states of the two sub-chains. The coefficient measures the similarity of states by comparing the proportion of nodes in the same state (e.g., both unrestricted or both restricted) in two sub-chains to the total number of nodes in both sub-chains. If the calculated similarity is greater than 85%, the node states of the two chains are directly merged. For example, 82% of the nodes in the East China-GPU pool are in the unrestricted state, and 79% of the nodes in the North China-GPU pool are in the unrestricted state. Since the similarity meets the condition, all nodes are set to unrestricted after merging. For weighted arbitration, a weight is assigned to each node based on its resource quantity. The weight calculation formula is weight = min(1, total node resources / 1000). For example, a node with 1500 cores has a weight of 1, while a node with 800 cores has a weight of 0.8. For manual arbitration, manual arbitration serves as a last resort. When automatic arbitration and weighted voting fail to resolve the issue, decisions are made based on the experience and judgment of the administrators.
[0056] After completing the burst detection and three-level arbitration, the method for merging multiple tainted states is as follows: First, freeze the new scheduling tasks on the two sub-chains. In order to prevent the new task scheduling from interfering with the node state during the state synchronization process, modify the node state according to the arbitration result.
[0057] S303, merging multiple subchains simultaneously during cross-chain operations;
[0058] Further, in the process of merging multiple sub-chains, the borrowed task resource isolation ratio is set to 100%, indicating that regardless of how the sub-chains are merged, the borrowed resources are in a completely independent state and will not be affected by the merging operation, ensuring that the service level agreement (SLA) of the borrowed resources is not disturbed, and the system dynamically adjusts the resource allocation priority of the merged chain. Specifically, it ensures that the CPU delay of the borrowed task always does not exceed 10ms, the system monitors the resource usage and performance indicators of the borrowed task in real time, dynamically allocates computing resources according to actual needs, prioritizes the demand for CPU resources of the borrowed task, avoids increasing CPU delay due to resource competition, and thus guarantees the real-time performance and response speed of the borrowed task; the state of the borrowed node remains independent during the sub-chain merging process and is not affected by the merging arbitration, regardless of whether it is before or after merging, the borrowed node maintains its original tainted state (such as being released or restricted), for example, if a borrowed node is in a released state before merging and is allowed to execute tasks normally, it will still remain in the released state after merging and will not be affected by the state changes of other nodes on the merged chain. This independent state management ensures the stability and reliability of the borrowed task, avoiding abnormal state of the borrowed node due to merging operation, and thus affecting the execution of the borrowed task; when merging sub-chains, only non-borrowed nodes are operated, the system identifies borrowed nodes and skips these nodes during merging, and only the state of non-borrowed nodes is evaluated and adjusted, which can avoid pollution of the state of borrowed nodes by merging operation and ensure that the state of borrowed nodes is not affected by the state changes of other nodes on the merged sub-chain. For example, when merging two sub-chains, the system will first select non-borrowed nodes, then merge and modify the state of these nodes according to the established arbitration rules, and the state of borrowed nodes remains unchanged; the system monitors the QoS indicators of the borrowed task after merging in real time, especially the CPU delay, and if the CPU delay exceeds 13ms (30% wider than the benchmark value), it is determined that the QoS fluctuates, triggering the alarm mechanism.
[0059] The technical solutions in the embodiments of the present application have at least the following technical effects or advantages: through the linkage of sub-chain merging and borrowing, conflict monitoring and three-level arbitration, efficient and stable collaboration of cross-chain resource borrowing and sub-chain merging is achieved, and through dynamic QoS guarantee and tainted state fusion, the problem of resource allocation conflict and state inconsistency is solved.
[0060] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. An artificial intelligence-based resource pool scheduling method, characterized in that, The method comprises the following steps of: S101, splitting or merging resource pools according to geographical areas and resource types, using an algorithm to build a double-chain architecture on a blockchain platform, wherein the resource pools are sub-chains; the double-chain architecture comprises a scheduling chain and a storage chain; S201, during the splitting or merging of the resource pools, assigning nodes to the split or merged resource pools, setting Pod migration rules, and identifying Pods that need to be synchronously migrated according to the migration rules; S202, before the splitting of the resource pools, generating a snapshot by capturing cross-chain transactions and resource state information, storing the generated snapshot, after the splitting of the resource pools, monitoring the cross-chain transactions and resource state information in real time, if it is found that the cross-chain transactions and resource state information are greater than a preset failure threshold, that is, the cross-chain transactions and resource state have been interrupted, otherwise, no interruption has occurred, and based on the generated snapshot, the interrupted cross-chain transactions and resource state are restored; S102, verifying the use state of the node resources through a resource proof strategy, writing the verified state into the scheduling chain, monitoring the resources in real time, and based on the obtained real-time resources, setting a threshold to allocate the resources in the same resource pool; S301, matching the resource lending chain and the resource borrowing chain according to the quality of service and the resource specifications, recording the matching results, and monitoring the service quality indicators of the lending and borrowing tasks in real time, wherein the resource pool includes the resource lending chain and the resource borrowing chain during the lending and borrowing; S302, detecting conflicts of multiple sub-chain nodes, and merging multiple tainted states through three-level arbitration; S303, merging multiple sub-chains while mobilizing across chains; S103, detecting the resource conditions of different resource pools, using the scheduling chain to realize cross-chain resource lending and borrowing, and if more than half of the resource pool nodes fail, starting a backup resource pool.
2. The artificial intelligence-based resource pool scheduling method of claim 1, wherein, The divided sub-chains are named according to the area + resource type + sub-pool identifier, and the sub-pool identifier is used to distinguish different uses of the same type of resources in the same area.
3. The artificial intelligence-based resource pool scheduling method of claim 2, wherein, The method for merging sub-chains is to find adjacent sub-chains with the same resource type in the same area, construct an adjacency relationship graph according to the resource attributes, and when it is detected that the real-time resource utilization rate of the adjacent sub-chains continuously falls below the utilization rate threshold, merge them into a new sub-chain.
4. The artificial intelligence-based resource pool scheduling method of claim 3, wherein, The resource proof strategy is a distributed consensus algorithm based on the actual resource contribution of the node, which verifies the storage resource use state of the node.
5. The artificial intelligence-based resource pool scheduling method of claim 4, wherein, The Pod is a container group that needs to be synchronously migrated and has a dependency relationship between each other when the system is split or merged.
6. The artificial intelligence-based resource pool scheduling method of claim 5, wherein, The migration order of the container groups with a dependency relationship is to migrate the dependent party first, and then migrate the dependent party.
7. The artificial intelligence-based resource pool scheduling method of claim 6, wherein, The snapshot is a structured capture and storage strategy for the system state at a specified time point, and the trigger condition for generating the snapshot is that the system obtains a pre-splitting notification of a sub-chain, identifies that the sub-chain will be split after a period of time, and at the same time, the system detects that the activity of the cross-chain transaction is less than a preset threshold.
Citation Information
Patent Citations
Power distribution network construction resource scheduling method and device, computer equipment and storage medium
CN117852803A
Intelligent contract driven workflow engine automatic execution method and system based on block chain
CN120179419A