Method and system for intelligent capacity planning in hybrid cloud environment
By adopting declarative description language and directed acyclic graph DAG in hybrid cloud environments, the problems of load volatility and resource demand uncertainty in hybrid cloud environments are solved, real-time monitoring and optimization scheduling of resources are realized, resource waste is reduced and configuration efficiency is improved.
Patent Information
- Application Number
- CN202510846886.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Traditional capacity planning methods cannot meet the dynamic scheduling needs of load volatility and resource demand uncertainty in hybrid cloud environments, resulting in waste of resources or insufficient performance.
Declarative description language is used to model heterogeneous resources, build directed acyclic graph DAG, perform tasks in parallel through distributed schedulers, synchronize states in real time, and introduce exception detection and automated rollback mechanisms, combined with machine learning optimization strategies.
Real-time monitoring and accurate prediction of resources in hybrid cloud environments are realized, reasonable resource scheduling strategies are generated, manual intervention is reduced, resource utilization efficiency and configuration optimization are improved.
Smart Images

Figure CN120353610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing technology, and particularly to a method and system for intelligent capacity planning in a hybrid cloud environment. Background Art
[0002] With the rapid development of cloud computing technology, more and more enterprises and organizations choose to adopt a hybrid cloud architecture to balance flexibility and cost control by combining the advantages of private clouds and public clouds.
[0003] In a hybrid cloud environment, due to load volatility, uncertainty in resource requirements, and differences in resource management methods among different cloud platforms, traditional capacity planning methods can no longer meet the resource scheduling requirements in a dynamic environment. Traditional capacity planning relies on static rules or manual intervention and cannot adjust cloud resources in real time, resulting in resource waste or insufficient performance.
[0004] With the development of machine learning, deep learning, and optimization algorithms, intelligent capacity planning has gradually become an effective solution to this problem. In order to flexibly respond to load fluctuations, reduce resource waste, and optimize cost-effectiveness, the present invention proposes a method and system for intelligent capacity planning in a hybrid cloud environment. Summary of the Invention
[0005] In order to make up for the defects of the prior art, the present invention provides a simple and efficient method and system for intelligent capacity planning in a hybrid cloud environment.
[0006] The present invention is realized by the following technical solutions: A method for intelligent capacity planning in a hybrid cloud environment includes the following steps: Step S1, resource modeling Model heterogeneous resources in the hybrid cloud, including computing, storage, network, and middleware, using a unified declarative description language, define the target state, configuration parameters, dependencies, and lifecycle hooks of the heterogeneous resources, and accurately describe the target state and configuration intent of the resources through a structured document; Step S2, dependency parsing and directed acyclic graph (DAG) generation In order to ensure the sequentiality and correctness of complex resources during deployment and change, after parsing the modeling data, automatically identify the dependencies between resources, and based on the dependencies between resources, use a topological sorting algorithm to construct a directed acyclic graph (DAG) to clarify the order of resource creation and change; Step S3, task scheduling and execution control Based on the directed acyclic graph (DAG), batch the resource operation tasks, parallelly execute tasks without dependencies through a distributed scheduler, and dynamically adjust the concurrency and execution rhythm of the tasks according to system load and cloud platform limitations; Step S4, Status Synchronization and Consistency Verification To ensure that the system view is consistent with the actual cloud resource status, during the resource operation process, the running status of each resource is synchronized in real time. An event-driven and polling combined mechanism is adopted to obtain status change information, and a state machine is used for status consistency verification and exception marking; Step S5, Exception Detection and Automatic Rollback When an exception occurs in the resource operation, the exception type is automatically identified. For transient exceptions, a retry mechanism is triggered, and for severe exceptions, a rollback process is triggered. The relevant resources are destroyed in the reverse topological sorting order to ensure the integrity of the dependency relationship; Step S6, Execution Feedback and Continuous Optimization After the resource orchestration process is completed, task execution data is automatically collected, including execution time consumption, failure rate, and retry count metrics. After summarizing the whole-process data, combined with the built-in analysis engine, performance evaluation and policy optimization of the resource operation are carried out to achieve intelligent evolution.
[0007] In the above-mentioned step S1, the data serialization format YAML or JSON format standard language is adopted to encapsulate the interface parameters of mainstream cloud platforms, support the unified description of various types of resources, including computing, network, storage, security group, load balancing, database, and middleware, and support model parameterization, templatization, public module reference, and resource combination modeling; Among them, model parameterization and templatization refer to achieving high reuse, dynamic adjustment, and flexible expansion of the resource model through variable definition, reference of common modules, and nested combination methods; A resource abstraction adapter is introduced to connect to the resource application programming interface (API) interfaces of different cloud providers (including Alibaba Cloud, Huawei Cloud, Amazon Web Services (AWS), etc.), automatically convert resource parameters, and achieve cross-platform modeling and unified delivery; A version management and auditing mechanism is built, and all resource description files are uniformly stored in a centralized version management system, supporting configuration comparison, change approval, version backtracking, and audit tracking to meet the requirements of security compliance and team collaboration; Support embedding lifecycle hook functions (pre_create, post_delete), compliance check items, resource quota control, and permission tags in the model to enhance the integrity, security, and governance capabilities of resource delivery.
[0008] In the above-mentioned step S2, the multi-level dependency recognition technology automatically parses explicit dependencies and implicit dependencies, constructs a complete dependency graph, uses the topological sorting algorithm to convert the directed acyclic graph (DAG) into a strictly ordered execution plan, and customizes the division of concurrent batches to improve the scheduling efficiency while ensuring the correct dependency relationship; In the directed acyclic graph (DAG) generation stage, key path nodes that affect the overall execution time are automatically identified, and priority scheduling and processing are performed on them to optimize the key path. In combination with the prediction model, the execution cost of tasks is dynamically evaluated, and the resource usage order is optimized; Support graph pruning or expansion of the directed acyclic graph (DAG) according to the actual running state of resources to improve processing performance and scenario adaptability; The generated directed acyclic graph (DAG) is visually displayed through a graphical interface for administrators to review, adjust, and simulate the execution path, so as to reduce the risk of execution failure caused by modeling errors.
[0009] In step S3, the distributed scheduler batches and executes tasks in the topological order of the directed acyclic graph (DAG). While ensuring the order of resource dependencies, parallel execution is carried out between independent nodes to maximize the execution efficiency; According to the rate limit of the cloud platform application programming interface (API), the processing characteristics of resource types, and the current system load, the task concurrency and execution rate are adaptively and dynamically adjusted to achieve fine-grained scheduling of resources and interfaces.
[0010] Support off-site deployment of execution nodes. The distributed scheduler assigns tasks nearby according to the cloud platform or region to which the resources belong to avoid delays and failures caused by cross-region calls; Each resource operation task has an idempotency mechanism, and the timeout threshold and exponential backoff retry strategy are custom-set to ensure that the task can recover by itself under short-term exceptions; The distributed scheduler has the ability to provide status feedback and can judge in real time whether the task is successfully advanced. If it is found that the status is blocked or in error, the rearrangement and compensation mechanism is actively triggered.
[0011] In step S4, a dual-channel mechanism combining the cloud platform native event bus and timed polling is adopted. By subscribing to event notifications from cloud providers, low-latency status updates are achieved; for resource types that do not support event push, a differential polling mechanism is enabled for status compensation; Build a standard resource state transition model, clearly define the state paths from "being created" to "running", and from "failed" to "rolling back", and support state-triggered automatic processing; Combined with the distributed consistency verification strategy, by comparing the resource description model with the actual status fields, state drift and configuration deviation are identified, and the synchronization, repair, or alarm process is automatically triggered to ensure the timeliness and integrity of synchronization; Use a distributed storage system to store status information to ensure data multi-copy, transaction consistency, and high-concurrency read and write; Record a snapshot of the resource status before critical operations. If the operation fails, compare the status at the rollback point and perform reverse operations or compensation processing to enhance the state consistency closed-loop.
[0012] In step S5, exceptions are classified into transient, structural, and policy types, and corresponding countermeasures are customized respectively; Among them, transient exceptions include network jitter and interface congestion, structural exceptions include missing dependencies and parameter errors, and policy exceptions include quota overrun and insufficient permissions; Retry and skip mechanism: For transient exceptions, automatic retry is performed according to the exponential backoff strategy; for non-critical resources that are custom-selected to tolerate failures, a skip strategy configured by the user is adopted to ensure the progress of the overall process; When the task fails and cannot be recovered, rollback tasks are executed in the reverse topological order of the directed acyclic graph (DAG) to avoid resource residue and configuration drift caused by uncleared factor resources; Support users to declare rollback hook functions in the model, inject custom recovery scripts or notification mechanisms to implement custom compensation logic and meet differentiated scenarios; All exception information, execution stack information, context variables, retry history, and rollback results are written into the audit log to achieve comprehensive audit records for post-event problem review and intelligent optimization analysis.
[0013] In step S6, the task execution data collection includes the following dimensions: task completion rate, average / maximum duration, concurrency utilization, application programming interface (API) failure rate, retry distribution, failure type distribution, and critical path length; Based on data analysis, performance bottleneck nodes and high-risk resources are identified, and the structure and scheduling strategy of the directed acyclic graph (DAG) are automatically adjusted to improve execution stability; Machine learning algorithms are used to model historical task execution data to build resource behavior models and deployment time estimation models, supporting optimal path prediction and concurrency setting before task scheduling; According to historical successful experiences, the optimal resource combination method, API call parameters, and concurrency control strategy are automatically recommended to achieve automatic process optimization; and the optimization results are automatically written into the model template and scheduling policy library to form the system's self-learning and policy self-evolution capabilities to support business scale expansion.
[0014] A system for intelligent capacity planning in a hybrid cloud environment for implementing the above method, including: A declarative resource modeling module responsible for modeling heterogeneous resources in the hybrid cloud, including computing, storage, network, and middleware, using a unified declarative description language, defining the target state, configuration parameters, dependency relationships, and lifecycle hooks of the heterogeneous resources, and accurately describing the target state and configuration intent of the resources through structured documents; A dependency relationship resolution module responsible for automatically identifying the dependency relationships between resources after parsing the modeling data; A Directed Acyclic Graph (DAG) generation module, which is responsible for constructing a Directed Acyclic Graph (DAG) based on the dependency relationships between resources using a topological sorting algorithm to clarify the order of resource creation and change; A distributed task scheduling module, which is responsible for batch-dividing resource operation tasks based on the Directed Acyclic Graph (DAG) and parallelly executing tasks without dependency relationships through a distributed scheduler; An execution control and status synchronization module, which is responsible for dynamically adjusting the concurrency and execution rhythm of tasks according to system load and cloud platform limitations, and during the resource operation process, real-time synchronizing the running status of each resource, obtaining status change information using a mechanism combining event-driven and polling, and performing status consistency verification and exception marking through a state machine; An exception handling and automatic rollback module, which is responsible for automatically identifying the exception type when an exception occurs in resource operation, triggering a retry mechanism for transient exceptions, triggering a rollback process for severe exceptions, and destroying relevant resources in the reverse topological sorting order to ensure the integrity of the dependency relationship; An execution feedback and optimization module, which is responsible for automatically collecting task execution data after the resource orchestration process is completed, including indicators such as execution time, failure rate, and retry times; after summarizing the whole-process data, combining with a built-in analysis engine to perform performance evaluation and policy optimization on resource operations to achieve intelligent evolution.
[0015] A device for intelligent capacity planning in a hybrid cloud environment, including a memory and a processor; the memory is used to store a computer program, and the processor is used to implement the above method steps when executing the computer program.
[0016] A readable storage medium, on which a computer program is stored, and the computer program implements the above method steps when executed by a processor.
[0017] The beneficial effects of the present invention are as follows: The method and system for intelligent capacity planning in the hybrid cloud environment can not only monitor and analyze the usage of resources in real time, accurately predict resource requirements, but also generate reasonable resource scheduling strategies based on performance and cost constraints, and reduce manual intervention through automated execution, thereby achieving elastic expansion and optimized configuration of cloud resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] APPENDIX Figure 1Schematic diagram of the method for intelligent capacity planning in the hybrid cloud environment of the present invention.
[0020] Appendix Figure 2 Schematic diagram of the system architecture for intelligent capacity planning in the hybrid cloud environment of the present invention. Detailed implementation manners
[0021] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0022] By learning historical data, the intelligent capacity planning method can predict future resource requirements according to the changes in the load, automatically make decisions on scaling up and down, thereby improving resource utilization efficiency, reducing operating costs, and ensuring high availability and performance requirements of the service.
[0023] The method for intelligent capacity planning in the hybrid cloud environment includes the following steps: Step S1, Resource modeling Model heterogeneous resources in the hybrid cloud, including computing, storage, network, and middleware, using a unified declarative description language, define the target state, configuration parameters, dependency relationships, and lifecycle hooks of the heterogeneous resources, and accurately describe the target state and configuration intent of the resources through structured documents; In the heterogeneous hybrid cloud environment, there are a large number of resource types, complex structures, and inconsistent configuration standards. The traditional script-based configuration method is difficult to meet the requirements of large-scale and rapidly changing resource orchestration. The declarative resource modeling method adopted in Step S1 fundamentally improves the portability and consistency of resource management.
[0024] Step S2, Dependency relationship parsing and directed acyclic graph DAG generation In order to ensure the sequence and correctness of complex resources during deployment and change, after parsing the modeling data, automatically identify the dependency relationships between resources, and based on the dependency relationships between resources, use a topological sorting algorithm to construct a directed acyclic graph DAG to clarify the order of resource creation and change; Step S3, Task scheduling and execution control Batch divide the resource operation tasks based on the directed acyclic graph DAG, parallelly execute the tasks without dependency relationships through a distributed scheduler, and dynamically adjust the concurrency and execution rhythm of the tasks according to the system load and cloud platform limitations; Step S4, State synchronization and consistency verification To ensure that the system view is consistent with the actual cloud resource status, during the resource operation process, the running status of each resource is synchronized in real time. An event-driven and polling combined mechanism is adopted to obtain status change information, and a state machine is used for status consistency verification and exception marking; Step S5, Exception Detection and Automatic Rollback When an exception occurs in the resource operation, the exception type is automatically identified. For transient exceptions, a retry mechanism is triggered, and for severe exceptions, a rollback process is triggered. Relevant resources are destroyed in the reverse topological sorting order to ensure the integrity of the dependency relationship; Step S6, Execution Feedback and Continuous Optimization After the resource orchestration process is completed, task execution data is automatically collected, including execution time consumption, failure rate, and retry count metrics. After summarizing the whole-process data, combined with the built-in analysis engine, performance evaluation and policy optimization of the resource operation are carried out to achieve intelligent evolution.
[0025] In the said step S1, the data serialization format YAML or JSON format standard language is adopted to encapsulate the interface parameters of mainstream cloud platforms, support the unified description of various types of resources, including computing, network, storage, security group, load balancing, database, and middleware, and support model parameterization, templatization, common module reference, and resource combination modeling; Among them, model parameterization and templatization refer to achieving high reuse, dynamic adjustment, and flexible expansion of the resource model through variable definition, reference of common modules, and nested combination methods; A resource abstraction adapter is introduced to connect to the resource application programming interface (API) interfaces of different cloud providers (including Alibaba Cloud, Huawei Cloud, Amazon Web Services AWS, etc.), automatically convert resource parameters, and achieve cross-platform modeling and unified delivery; A version management and auditing mechanism is built to uniformly store all resource description files in a centralized version management system, support configuration comparison, change approval, version backtracking, and audit tracking to meet the requirements of security compliance and team collaboration; Support embedding lifecycle hook functions (pre_create, post_delete), compliance check items, resource quota control, and permission tags in the model to enhance the integrity, security, and governance capabilities of resource delivery.
[0026] In the said step S2, the multi-level dependency recognition technology automatically parses explicit dependencies and implicit dependencies, constructs a complete dependency graph, uses the topological sorting algorithm to convert the directed acyclic graph (DAG) into a strictly ordered execution plan, and customizes the division of concurrent batches to improve the scheduling efficiency while ensuring the correct dependency relationship; During the directed acyclic graph (DAG) generation phase, the critical path nodes that affect the overall execution time are automatically identified, and priority scheduling is performed on them to optimize the critical path. In combination with the prediction model, the execution cost of tasks is dynamically evaluated, and the resource usage order is optimized. Support graph pruning or expansion of the directed acyclic graph (DAG) according to the actual running state of resources to improve processing performance and scenario adaptability. The generated directed acyclic graph (DAG) is visually displayed through a graphical interface for administrators to review, adjust, and simulate the execution path, so as to reduce the risk of execution failure caused by modeling errors.
[0027] In step S3, the distributed scheduler batches and executes tasks in the topological order of the directed acyclic graph (DAG). While ensuring the order of resource dependencies, parallel execution is carried out between independent nodes to maximize the execution efficiency. According to the rate limit of the cloud platform application programming interface (API), the processing characteristics of resource types, and the current system load, the task concurrency and execution rate are adaptively and dynamically adjusted to achieve fine-grained scheduling of resources and interfaces.
[0028] Support remote deployment of execution nodes. The distributed scheduler assigns tasks nearby according to the cloud platform or region to which the resources belong to avoid delays and failures caused by cross-region calls. Each resource operation task has an idempotency mechanism, and the timeout threshold and exponential backoff retry strategy are customarily set to ensure that the task can recover by itself under short-term exceptions. The distributed scheduler has the ability to provide status feedback, and can judge in real time whether the task is successfully advanced. If it is found that the status is blocked or in error, the rearrangement and compensation mechanism is actively triggered.
[0029] In step S4, a dual-channel mechanism combining the cloud platform native event bus and timed polling is adopted. By subscribing to the event notifications of cloud providers, low-latency status updates are achieved. For resource types that do not support event push, a differentiated polling mechanism is enabled for status compensation. Build a standard resource state transition model, clearly define the state paths from "being created" to "running", and from "failed" to "rolling back", and support state-triggered automatic processing. Combined with the distributed consistency verification strategy, by comparing the resource description model with the actual status fields, state drift and configuration deviation are identified, and the synchronization, repair, or alarm process is automatically triggered to ensure the timeliness and integrity of synchronization. Use a distributed storage system to store status information to ensure data multi-copy, transaction consistency, and high-concurrency read and write. Record the resource status snapshot before critical operations. If the operation fails, perform reverse operations or compensation processing by comparing the status at the rollback point to enhance the state consistency closed-loop.
[0030] In step S5, exceptions are classified into transient, structural, and policy types, and corresponding handling strategies are customized respectively; Among them, transient exceptions include network jitter and interface congestion, structural exceptions include missing dependencies and parameter errors, and policy exceptions include quota overrun and insufficient permissions; Retry and skip mechanism: For transient exceptions, retry automatically according to the exponential backoff strategy; for non-critical resources that are custom-selected to tolerate failures, adopt the skip strategy configured by the user to ensure the progress of the overall process; When the task fails and cannot be recovered, execute the rollback task in the reverse topological order of the directed acyclic graph (DAG) to avoid resource residue and configuration drift caused by uncleared factor resources; Support users to declare rollback hook functions in the model, inject custom recovery scripts or notification mechanisms to implement custom compensation logic and meet differentiated scenarios; Write all exception information, execution stack information, context variables, retry history, and rollback results into the audit log to achieve comprehensive audit records for post-event problem review and intelligent optimization analysis.
[0031] In step S6, the task execution data collection includes the following dimensions: task completion rate, average / maximum duration, concurrency utilization rate, application programming interface (API) failure rate, retry distribution, failure type distribution, and critical path length; Based on data analysis, identify performance bottleneck nodes and high-risk resources, and automatically adjust the structure of the directed acyclic graph (DAG) and scheduling strategy to improve execution stability; Use machine learning algorithms to model historical task execution data, construct resource behavior models and deployment time estimation models to support optimal path prediction and concurrency setting before task scheduling; According to historical successful experiences, automatically recommend the optimal resource combination method, application programming interface (API) call parameters, and concurrency control strategy to achieve automatic process optimization; and automatically write the optimization results into the model template and scheduling policy library to form the system's self-learning and policy self-evolution capabilities to support business scale expansion.
[0032] The intelligent capacity planning system in the hybrid cloud environment for implementing the above method includes: The declarative resource modeling module is responsible for modeling heterogeneous resources in the hybrid cloud, including computing, storage, network, and middleware, using a unified declarative description language, defining the target state, configuration parameters, dependency relationships, and lifecycle hooks of heterogeneous resources, and accurately describing the target state and configuration intent of resources through structured documents; The dependency relationship parsing module is responsible for automatically identifying the dependency relationships between resources after parsing the modeling data; A Directed Acyclic Graph (DAG) generation module, responsible for constructing a directed acyclic graph (DAG) based on the dependency relationships between resources using a topological sorting algorithm to clarify the order of resource creation and change; A distributed task scheduling module, responsible for batch partitioning resource operation tasks based on the directed acyclic graph (DAG) and parallelly executing tasks without dependencies through a distributed scheduler; An execution control and status synchronization module, responsible for dynamically adjusting the concurrency and execution rhythm of tasks according to system load and cloud platform limitations, and during the resource operation process, real-time synchronizing the running status of each resource, using a mechanism that combines event-driven and polling to obtain status change information, and performing status consistency verification and exception marking through a state machine; An exception handling and automatic rollback module, responsible for automatically identifying the exception type when an exception occurs during resource operation, triggering a retry mechanism for transient exceptions, triggering a rollback process for severe exceptions, and destroying related resources in reverse topological sorting order to ensure the integrity of the dependency relationship; An execution feedback and optimization module, responsible for automatically collecting task execution data after the resource orchestration process is completed, including metrics such as execution time, failure rate, and retry count; after summarizing the whole-process data, combining with a built-in analysis engine to perform performance evaluation and policy optimization on resource operations to achieve intelligent evolution.
[0033] The device for intelligent capacity planning in this hybrid cloud environment includes a memory and a processor; the memory is used to store computer programs, and the processor is used to implement the above method steps when executing the computer programs.
[0034] A computer program is stored on the readable storage medium, and when the computer program is executed by the processor, the above method steps are implemented.
[0035] Compared with the prior art, the method and system for intelligent capacity planning in this hybrid cloud environment have the following characteristics: 1). Improve the delivery efficiency and accuracy of flow path resources: Through declarative modeling and automatic parsing of dependency relationships, the automated processing of resources from definition to deployment is realized, which can avoid human configuration errors and repetitive labor, and greatly improve the resource delivery efficiency and accuracy.
[0036] 2). Enhance the compatibility and unity across cloud platforms: Build a unified modeling and scheduling framework, shield the differences in application programming interfaces (APIs) of different cloud providers, realize the centralized modeling, unified orchestration, and undifferentiated execution of multi-cloud resources, and improve the heterogeneous adaptation ability of the system.
[0037] 3). Achieved highly available and elastic resource scheduling control: Adopted a distributed task scheduling architecture, supported high-concurrency and scalable resource creation and change processes, and ensured the high availability and stability of resource operations through dynamic concurrency control and exception handling mechanisms.
[0038] 4). Ensured resource state consistency and automatic self-healing: Introduced a real-time state synchronization mechanism and a consistency verification process to ensure that the system view is always synchronized with the actual resource state, supported automatic discovery and repair of state drift problems, and significantly improved the reliability of resource management.
[0039] 5). Improved the ability to handle exceptions and the efficiency of fault recovery: Equipped with a complete automatic retry, classification processing, and rollback mechanism, which can automatically adopt corresponding strategies according to the exception type, effectively reducing the deployment failure rate, shortening the fault recovery time, and enhancing the robustness of the system.
[0040] 6). Achieved intelligent optimization and continuous evolution of resource management: By collecting and analyzing various data during the execution process, dynamically optimized the task scheduling strategy, and formed intelligent recommendation and automatic learning capabilities based on historical data, enabling the resource orchestration system to continuously evolve and self-optimize.
[0041] 7). Supported the closed-loop management of the entire resource life cycle: This solution covers the entire life cycle processes of resource definition, deployment, monitoring, exception handling, optimization feedback, etc., truly realizing end-to-end automation and intelligent closed-loop management from resource intention to the final state.
[0042] The above-described embodiments are only one of the specific implementation manners of the present invention, and the common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for intelligent capacity planning in a hybrid cloud environment, characterized in that: It includes the following steps: Step S1, Resource Modeling Use a unified declarative description language to model heterogeneous resources in the hybrid cloud, including computing, storage, network, and middleware, define the target state, configuration parameters, dependencies, and lifecycle hooks of the heterogeneous resources, and describe the target state and configuration intent of the resources through structured documents; Step S2, Dependency Resolution and Directed Acyclic Graph (DAG) Generation To ensure the sequentiality and correctness of complex resources during deployment and change processes, after parsing the modeling data, automatically identify the dependencies between resources, and based on the dependencies between resources, use a topological sorting algorithm to construct a directed acyclic graph (DAG) to clarify the order of resource creation and change; Step S3, Task Scheduling and Execution Control Based on the directed acyclic graph (DAG), batch the resource operation tasks, parallelly execute the tasks without dependencies through a distributed scheduler, and dynamically adjust the concurrency and execution rhythm of the tasks according to the system load and cloud platform limitations; Step S4, State Synchronization and Consistency Verification To ensure that the system view is consistent with the actual cloud resource state, during the resource operation process, real-time synchronize the running states of each resource, use a mechanism combining event-driven and polling to obtain state change information, and perform state consistency verification and exception marking through a state machine; Step S5, Exception Detection and Automated Rollback When an exception occurs in the resource operation, automatically identify the exception type, trigger a retry mechanism for transient exceptions, trigger a rollback process for severe exceptions, and destroy the relevant resources in the reverse topological sorting order to ensure the integrity of the dependencies; Step S6, Execution Feedback and Continuous Optimization After the resource orchestration process is completed, automatically collect task execution data, including execution time, failure rate, and retry count metrics; after summarizing the whole-process data, combine with the built-in analysis engine to perform performance evaluation and policy optimization on the resource operation to achieve intelligent evolution.
2. The method for intelligent capacity planning in a hybrid cloud environment according to claim 1, wherein: In the said Step S1, use the data serialization format YAML or JSON format standard language to encapsulate the interface parameters of the cloud platform, support unified description of various types of resources, including computing, network, storage, security group, load balancer, database, and middleware, support model parameterization, templatization, common module reference, and resource combination modeling; Among them, model parameterization and templatization refer to realizing the reuse, dynamic adjustment, and flexible expansion of the resource model through variable definition, reference of common modules, and nested combination methods; Introduce a resource abstraction adapter to dock with the resource application programming interface (API) interfaces of different cloud providers, automatically convert resource parameters, and achieve cross-platform modeling and unified delivery; Build a version management and auditing mechanism, uniformly store all resource description files in a centralized version management system, support configuration comparison, change approval, version rollback, and audit tracking to meet the requirements of security compliance and team collaboration; Support embedding lifecycle hook functions, compliance check items, resource quota control, and permission tags in the model to enhance the integrity, security, and governance capabilities of resource delivery.
3. The method for intelligent capacity planning in a hybrid cloud environment according to claim 1, wherein: In step S2, the multi-level dependency recognition technology automatically parses explicit and implicit dependencies, constructs a complete dependency graph, uses the topological sorting algorithm to transform the directed acyclic graph (DAG) into an ordered execution plan, and customizes the division of concurrent batches to improve the scheduling efficiency while ensuring the correct dependency relationship; In the stage of generating the directed acyclic graph (DAG), it automatically identifies the key path nodes that affect the overall execution time, preferentially schedules and processes them to achieve key path optimization, and combines the prediction model to dynamically evaluate the task execution cost and optimize the resource usage order; It supports graph pruning or expansion of the directed acyclic graph (DAG) according to the actual running state of the resources to improve the processing performance and scenario adaptability; The generated directed acyclic graph (DAG) is visually displayed through a graphical interface for administrators to review, adjust, and simulate the execution path to reduce the risk of execution failure caused by modeling errors.
4. The method for intelligent capacity planning in a hybrid cloud environment according to claim 1, wherein: In step S3, the distributed scheduler executes tasks in batches according to the topological order of the directed acyclic graph (DAG). While ensuring the order of resource dependencies, it executes in parallel between independent nodes to maximize the execution efficiency; According to the rate limit of the cloud platform application programming interface (API), the processing characteristics of resource types, and the current system load, it adaptively and dynamically adjusts the task concurrency and execution rate to achieve fine-grained scheduling of resources and interfaces; It supports deploying execution nodes remotely. The distributed scheduler dispatches tasks nearby according to the cloud platform or region to which the resources belong to avoid delays and failures caused by cross-region calls; Each resource operation task has an idempotency mechanism, and customizes the timeout threshold and exponential backoff retry strategy to ensure that the task can recover itself under short-term exceptions; The distributed scheduler has the ability to provide status feedback, and can judge in real time whether the task is successfully advanced. If it finds that the status is blocked or in error, it will actively trigger the rearrangement and compensation mechanism.
5. The method for intelligent capacity planning in a hybrid cloud environment according to claim 1, wherein: In step S4, a dual-channel mechanism combining the cloud platform native event bus and timed polling is adopted. By subscribing to the event notifications of cloud providers, low-latency status updates are achieved; for resource types that do not support event pushing, a differential polling mechanism is enabled for status compensation; A standard resource state transition model is constructed, which clearly defines the state paths from "creating" to "running" and from "failed" to "rolling back" to support state-triggered automatic processing; Combined with the distributed consistency verification strategy, by comparing the resource description model with the actual status fields, status drift and configuration deviation are identified, and the synchronization, repair, or alarm process is automatically triggered to ensure the timeliness and integrity of synchronization; A distributed storage system is used to store status information to ensure data multi-copy, transaction consistency, and high-concurrency read and write; Before critical operations, a resource status snapshot is recorded. If the operation fails, the reverse operation or compensation process is executed by comparing the status at the rollback point to enhance the state consistency closed-loop.
6. The method for intelligent capacity planning in a hybrid cloud environment according to claim 1, wherein: In step S5, exceptions are divided into transient, structural, and policy types, and corresponding countermeasures are customized respectively; Among them, transient exceptions include network jitter and interface congestion, structural exceptions include dependency missing and parameter errors, and policy exceptions include quota overrun and insufficient permissions; Retry and Skip Mechanism: For transient exceptions, automatic retry is performed according to the exponential backoff strategy; for non-critical resources that are custom-selected to tolerate failures, the skip strategy configured by the user is adopted to ensure the progress of the overall process; When the task fails and cannot be recovered, the rollback tasks are executed in the reverse topological order of the directed acyclic graph (DAG) to avoid resource residues and configuration drifts caused by uncleared factor resources; Support the user to declare a rollback hook function in the model and inject a custom recovery script or notification mechanism to implement custom compensation logic and meet differentiated scenarios; Write all exception information, execution stack information, context variables, retry history, and rollback results into the audit log to achieve comprehensive audit records for post-event problem review and intelligent optimization analysis.
7. The method for intelligent capacity planning in a hybrid cloud environment according to claim 1, wherein: In step S6, the collection of task execution data includes the following dimensions: task completion rate, average / maximum time consumption, concurrency utilization rate, application programming interface (API) failure rate, retry distribution, failure type distribution, and critical path length; Based on data analysis, identify performance bottleneck nodes and high-risk resources, and automatically adjust the structure of the directed acyclic graph (DAG) and the scheduling strategy to improve execution stability; Use machine learning algorithms to model historical task execution data, construct resource behavior models and deployment time prediction models, and support optimal path prediction and concurrency setting before task scheduling; According to historical successful experiences, automatically recommend the optimal resource combination method, application programming interface (API) call parameters, and concurrency control strategy to achieve automatic process optimization; and automatically write the optimization results into the model template and the scheduling strategy library to form the system's self-learning and strategy self-evolution capabilities to support the expansion of the business scale.
8. A system for intelligent capacity planning in a hybrid cloud environment, characterized in that: For implementing the method described in any one of claims 1 to 7, including: A declarative resource modeling module responsible for modeling heterogeneous resources in the hybrid cloud, including computing, storage, network, and middleware, using a unified declarative description language, defining the target state, configuration parameters, dependency relationships, and lifecycle hooks of the heterogeneous resources, and describing the target state and configuration intent of the resources through structured documents; A dependency relationship parsing module responsible for automatically identifying the dependency relationships between resources after parsing the modeling data; A directed acyclic graph (DAG) generation module responsible for constructing a directed acyclic graph (DAG) using the topological sorting algorithm based on the dependency relationships between resources to clarify the order of resource creation and change; A distributed task scheduling module responsible for batch-dividing resource operation tasks based on the directed acyclic graph (DAG) and parallelly executing tasks without dependency relationships through a distributed scheduler; An execution control and status synchronization module responsible for dynamically adjusting the concurrency and execution rhythm of tasks according to system load and cloud platform limitations, and in the process of resource operation, real-time synchronize the running status of each resource, adopt a mechanism combining event-driven and polling to obtain status change information, and perform status consistency verification and exception marking through a state machine; The exception handling and automatic rollback module is responsible for automatically identifying the exception type when an exception occurs in resource operations, triggering a retry mechanism for transient exceptions, triggering a rollback process for severe exceptions, and destroying related resources in the reverse topological sorting order to ensure the integrity of the dependency relationship; The execution feedback and optimization module is responsible for automatically collecting task execution data after the resource orchestration process is completed, including execution time, failure rate, and retry count metrics; after aggregating the whole-process data, it combines the built-in analysis engine to perform performance evaluation and policy optimization on resource operations to achieve intelligent evolution.
9. An apparatus for intelligent capacity planning in a hybrid cloud environment, characterized in that: It includes a memory and a processor; the memory is used to store a computer program, and the processor is used to implement the method according to any one of claims 1 to 7 when executing the computer program.
10. A readable storage medium, characterized in that: A computer program is stored on the readable storage medium, and when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cloud platform stream processing resource allocation method based on dynamic optimization model
CN115185683A
Multi-source computing power data integration and intelligent scheduling system and method
CN118916147A
Multi-cloud platform configuration management method
CN119847726A
Cloud financial system and method based on artificial intelligence
CN119902896A
Edge computing scheduling method and system for heterogeneous multi-source sensor
CN119960950A
Cited By
Cloud resource management method and device based on large model knowledge base, equipment and storage medium
CN120631547A
Distributed simulation platform elastic scheduling method and system based on cloud computing
CN120821547A
New plastic material research and development data processing method based on cloud platform
CN120952719A
Deep learning-based management and education monitoring system
CN121050907A
Disclosed is a deep learning-based pipe teaching monitoring system.
CN121050907B