Dynamic resource allocation method and system under micro-service architecture

Through real-time monitoring and TimesNet timing model analysis, combined with business priority rules, dynamic adjustment of resource allocation solves the problem of waste and shortage of resource allocation in microservice architecture, realizes flexible allocation and efficient utilization of resources, and improves the stability and reliability of the system.

CN120704900AInactive Publication Date: 2025-09-26INSPUR SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202511205025.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing resource allocation methods in microservice architectures have problems such as idle resources and waste and insufficient resources under high load conditions. In addition, the existing dynamic allocation solutions lack flexibility and are difficult to meet complex and changing resource requirements.

Method used

By monitoring microservice resource usage in real time, using the TimesNet time series model to analyze resource trends and combining business priority rules to generate elastic scaling recommendations, resource allocation is dynamically adjusted, and flexible resource allocation is achieved by combining the Kubernetes API.

Benefits of technology

It improves resource utilization, ensures the stable performance and efficient operation of microservices, reduces resource waste and operating costs, adapts to complex business scenarios, and improves system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704900A_ABST
    Figure CN120704900A_ABST
Patent Text Reader

Abstract

The invention discloses a resource dynamic allocation method and system under a micro-service architecture, relates to the technical field of resource allocation, and aims to overcome the defect that an existing dynamic resource allocation scheme cannot meet operation requirements of micro-services, and the adopted scheme comprises the following steps: acquiring resource use data in a service operation process in real time; cleaning and filtering the collected data of the data collection module to generate a standardized time sequence data set; for the time sequence data set, analyzing a resource use trend and predicting a future resource demand by adopting a time sequence model, and then generating an elastic capacity expansion and contraction suggestion in combination with a preset service priority rule to provide data support for subsequent resource scheduling; and on the basis of the elastic capacity expansion and contraction suggestion, carrying out actual allocation operation on the resources in combination with a pre-formulated resource allocation strategy. According to the method, the resource use condition of the micro-service is monitored in real time, and a scientific and reasonable dynamic adjustment strategy is combined, so that flexible allocation of resources is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource allocation, and in particular to a method and system for dynamic resource allocation under a microservice architecture. Background Art

[0002] In a microservice architecture, different microservices have very different requirements for resources such as memory, CPU, and GPU, and their requirements will change dynamically with factors such as computing tasks.

[0003] Traditional resource allocation methods often use a static allocation strategy, pre-allocating a fixed resource quota for each microservice. This approach has numerous drawbacks. For one thing, when microservices are under low load, resources go unused and wasteful. For another, under high load, the fixed resource quotas can't meet the operational needs of microservices, leading to performance degradation or even service interruptions.

[0004] In addition, some existing dynamic resource allocation schemes lack flexibility in allocation strategies, making it difficult to quickly adapt to the complex and changing resource requirements of microservices, and unable to effectively ensure the stable and efficient operation of microservices. Summary of the Invention

[0005] In response to the defect that existing dynamic resource allocation schemes cannot meet the operating requirements of microservices, the present invention provides a dynamic resource allocation method and system under a microservice architecture. By real-time monitoring of the resource usage of microservices and combining scientific and reasonable dynamic adjustment strategies, flexible allocation of resources such as memory, CPU, and GPU can be achieved, resource utilization can be improved, the stable performance and efficient operation of microservices can be guaranteed, and resource waste and operating costs can be reduced.

[0006] In a first aspect, the present invention provides a method for dynamic resource allocation under a microservice architecture, and the technical solutions adopted to solve the above technical problems are as follows: A method for dynamic resource allocation under a microservice architecture comprises the following steps: S1. Real-time collection of resource usage data during service operation; S2. Clean and filter the collected data from the data acquisition module to generate a standardized time series data set; S3: For time series datasets, the TimesNet time series model is used to analyze resource usage trends and predict future resource demand. Then, based on pre-set business priority rules, elastic scaling recommendations are generated to provide data support for subsequent resource scheduling. S4. Based on the elastic scaling recommendations and in combination with the pre-defined resource allocation strategy, actual resource allocation operations are performed.

[0007] Optionally, step S1 specifically includes the following operations: Deploy monitoring probes in the K8S cluster operating environment to obtain resource usage data at a preset sampling frequency, where: Deploy the Prometheus Operator in the Kubernetes cluster using Helm, and deploy a NodeExporter DaemonSet on each worker node to collect node-level metrics such as CPU usage, memory usage, and disk usage. Use the NVIDIA tool DCGM to collect GPU memory usage, utilization, and temperature-related metrics. Transfer the collected resource data to the next process.

[0008] Optionally, step S2 specifically includes the following operations: Use Prometheus's Recording Rules to process the raw data to remove erroneous data and complete data cleaning; Exponential smoothing method is used to smooth the data, and polynomial interpolation filling method is used to handle missing data to ensure the integrity and availability of the data set. Data filtering and standardization are completed to finally form a standardized time series data set.

[0009] Optionally, the TimesNet timing model analyzes resource usage trends and predicts future resource requirements through the following steps: The Fourier transform is used to convert the time series data contained in the time series data set from the time domain to the frequency domain. By calculating the amplitudes of different frequency components, multiple cycles in the data that conform to the actual load pattern are screened out. Based on these cycles, the data is preprocessed in a targeted manner so that the preprocessed data retains the multi-cycle essential characteristics of the original time series data. A deep residual network is introduced to perform deep feature extraction on preprocessed data tensors, effectively capturing the complex patterns and dynamic changes hidden in time series data; Feature fusion is performed based on the importance weight of each cycle in resource changes, and the resource demand prediction result is finally output to adapt to the mixed cycle characteristics of microservice load.

[0010] Optionally, step S3 is performed to generate elastic scaling recommendations based on the resource demand forecast results of the TimesNet time series model and preset service priority rules. Specifically, the following operations are performed: Load the preset business priority rules, associate the resource demand forecast results output by the TimesNet timing model with the priority labels of the corresponding businesses, and clarify the resource demand trends of businesses with different priorities; Based on business priority rules, resource guarantee thresholds and scaling trigger conditions are set for businesses of different priorities to ensure that resource requirements of high-priority businesses are met first. Combining the prediction results, priority rules, and resource thresholds, it generates preliminary scaling recommendations. It then optimizes the recommendations using preset logic, ultimately generating elastic scaling recommendations that are output to the resource scheduling module in a structured form. At the same time, it continuously monitors the deviation between actual resource usage and the prediction results, and dynamically adjusts the scaling recommendations based on real-time changes in business priorities.

[0011] Optionally, based on business priority classification and combined with predicted resource demand trends and business characteristics, differentiated resource allocation strategies can be developed in advance. Specific details are as follows: Define AI training / inference tasks as high-priority tasks, and define non-real-time data processing services, non-critical business auxiliary services, and backend analysis and report generation services as low-priority tasks; When the resource demand forecast output by the intelligent prediction module is greater than 1.1 times the current resource limit, resource adjustments are triggered for high-priority tasks: 80% of the forecasted resource demand is reserved for the task as a fixed quota, and an additional 20% of resources is allowed for elastic burst during peak demand. When the resource demand forecast output by the intelligent prediction module is greater than 1.3 times the current resource limit, resource adjustment is triggered for low-priority tasks: resources are dynamically allocated based on the actual forecast demand of the task, but the allocation upper limit does not exceed 130% of the forecast demand; soft limits and hard limits are set for the task: soft limits provide basic resource guarantees for task operation, and hard limits constrain the upper limit of resource usage; when cluster resources are tight, the resources of low-priority tasks that exceed the soft limit are reclaimed, but the task is not forcibly terminated.

[0012] Further optionally, execute step S4, and according to the pre-established resource allocation strategy, interact with the underlying platform by calling the Kubernetes API to dynamically adjust the resource quota of the Pod, realize real-time allocation, elastic scaling and on-demand recycling of resources, and ultimately achieve dynamic optimization and efficient utilization of cluster resources, ensuring that the resource requirements of high-priority tasks are met first.

[0013] In a second aspect, the present invention provides a dynamic resource allocation system under a microservice architecture, and the technical solutions adopted to solve the above technical problems are as follows: A dynamic resource allocation system under a microservice architecture, used to implement the method described in the first aspect, comprising: Data collection module, used to collect resource usage data in real time during service operation; The data processing module is used to clean and filter the data collected by the data acquisition module to generate a standardized time series data set; The intelligent prediction module uses the TimesNet time series model to analyze resource usage trends and predict future resource requirements for time series datasets. It then generates elastic scaling recommendations based on pre-set business priority rules, providing data support for resource scheduling. The resource scheduling module is used to allocate resources based on the elastic scaling recommendations generated by the intelligent prediction module and the pre-defined resource allocation strategy.

[0014] In a third aspect, the present invention further provides a device for dynamically allocating resources under a microservice architecture, comprising: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to implement the method described in the first aspect.

[0015] In a fourth aspect, the present invention further provides a computer-readable medium having computer instructions stored thereon, and when the computer instructions are executed by a processor, the method described in the first aspect can be implemented.

[0016] The method and system for dynamic resource allocation under a microservice architecture of the present invention have the following beneficial effects compared with the prior art: 1. This invention monitors the resource usage of microservices in real time and combines scientific and reasonable dynamic adjustment strategies to achieve flexible allocation of memory, CPU, GPU and other resources, improve resource utilization, ensure the stable performance and efficient operation of microservices, and reduce resource waste and operating costs; 2. The present invention can avoid idle waste and shortage of resources, significantly improve resource utilization and reduce operating costs through dynamic and flexible allocation of resources; combine time series prediction models with resource allocation strategies, adapt to complex business scenarios, ensure stable and efficient operation of microservices, and improve system reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Attachment Figure 1 is a flow chart of a method according to embodiment 1 of the present invention; Attachment Figure 2 This is a module connection block diagram of the second embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to make the technical solution, the technical problems solved and the technical effects of the present invention more clear, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.

[0019] Example 1: Combined with the attached Figure 1 This embodiment proposes a method for dynamic resource allocation under a microservice architecture, which includes the following steps: S1. Real-time collection of resource usage data during service operation, including the following operations: Deploy monitoring probes in the K8S cluster operating environment to obtain resource usage data at a preset sampling frequency, where: Deploy the Prometheus Operator in the Kubernetes cluster using Helm, and deploy a NodeExporter DaemonSet on each worker node to collect node-level metrics such as CPU usage, memory usage, and disk usage. Use the NVIDIA tool DCGM to collect GPU memory usage, utilization, and temperature-related metrics. Transfer the collected resource data to the next process.

[0020] S2. Clean and filter the data collected by the data acquisition module to generate a standardized time series data set. The specific operations include the following: Use Prometheus's Recording Rules to process the raw data to remove erroneous data and complete data cleaning; Exponential smoothing method is used to smooth the data, and polynomial interpolation filling method is used to handle missing data to ensure the integrity and availability of the data set. Data filtering and standardization are completed to finally form a standardized time series data set.

[0021] S3. For time series data sets, the TimesNet time series model is used to analyze resource usage trends and predict future resource requirements. Then, elastic scaling recommendations are generated based on preset business priority rules to provide data support for subsequent resource scheduling.

[0022] The TimesNet timing model involved analyzes resource usage trends and predicts future resource requirements through the following steps: (i) Use Fourier transform to convert the time series data contained in the time series dataset from the time domain to the frequency domain. By calculating the amplitudes of different frequency components, multiple cycles in the data that conform to the actual load pattern are screened out. Based on these cycles, the data is subjected to targeted preprocessing so that the preprocessed data retains the multi-cycle essential characteristics of the original time series data. Targeted preprocessing includes cycle alignment and segmentation, intra-cycle trend extraction and noise filtering, multi-cycle feature separation and enhancement, cycle consistency verification, and missing value repair. (ii) Introducing a deep residual network to perform deep feature extraction on the preprocessed data tensor, effectively capturing the complex patterns and dynamic changes hidden in time series data; (iii) Feature fusion is performed based on the importance weight of each cycle in resource changes, and the resource demand prediction result is finally output to adapt to the mixed cycle characteristics of microservice load.

[0023] This step generates elastic scaling recommendations based on the resource demand forecast results of the TimesNet timing model and preset service priority rules. The specific steps include the following: S3.1. Load the preset business priority rules, associate the resource demand forecast results output by the TimesNet timing model with the priority labels of the corresponding businesses, and clarify the resource demand trends of businesses with different priorities.

[0024] Pre-defined service priority rules are typically defined by users or system administrators based on factors such as business importance, service-level agreements (SLAs), and cost sensitivity. For example: a) Prioritization by business type: core transaction services (such as payments and orders) are assigned "P0," while non-core services (such as log analysis and report generation) are assigned "P1." b) Prioritization by SLA requirements: services requiring "99.99% availability" are prioritized over services requiring "99.9% availability." c) Prioritization by resource preemption rules: high-priority services receive priority in resource competition, while low-priority services may be temporarily downgraded or throttled. These rules are quantified into computable metrics (such as priority weights and resource guarantee thresholds) and stored in the rule engine.

[0025] The resource demand forecasts output by the TimesNet model (such as the peak and valley CPU / memory / GPU demand for a service within the next hour) are associated with the corresponding service priority labels to clarify the resource demand trends of services of different priorities. For example: i) A P0 service is predicted to have CPU demand increasing from 80% to 95% (approaching a resource bottleneck) within the next 30 minutes; ii) A P1 service is predicted to have memory demand remaining stable at around 50% within the next hour.

[0026] S3.2. Based on the business priority rules, set resource guarantee thresholds and scaling trigger conditions for different priority businesses to ensure that the resource requirements of high-priority businesses are met first.

[0027] For example: For P0 services: set "capacity expansion when CPU usage ≥ 80%" and "capacity reduction when memory usage ≤ 20%", and the idle resources of P1 services can be preempted during capacity expansion; For P1 services: Set "Capacity expansion triggered when CPU usage ≥ 90%" and "Capacity reduction triggered when memory usage ≤ 10%". Capacity expansion must be performed after P0 service resources are sufficient. When reducing capacity, non-essential resources must be released first.

[0028] S3.3. Combine the prediction results, priority rules, and resource thresholds to generate preliminary scaling recommendations. The recommendations are then optimized using preset logic to ultimately generate elastic scaling recommendations. These recommendations are then output to the resource scheduling module in a structured format (such as JSON instructions or API call parameters). Furthermore, the deviation between actual resource usage and the prediction results is continuously monitored. The scaling recommendations are adjusted based on real-time changes in business priorities (such as temporarily increasing the priority of a business) to ensure the flexibility and reliability of resource scheduling.

[0029] Combine the prediction results, priority rules, and resource thresholds to generate preliminary scaling recommendations. Then, optimize the scaling recommendations using pre-set logic. For example: For capacity expansion scenarios: If the resource demand forecast for a P0-level service is about to exceed the threshold, priority will be given to generating capacity expansion suggestions for it (such as increasing the number of Pod replicas or increasing resource quotas). If the current cluster resources are insufficient, it may be recommended to temporarily limit non-essential resource usage for P1-level services (such as reducing the number of replicas). For scaling-down scenarios: If the predicted resource demand of a low-priority service (such as P1) remains below the threshold for a long period of time, scaling-down recommendations are prioritized to free up resources. For high-priority services (such as P0), even if resource utilization is low, a certain amount of redundant resources must be retained to cope with sudden demand. Consider the balance between cost and efficiency: For services with lower priority but high resource consumption, scaling down should be triggered during low-business hours (such as nighttime) to avoid affecting normal business operations.

[0030] S4. Based on the elastic scaling recommendations and in combination with the pre-defined resource allocation strategy, actual resource allocation operations are performed.

[0031] Based on business priority classification, combined with predicted resource demand trends and business characteristics, differentiated resource allocation strategies are formulated in advance. The details are as follows: Define AI training / inference tasks as high-priority tasks, and define non-real-time data processing services, non-critical business auxiliary services, and backend analysis and report generation services as low-priority tasks; When the resource demand forecast output by the intelligent prediction module is greater than 1.1 times the current resource limit, resource adjustments are triggered for high-priority tasks: 80% of the forecasted resource demand is reserved for the task as a fixed quota, and an additional 20% of resources are allowed for elastic burst during peak demand. These resources come from the shared resource pool in the cluster, and elastic resources are managed through Kubernetes' LimitRange and ResourceQuota mechanisms. When the resource demand forecast output by the intelligent prediction module is greater than 1.3 times the current resource limit, resource adjustment is triggered for low-priority tasks: resources are dynamically allocated based on the actual forecast demand of the task, but the allocation upper limit does not exceed 130% of the forecast demand; soft limits and hard limits are set for the task: soft limits provide basic resource guarantees for task operation, and hard limits constrain the upper limit of resource usage; when cluster resources are tight, the resources of low-priority tasks that exceed the soft limit are reclaimed, but the task is not forcibly terminated.

[0032] This step dynamically adjusts the Pod's resource quota based on the pre-defined resource allocation strategy by calling the Kubernetes API and interacting with the underlying platform. This allows for real-time resource allocation, elastic scaling, and on-demand recycling. Ultimately, this results in dynamic optimization and efficient utilization of cluster resources, ensuring that resource requirements for high-priority tasks are met first.

[0033] Example 2: Combined with the attached Figure 2 This embodiment proposes a dynamic resource allocation system under a microservice architecture, which is used to implement the method described in Example 1, including: Data collection module, used to collect resource usage data in real time during service operation; The data processing module is used to clean and filter the data collected by the data acquisition module to generate a standardized time series data set; The intelligent prediction module uses the TimesNet time series model to analyze resource usage trends and predict future resource requirements for time series datasets. It then generates elastic scaling recommendations based on pre-set business priority rules, providing data support for resource scheduling. The resource scheduling module is used to allocate resources based on the elastic scaling recommendations generated by the intelligent prediction module and the pre-defined resource allocation strategy.

[0034] Embodiment 3: This embodiment also proposes a dynamic resource allocation device under a microservice architecture, which includes: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to implement the method described in the first embodiment.

[0035] Embodiment 4: This embodiment further provides a computer-readable medium having computer instructions stored thereon. When executed by a processor, the computer instructions can implement the method described in Embodiment 1. Specifically, a system or device can be provided with a storage medium. The storage medium stores software program code that implements the functions of any of the above embodiments, and the computer (or CPU or MPU) of the system or device can read and execute the program code stored in the storage medium.

[0036] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.

[0037] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, and DVD+RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer via a communications network.

[0038] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.

[0039] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.

[0040] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the scope of protection of the present invention.

Claims

1. A method for dynamic resource allocation under a microservice architecture, characterized in that: It includes the following steps: S1. Real-time collection of resource usage data during service operation; S2. Clean and filter the collected data from the data acquisition module to generate a standardized time series data set; S3: For time series datasets, the TimesNet time series model is used to analyze resource usage trends and predict future resource demand. Then, based on pre-set business priority rules, elastic scaling recommendations are generated to provide data support for subsequent resource scheduling. S4. Based on the elastic scaling recommendations and in combination with the pre-defined resource allocation strategy, actual resource allocation operations are performed.

2. The method for dynamic resource allocation under a microservice architecture according to claim 1, characterized in that: The step S1 specifically includes the following operations: Deploy monitoring probes in the K8S cluster operating environment to obtain resource usage data at a preset sampling frequency, where: Deploy the Prometheus Operator in the Kubernetes cluster using Helm, and deploy a NodeExporter DaemonSet on each worker node to collect node-level metrics such as CPU usage, memory usage, and disk usage. Use the NVIDIA tool DCGM to collect GPU memory usage, utilization, and temperature-related metrics. Transfer the collected resource data to the next process.

3. The method for dynamic resource allocation under a microservice architecture according to claim 1, characterized in that: The step S2 specifically includes the following operations: Use Prometheus's Recording Rules to process the raw data to remove erroneous data and complete data cleaning; Exponential smoothing method is used to smooth the data, and polynomial interpolation filling method is used to handle missing data to ensure the integrity and availability of the data set. Data filtering and standardization are completed to finally form a standardized time series data set.

4. The method for dynamic resource allocation under a microservice architecture according to claim 1, characterized in that: The TimesNet timing model analyzes resource usage trends and predicts future resource requirements through the following steps: The Fourier transform is used to convert the time series data contained in the time series data set from the time domain to the frequency domain. By calculating the amplitudes of different frequency components, multiple cycles in the data that conform to the actual load pattern are screened out. Based on these cycles, the data is preprocessed in a targeted manner so that the preprocessed data retains the multi-cycle essential characteristics of the original time series data. A deep residual network is introduced to perform deep feature extraction on preprocessed data tensors, effectively capturing the complex patterns and dynamic changes hidden in time series data; Feature fusion is performed based on the importance weight of each cycle in resource changes, and the resource demand prediction result is finally output to adapt to the mixed cycle characteristics of microservice load.

5. The method for dynamic resource allocation under a microservice architecture according to claim 4, characterized in that: Execute step S3 to generate elastic scaling recommendations based on the resource demand forecast results of the TimesNet timing model and the preset service priority rules. The specific operations include the following: Load the preset business priority rules, associate the resource demand forecast results output by the TimesNet timing model with the priority labels of the corresponding businesses, and clarify the resource demand trends of businesses with different priorities; Based on business priority rules, resource guarantee thresholds and scaling trigger conditions are set for businesses of different priorities to ensure that resource requirements of high-priority businesses are met first. Combining the prediction results, priority rules, and resource thresholds, it generates preliminary scaling recommendations. It then optimizes the recommendations using preset logic, ultimately generating elastic scaling recommendations that are output to the resource scheduling module in a structured form. At the same time, it continuously monitors the deviation between actual resource usage and the prediction results, and dynamically adjusts the scaling recommendations based on real-time changes in business priorities.

6. The method for dynamic resource allocation under a microservice architecture according to claim 1, characterized in that: Based on business priority classification, combined with predicted resource demand trends and business characteristics, differentiated resource allocation strategies are formulated in advance. The details are as follows: Define AI training / inference tasks as high-priority tasks, and define non-real-time data processing services, non-critical business auxiliary services, and backend analysis and report generation services as low-priority tasks; When the resource demand forecast output by the intelligent prediction module is greater than 1.1 times the current resource limit, resource adjustments are triggered for high-priority tasks: 80% of the forecasted resource demand is reserved for the task as a fixed quota, and an additional 20% of resources is allowed for elastic burst during peak demand. When the resource demand forecast output by the intelligent prediction module is greater than 1.3 times the current resource limit, resource adjustment is triggered for low-priority tasks: resources are dynamically allocated based on the actual forecast demand of the task, but the allocation upper limit does not exceed 130% of the forecast demand; soft limits and hard limits are set for the task: soft limits provide basic resource guarantees for task operation, and hard limits constrain the upper limit of resource usage; when cluster resources are tight, the resources of low-priority tasks that exceed the soft limit are reclaimed, but the task is not forcibly terminated.

7. The method for dynamic resource allocation under a microservice architecture according to claim 6, characterized in that: Execute step S4. Based on the pre-defined resource allocation strategy, the Kubernetes API is called to interact with the underlying platform to dynamically adjust the Pod's resource quota, achieve real-time resource allocation, elastic scaling, and on-demand recycling, and ultimately achieve dynamic optimization and efficient utilization of cluster resources, ensuring that the resource requirements of high-priority tasks are met first.

8. A dynamic resource allocation system under a microservice architecture, characterized in that: It is used to implement the method according to any one of claims 1 to 7, comprising: Data collection module, used to collect resource usage data in real time during service operation; The data processing module is used to clean and filter the data collected by the data acquisition module to generate a standardized time series data set; The intelligent prediction module uses the TimesNet time series model to analyze resource usage trends and predict future resource requirements for time series datasets. It then generates elastic scaling recommendations based on pre-set business priority rules, providing data support for resource scheduling. The resource scheduling module is used to allocate resources based on the elastic scaling recommendations generated by the intelligent prediction module and the pre-defined resource allocation strategy.

9. A dynamic resource allocation device under a microservice architecture, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to implement the method according to any one of claims 1 to 7.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, can implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Micro-service workload capacity expansion and contraction method and system in hybrid cloud environment

    CN115686828A

  • Dynamic interval elastic capacity expansion and contraction method and system based on long-time sequence prediction

    CN117056021A

  • Container energy-saving elastic capacity expansion and contraction method and system based on time sequence prediction and medium

    CN118260021A

  • Cloud data center load prediction method and system based on improved TimesNet

    CN119938463A

  • High-efficiency computer software system resource scheduling method, system, equipment and medium

    CN120315873A

Cited By

  • Elastic resource beforehand early warning and scheduling method and system based on load prediction

    CN120909742A

  • Intelligent resource scheduling system and method based on elastic threshold and AI prediction

    CN120909743A

  • An elastic threshold and AI prediction-based resource intelligent scheduling system and method

    CN120909743B

  • Stream batch integrated task dynamic preemption scheduling method and device

    CN121166385A

  • Server resource processing method and electronic equipment

    CN121579179A