Method and system for dynamically allocating cloud platform resources

By monitoring the resource usage data of virtual machines in the cloud platform, evaluating timely indicators and screening and coordinating virtual machines, resource allocation when the second exceeds the standard is achieved, the problem of lagging resource adjustment in the existing technology is solved, and resource utilization and service quality are improved.

CN120179424AActive Publication Date: 2025-06-20YUNBIAN CLOUD TECHNOLOGY (SHANGHAI) CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510660868.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The dynamic allocation method of existing cloud platform resources relies on real-time monitoring data, and cannot predict future changes in resource demand, resulting in lagging resource adjustments in the face of sudden business peaks, affecting virtual machine performance and terminal service quality.

Method used

By monitoring the resource usage data of the virtual machine, when the first exceeds the standard condition is met, the access terminal operation data of other virtual machines is obtained, the timeliness indicators are evaluated, the coordination virtual machine is filtered, and when the second exceeds the standard condition is met, the coordination virtual machine is controlled to allocate computing resources to the virtual machine that needs resources.

Benefits of technology

It has realized forward-looking resource allocation, reduced the lag of resource adjustment, improved resource utilization, enhanced system adaptability, and effectively guaranteed the service quality of the cloud platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179424A_ABST
    Figure CN120179424A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of cloud computing, and provides a cloud platform resource dynamic allocation method and allocation system. The method comprises the steps that a cloud platform monitors resource use data of all virtual machines, and when a first virtual machine with the resource use data meeting a first standard exceeding condition appears, operation data of access terminals of other virtual machines are acquired; obtaining a timeliness index through evaluation based on each operation data corresponding to other virtual machines, and screening a plurality of virtual machines with the timeliness indexes lower than an index threshold value as coordinated virtual machines; and the cloud platform continuously monitors the first virtual machine according to a set monitoring interval, and controls each coordination virtual machine to allocate computing resources to the first virtual machine when the resource use data of the first virtual machine meets a second standard exceeding condition. According to the method, prospective resource allocation can be realized, hysteresis is reduced, and the service quality of the cloud platform is effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and in particular, to a method and system for dynamically allocating cloud platform resources. Background Art

[0002] At present, with the rapid development of cloud computing technology, cloud platforms have become the core infrastructure for many enterprises and users to process data and run application programs. Currently, the architecture model widely adopted by cloud platforms to provide data processing services for each terminal is as follows: the cloud platform divides physical servers into multiple virtual machines, each virtual machine manages at least one terminal, and the cloud platform can monitor each virtual machine in real time.

[0003] To ensure the performance of virtual machines, the cloud platform will monitor the resource usage of each virtual machine in real time. When it detects that the resource usage of a certain virtual machine exceeds a preset threshold, it will enable a resource scheduling algorithm to dynamically allocate other idle resources to this virtual machine. Common dynamic allocation algorithms, such as the greedy algorithm, genetic algorithm, simulated annealing algorithm, etc., all perform resource allocation based on the real-time monitored resource usage status.

[0004] However, the above-mentioned resource dynamic allocation method has significant drawbacks. Since it completely relies on real-time monitoring data and cannot predict future resource demand changes and cannot pre-adjust resource allocation, when facing sudden business peaks or resource demand changes, resource adjustment is often lagging, resulting in poor resource adjustment effects. For example, in business scenarios such as e-commerce promotions and live streaming with goods, the number of terminal accesses will surge instantly. Due to the lack of pre-adjustment in the existing dynamic allocation method, resource supply is often not timely, resulting in a decline in the performance of virtual machines, which in turn affects the quality of terminal services, causing problems such as poor user experience and business losses.

[0005] Therefore, there is an urgent need for a cloud platform resource dynamic allocation method that can effectively solve the above problems. Summary of the Invention

[0006] In view of the above technical problems, the present invention provides a method, system, electronic device, computer storage medium, and computer program product for dynamically allocating cloud platform resources.

[0007] The present invention discloses a method for dynamically allocating cloud platform resources, and the method includes the following steps: The cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine whose resource usage data meets the first over-standard condition appears, it obtains the operation data of the access terminals of other virtual machines; Based on the respective operation data corresponding to other virtual machines, an timeliness index used to characterize the expectation of the access terminal's response speed to the virtual machine is evaluated, and several virtual machines with timeliness indexes lower than the index threshold are selected as coordinated virtual machines; The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.

[0008] The present invention also discloses a cloud platform resource dynamic allocation system, which includes a processing device and a storage device. The computer code stored in the storage device is called and executed by the processing device to implement the following steps: The cloud platform monitors the resource usage data of each virtual machine. When there is a first virtual machine whose resource usage data meets the first over-standard condition, it obtains the operation data of the access terminals of other virtual machines respectively; Based on the respective operation data corresponding to other virtual machines, an timeliness index used to characterize the expectation of the access terminal's response speed to the virtual machine is evaluated, and several virtual machines with timeliness indexes lower than the index threshold are selected as coordinated virtual machines; The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.

[0009] The present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. The processor executes the computer program to implement the method as described in any one of the foregoing.

[0010] The present invention also discloses a computer storage medium, which stores a computer program. The computer program is executed by a processor to implement the method as described in any one of the foregoing.

[0011] The present invention also discloses a computer program product, which contains computer code. When the computer code is executed by the processor of an electronic device, it implements the method as described in any one of the foregoing.

[0012] The beneficial effects of the present invention are at least as follows: The present invention first identifies a first virtual machine with a trend of resource tension in advance according to a first over-standard condition, screens out a coordinated virtual machine by evaluating the operation data of the access terminal, and then triggers the allocation of resources according to a second over-standard condition. The present invention realizes forward-looking resource allocation, reduces lag, improves resource utilization rate based on the needs of the access terminal, enhances system adaptability through multi-dimensional monitoring, and effectively guarantees the service quality of the cloud platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 is a flowchart showing the method for dynamically allocating cloud platform resources disclosed in the embodiments of the present invention; Figure 2 is a flowchart showing the process of extracting each operation feature from the target operation data using a combination algorithm disclosed in the embodiments of the present invention; Figure 3 is a structural diagram of a system for dynamically allocating cloud platform resources disclosed in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.

[0016] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.

[0017] In the field of dynamic allocation of cloud platform resources, the prior art completely relies on real-time monitoring data and cannot predict future resource demand changes, resulting in a lag in resource adjustment when facing sudden business peaks. For example, during a major e-commerce promotion, the CPU usage rate of a certain virtual machine instantly soars from 30% to 80%. The traditional method can only start resource allocation when the usage rate exceeds a preset threshold (such as 85%). During this period, service response may be slow due to insufficient resources, affecting the user experience.

[0018] To solve this problem, the present invention proposes a new method for dynamic resource allocation, and it is based on the existing architecture model in which physical servers are divided into multiple virtual machines on a cloud platform, and each virtual machine manages at least one terminal and monitors the virtual machines (especially the activation status) in real time.

[0019] As Figure 1 shown, an embodiment of the present invention discloses a method for dynamic resource allocation on a cloud platform, and the method includes the following steps: S101, the cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine whose resource usage data meets the first over-standard condition appears, obtain the running data of the access terminals of other virtual machines respectively.

[0020] The cloud platform continuously monitors the resource usage data of each virtual machine (such as CPU usage rate, memory occupancy, disk I / O read and write rate, network bandwidth utilization rate, etc.). When the resource usage data of a certain virtual machine meets the first over-standard condition, it is designated as the first virtual machine and further operations are triggered. For example: (1) In an online education cloud platform, virtual machine A responsible for real-time course live broadcast undertakes the tasks of live streaming and interactive data processing for multiple classes. During the prime learning period in the evening, a large number of students watch the courses online at the same time and ask questions in real time, resulting in a continuous increase in the CPU usage rate of virtual machine A. When the CPU usage rate exceeds 70% set in the first over-standard condition for 10 consecutive minutes, for example, reaches 75%, at this time, the resource usage data of virtual machine A meets the first over-standard condition. Excessive CPU usage rate will cause problems such as stuttering of the live broadcast screen and delay of interactive messages, affecting the learning experience of students, and further resource allocation is required.

[0021] (2) During the big promotion activity of an e-commerce platform, virtual machine B used for order processing and inventory management has its memory occupancy continuously rising due to a large number of order placement requests and inventory query operations pouring in within a short period of time. Virtual machine B is initially allocated 8GB of memory. When the memory occupancy exceeds 6.4GB (i.e., 80% of the memory occupancy ratio) set in the first over-standard condition for 15 consecutive minutes and reaches 6.6GB, it indicates that the memory resources of this virtual machine are in short supply. Insufficient memory may lead to slower order processing speed, and even serious problems such as order loss and inventory data confusion. Therefore, after meeting the first over-standard condition, the subsequent resource allocation process needs to be started.

[0022] (3) In the content storage and distribution system of the short video platform, virtual machine C is mainly responsible for storing the videos uploaded by users and distributing popular videos. When a certain popular video goes viral, a large number of users request to play the video simultaneously, causing the disk I / O read rate of virtual machine C to increase sharply. When the disk I / O read rate exceeds 400 MB / s set in the first over-standard condition for 20 consecutive minutes and reaches 450 MB / s, it indicates that the disk read and write pressure is too high. An excessively high disk I / O rate will reduce the video loading speed, resulting in frequent buffering when users watch videos. At this time, the resource usage data of virtual machine C meets the first over-standard condition, and resource adjustment is required.

[0023] (4) In the game cloud platform, virtual machine D provides game data interaction services for numerous online players. When a new game version is launched, a large number of players log in to the game simultaneously for updates and play, causing the network bandwidth utilization rate of virtual machine D to rise rapidly. When the network bandwidth utilization rate exceeds 80% set in the first over-standard condition for 30 consecutive minutes and reaches 85%, it indicates that the network transmission capacity is approaching saturation. Tight network bandwidth will cause excessive game latency, untimely response to player operations, and even disconnection. Therefore, virtual machine D meets the first over-standard condition, and the cloud platform needs to obtain the operation data of other virtual machine terminals and prepare for dynamic resource allocation.

[0024] After the resource usage data of the first virtual machine meets the first over-standard condition, the cloud platform obtains the operation data of the access terminals of other virtual machines respectively to provide a more comprehensive decision-making basis for subsequent resource scheduling. By obtaining the operation data of the access terminals, the cloud platform can more accurately evaluate the resource allocation potential of each virtual machine, so as to screen out suitable coordinating virtual machines in the subsequent steps, avoid blindly allocating resources and affecting the normal operation of other services, and achieve more reasonable resource allocation.

[0025] S102, based on the respective operation data corresponding to other virtual machines, evaluate a timeliness index used to represent the expectation of the response speed of the access terminal to the virtual machine, and screen out several virtual machines with timeliness indexes lower than the index threshold as coordinating virtual machines.

[0026] Based on the operation data corresponding to each other virtual machine, evaluate the timeliness index. Among them, the timeliness index represents the speed at which each access terminal requires the virtual machine to respond to its data processing and data acquisition requests (which can be represented by the response duration), and this requirement for the response speed is obtained based on the real-time operation data analysis of each terminal. That is, in different operation states, the timeliness index is actually different and needs to be analyzed in real time.

[0027] Under the existing architecture, the mode of each virtual machine management terminal makes the terminal operation data closely associated with the virtual machine. The cloud platform can accurately identify which virtual machines can be used to allocate some resources to the first virtual machine without affecting the quality of its terminal services by screening and coordinating the virtual machines through calculating the timeliness index, so as to ensure that the first virtual machine can provide computing services for each access terminal accessing it. Among them, the screening of the coordinated virtual machines should follow the principle of being sufficient, and over-screening should be avoided to reduce the impact on the normal business operation of other virtual machines.

[0028] S103. The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.

[0029] In the aforementioned step S102, the cloud platform has screened out each coordinated virtual machine that can be used to allocate resources to the first virtual machine. However, the first over-standard condition is actually set slightly lower to realize earlier screening of the coordinated virtual machines for the first virtual machine. After that, the cloud platform continues to monitor the first virtual machine. When its resource usage data meets a higher second over-standard condition (i.e., more serious over-standard), it controls each coordinated virtual machine to allocate computing resources to the first virtual machine. Among them, the first over-standard condition and the second over-standard condition can be comprehensively determined based on the historical operation data and management requirements of the cloud platform, and will not be elaborated here.

[0030] The cloud platform dynamically adjusts the allocation of virtual machines according to the actual needs of the terminal access type and quantity. For example, the cloud platform divides out a sufficient number of virtual machines at one time, but only activates a part of them based on the actual terminal access type and quantity; as the number of accessed terminals increases, the cloud platform will gradually activate some dormant virtual machines, or re-divide and activate some new virtual machines to meet the actual needs. And these newly added (i.e., newly activated) virtual machines will make the cloud platform have a greater redundancy when coordinating resources for the first virtual machine. The present invention sets the monitoring interval for adjusting whether or when the first virtual machine reaches the second over-standard condition in this stage of the first over-standard condition and the second over-standard condition according to this redundancy. It can be understood that the monitoring interval is positively correlated with the number of newly added virtual machines.

[0031] The present invention first identifies the first virtual machine with a resource shortage trend in advance with the first over-standard condition, screens out the coordinated virtual machines by evaluating the operation data of the access terminals, and then triggers the allocation of resources according to the second over-standard condition. The present invention realizes forward-looking resource allocation, reduces lag, improves resource utilization rate based on the needs of the access terminals, enhances system adaptability through multi-dimensional monitoring, and effectively guarantees the service quality of the cloud platform.

[0032] Optionally, before obtaining the running data of the access terminals of other virtual machines respectively, the method further includes: Determine the newly added virtual machines that are in an active state and available for invocation on the cloud platform, calculate the total resources of these newly added virtual machines, and if the total resources can ensure that the resource usage data of the first virtual machine does not meet the second over-standard condition, then no trigger signal is generated; otherwise, a trigger signal is generated. The trigger signal is used to trigger the obtaining of the running data of the access terminals of other virtual machines respectively.

[0033] In this embodiment, during the operation of the cloud platform, the virtual machine allocation is dynamically adjusted according to the terminal access type and quantity. For example, dormant virtual machines are gradually activated or new virtual machines are partitioned, and the virtual machines in an active state without terminal access for the time being also belong to the newly added virtual machines. At this time, the cloud platform first determines the newly added virtual machines that are currently available for invocation and in an active state. For example, before a major e-commerce promotion, the cloud platform pre-activates 10 virtual machines for coping with traffic peaks, and these 10 are the newly added virtual machines. Then, the cloud platform calculates the total resources of these newly added virtual machines (partitioned by the cloud platform when they are generated), including resource data such as the total number of CPU cores, the total memory capacity, the available disk space, and the total network bandwidth. Suppose these 10 newly added virtual machines have a total of 20 CPU cores, 80 GB of memory, 1 TB of disk space, and 10 Gbps of network bandwidth.

[0034] Associate and evaluate the calculated total resources of the newly added virtual machines with the resource requirements of the first virtual machine (the difference between the predicted resource demand of the first virtual machine and the total resources allocated to the first virtual machine). The basis for the evaluation is to determine whether these newly added resources can ensure that the resource usage data of the first virtual machine does not meet the second over-standard condition. If the total resources of the newly added virtual machines are sufficient and can ensure that the resource usage data of the first virtual machine will not reach the standard of the second over-standard condition within a foreseeable time range (such as 10 minutes) through resource supplementation, then no trigger signal is generated. At this time, there is no need to further obtain the running data of other virtual machine terminals and start the subsequent complex resource allocation process.

[0035] For example, if the second over-standard condition is set that the CPU usage rate exceeds 85%, and according to the calculation, after allocating the 20 CPU cores of the newly added virtual machines to the first virtual machine, its CPU usage rate can be controlled below 80%, that is, it does not meet the second over-standard condition, and at this time, no trigger signal is generated.

[0036] On the contrary, if the total resources of the newly added virtual machines are not sufficient to prevent the first virtual machine from reaching the second over-standard condition, then a trigger signal is generated, and subsequent operations such as obtaining the running data of the access terminals of other virtual machines respectively will be executed to further screen and coordinate virtual machines and allocate resources to ensure the performance of the first virtual machine.

[0037] Optionally, the timeliness metric used to characterize the expected response speed of the terminal to the virtual machine, which is evaluated based on the respective operation data corresponding to other virtual machines, includes: The target operation data is obtained by intercepting the operation data of each access terminal corresponding to other virtual machines according to the acquisition duration, and several data types are matched according to the pre-configured tags of the other virtual machines; wherein, the acquisition duration is determined based on the terminal type of the access terminal; The combination algorithm is used to extract the operation characteristics corresponding to each data type from the target operation data, and each operation characteristic is spliced into a target operation characteristic; The GNN-Transformer evaluator is used to classify and evaluate the target operation characteristics to obtain the timeliness metric used to characterize the expected response speed of the terminal to the virtual machine.

[0038] In this embodiment, the cloud platform or the virtual machine continuously records the data interaction records between the virtual machine and each access terminal to form operation data. However, there are a lot of early data in these operation data, which is not helpful for analyzing the current stage's expectation of the response speed of the access terminal. Therefore, it is necessary to determine an appropriate acquisition duration, and intercept a section from the operation data according to this acquisition duration (the interception end point is the current moment, and the interception start point is the historical moment), that is, the target operation data.

[0039] Among them, the acquisition duration is related to the terminal type of the access terminal. For example, for Web terminals with high-frequency real-time interaction (such as e-commerce real-time transaction terminals), due to their business characteristics, they have high requirements for instant response, and early data (such as request delay records more than 5 minutes ago) are difficult to reflect the current user experience requirements. Therefore, a 5-minute acquisition duration is set, and only data such as request interval sequences and timeout retry times within the most recent 5 minutes are intercepted to focus on real-time interaction characteristics. For IoT terminals with low-frequency batch processing (such as sensor data upload terminals), their data interaction is periodic and has a higher tolerance for response latency. Early data (such as transmission time-consuming records within 30 minutes) can still reflect the current stage's periodic law. Therefore, a 30-minute acquisition duration is set, and data such as batch data transmission success rate and processing time-consuming within this period are intercepted to retain trend characteristics.

[0040] By binding the acquisition duration to the terminal type, the cloud platform can accurately filter out invalid historical data, ensuring that the target operation data used to evaluate the timeliness metric only contains effective information highly relevant to the current terminal requirements, and avoiding evaluation deviation caused by mixing in outdated data. It can be understood that the terminal type is not limited to the above types, and there are also multiple corresponding acquisition durations. A terminal type - acquisition duration comparison table can be established in advance, which will not be elaborated here.

[0041] Next, after filtering out the target operation data by acquisition duration, first use a combination algorithm to extract the operation characteristics of sub-data corresponding to the specified data type from each target operation data. The data types are shown in Table 1 below: Among them, when the cloud platform generates a virtual machine, it can configure one or more pre-configured tags representing different interaction attributes for it. These tags include business real-time tags, data interaction mode tags, user experience sensitive tags, resource allocation policy tags, etc. (each tag corresponds to multiple sub-tags, such as real-time business tags, non-real-time business tags). Based on the association relationship between these tags and data types, multiple data types corresponding to specific other virtual machines can be determined for which operation characteristics need to be extracted. Specifically as follows: 1. Business real-time tag Businesses with strong real-time requirements (such as online live broadcasts, financial transactions): need to focus on millisecond-level response speed, so pay attention to data directly reflecting immediate interaction such as request interval sequences, timeout retry counts, and real-time business tags.

[0042] Non-real-time businesses (such as batch data processing, log archiving): allow a certain delay and are more concerned with trend indicators (such as batch transmission success rate, processing time), and do not require high-frequency real-time data.

[0043] 2. Data interaction mode tag High-frequency burst interaction businesses (such as e-commerce flash sales, game servers): need to analyze the burst request frequency (QPS peak) and user operation sequences to identify resource preemption risks and optimize response priorities.

[0044] Low-frequency periodic businesses (such as IoT sensor data upload): focus on periodic transmission time and historical feedback data to evaluate long-term stability rather than instantaneous response.

[0045] 3. User experience sensitive tag Interaction-intensive businesses (such as online collaboration tools): users are sensitive to the continuous operation interval (such as editing delay), and it is necessary to monitor user operation sequences and terminal device attributes (such as network mode) to avoid experience degradation caused by device performance or network fluctuations.

[0046] Non-interaction businesses (such as background data synchronization): are more concerned with business type tags (non-real-time) and historical feedback data (such as task failure rate in implicit feedback) to ensure task integrity rather than real-time response.

[0047] 4. Resource allocation policy tag High-priority services (such as medical image analysis): It is necessary to judge the impact of response latency on the service through the number of timeout retries and historical feedback data, and allocate computing resources preferentially to ensure timeliness.

[0048] Low-priority services (such as advertising push): Higher latency can be tolerated, and key metrics related to cost (such as resource utilization in the request interval sequence) are monitored, and the response speed and resource consumption are balanced.

[0049] After extracting the running characteristics of each access terminal, these running characteristics are spliced into target running characteristics by means such as weighted splicing and sliding window splicing. Then, a GNN-Transformer evaluator is used to classify and evaluate the target running characteristics to obtain a timeliness index for characterizing the expected response speed of the terminal to the virtual machine.

[0050] Optionally, if the combined algorithm is wavelet transform + CNN + Transformer, the combined algorithm is used to extract the running characteristics corresponding to each data type from the target running data, including: Using wavelet transform to perform multi-scale decomposition on the sub-data corresponding to the data type in the target running data, extracting the high-frequency component and the low-frequency component, and combining the components after normalization to form preliminary time-series features; Adopting a dynamic adjustment of the number of convolutional kernels of the one-dimensional convolutional layer to capture patterns in different time windows, using the one-dimensional convolutional layer of CNN to extract local features from the time-series features, and then reducing the dimension through a pooling layer to retain key features, obtaining target time-series features; Using a Transformer encoder to calculate the correlation weights between the feature dimensions in the target time-series features by means of the multi-head self-attention mechanism, generating a context-aware feature vector that fuses time-series and semantic information; Splicing the target time-series features and the context-aware feature vector into running features according to the data type.

[0051] In this embodiment, since the target running features obtained in the foregoing steps contain multi-type and multi-granularity data, such as time-series metric features (request interval sequence, response time fluctuation, number of timeout retries, etc.), device attribute features (terminal type (mobile phone / PC), network mode (5G / Wi-Fi), etc.), and some data also has characteristics such as long-distance dependence. Traditional single algorithms are difficult to capture high-frequency burst features (such as the instantaneous QPS peak in a flash sale scenario), low-frequency trend features (such as the periodic data transmission law of IoT terminals), and global correlation relationships (such as the impact of the time difference between "adding to cart - settlement" in the user operation sequence on resource requirements) at the same time. Thus, as Figure 2As shown in the figure, the combined algorithm composed of wavelet transform + CNN + Transformer in the present invention extracts features with different granularities layer by layer from the target operation features through multi-stage processing, and then obtains operation features that can better reflect the requirements of response timeliness. Specifically: First, the sub-data corresponding to each data type in the target operation data is decomposed into high-frequency components (reflecting short-term fluctuations and sudden changes) and low-frequency components (reflecting long-term trends and periodic patterns) through wavelet transform. For example, in the request interval sequence during a major e-commerce promotion, the high-frequency component corresponds to the intensive requests at the moment of flash sales, and the low-frequency component corresponds to the overall traffic trend during the promotion period. Normalize each component (for example, using Z-score standardization) to eliminate the dimension difference, and then combine them into preliminary time series features containing multi-scale information.

[0052] After wavelet transform processing, signals with different time scales can be separated, avoiding high-frequency noise masking low-frequency trends or low-frequency trends blurring high-frequency details, and providing a cleaner multi-granularity time series input for subsequent CNN and Transformer.

[0053] Then, dynamically adjust the number of convolutional kernels in the one-dimensional convolutional layer of CNN, and scan the preliminary time series features through sliding windows of different sizes (such as kernel size = 3, 5, 7) to extract local features, such as continuous high-frequency requests and periodic retry peaks. Use max pooling or average pooling to reduce the feature dimension and retain key features, such as the maximum QPS value and average response latency within a certain time window. Take the key features as the target time series features.

[0054] Next, use the Transformer encoder to perform global dependency modeling on the target time series features, calculate the correlation weights of each time point and each dimension (such as request interval, retry count, business type label) in the feature vector using the multi-head self-attention mechanism, and capture long-distance dependency relationships (such as the resource demand correlation between the user's "add to cart" operation and the "checkout" operation 10 minutes later). At the same time, inject position information (such as timestamp) into the time series features, that is, generate a context-aware feature vector that fuses time series and semantic information. This context-aware feature vector fuses global time series dependencies and semantic associations (such as "there is a strong positive correlation between high-frequency request periods and continuous user operations, and resources need to be guaranteed first").

[0055] Finally, splice the target time series features (local features) extracted by CNN and the context-aware feature vector (global dependency features) generated by Transformer according to the data type (such as request interval sequence, timeout retry count) to obtain operation features. After obtaining the operation features corresponding to all access terminals, splice them end to end to obtain the target operation features.

[0056] Optionally, the method of dynamically adjusting the number of convolutional kernels of the one-dimensional convolutional layer to capture patterns in different time windows includes: Determine an adjustment span based on the time scale type and fluctuation frequency type of the sub-data, and dynamically adjust the number of convolutional kernels of the one-dimensional convolutional layer according to the adjustment span to capture patterns in different time windows; wherein, the time scale type includes long-period fluctuations, medium-period fluctuations, and short-period fluctuations, and the frequency type includes high frequency, medium frequency, and low frequency.

[0057] In this embodiment, the time scale characteristics and fluctuation frequency characteristics of sub-data of different data types are different. Refer to the example shown in Table 2 below: Table 2 The adjustment span (for example, the step size ΔK when adjusting from the initial value K0 to K1) of the number of convolutional kernels set in the present invention needs to match the time scale type and fluctuation frequency type of the data type. It can be understood that Table 2 above is only an example for some data types and is not used to limit the data types to only those shown in Table 2.

[0058] For example: For high-frequency short-period data, use small-span dynamic adjustment (for example, ΔK = 2) to ensure fine changes in the number of convolutional kernels and cover more small-window patterns. For low-frequency long-period data (such as device attributes), use large-span dynamic adjustment (for example, ΔK = 3) to reduce redundant calculations and focus on key long-window patterns. For medium-frequency medium-period data (such as business behaviors), use medium-span dynamic adjustment (for example, ΔK = 4) to balance calculation efficiency and feature coverage.

[0059] It should be noted that to avoid missing features due to too large a span setting (for example, for low-frequency data, if the span is 50% of K0, key windows may be skipped), batch normalization can be combined to alleviate the distribution shift problem caused by the adjustment of the number of kernels.

[0060] As Figure 3 shown, the embodiment of the present invention also discloses a system 100 for dynamic allocation of cloud platform resources. The system includes a processing device 200 and a storage device 300. The computer code stored in the storage device 300 is called and executed by the processing device 200 to implement the following steps: The cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine whose resource usage data meets the first over-standard condition appears, obtain the running data of the access terminals of other virtual machines respectively; Based on the respective running data corresponding to other virtual machines, evaluate a timeliness index used to characterize the expected response speed of the access terminal to the virtual machine, and screen several virtual machines with timeliness indexes lower than the index threshold as coordinated virtual machines; The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine. Wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.

[0061] An embodiment of the present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, where the processor executes the computer program to implement the method as described in the foregoing embodiments.

[0062] An embodiment of the present invention also discloses a computer storage medium, where the computer storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in the foregoing embodiments.

[0063] An embodiment of the present invention also discloses a computer program product, where the computer program product contains computer code, and when the computer code is executed by a processor of an electronic device, it implements the method as described in the foregoing embodiments.

[0064] The above-mentioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the above. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0065] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitations are imposed herein.

[0066] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for dynamic allocation of cloud platform resources, characterized in that: The method includes the following steps: The cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine whose resource usage data meets the first over-standard condition appears, it obtains the running data of the access terminals of the other virtual machines respectively; Based on the respective running data corresponding to the other virtual machines, an timeliness index for characterizing the expected response speed of the access terminal to the virtual machine is evaluated, and several virtual machines with timeliness indexes lower than the index threshold are selected as coordinated virtual machines; The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.

2. The method for dynamic allocation of cloud platform resources according to claim 1, characterized in that: Before obtaining the running data of the access terminals of the other virtual machines respectively, the method further includes: Determine the newly added virtual machines in the cloud platform that are in an active state and available for invocation, calculate the total resources of these newly added virtual machines. If the total resources can make the resource usage data of the first virtual machine not meet the second over-standard condition, no trigger signal is generated, otherwise a trigger signal is generated; the trigger signal is used to trigger the obtaining of the running data of the access terminals of the other virtual machines respectively.

3. The method for dynamic allocation of cloud platform resources according to claim 1, characterized in that: Evaluating an timeliness index for characterizing the expected response speed of the terminal to the virtual machine based on the respective running data corresponding to the other virtual machines includes: Obtain target running data by intercepting from the running data of the access terminals corresponding to the other virtual machines according to the acquisition duration, and match several data types according to the pre-configured tags of the other virtual machines; wherein, the acquisition duration is determined based on the terminal type of the access terminal; Use a combination algorithm to extract the running features corresponding to each data type from the target running data, and splice the running features into a target running feature; Use a GNN-Transformer evaluator to classify and evaluate the target running feature to obtain an timeliness index for characterizing the expected response speed of the terminal to the virtual machine.

4. The method for dynamic allocation of cloud platform resources according to claim 3, characterized in that: If the combination algorithm is wavelet transform + CNN + Transformer, then using the combination algorithm to extract the running features corresponding to each data type from the target running data includes: Use wavelet transform to perform multi-scale decomposition on the sub-data corresponding to the data type in the target running data, extract the high-frequency component and the low-frequency component, and combine the components after normalization processing into a preliminary time series feature; Adopt a dynamic adjustment of the convolution kernel number of the one-dimensional convolutional layer to capture the patterns of different time windows, use the one-dimensional convolutional layer of CNN to extract local features from the time series feature, and then reduce the dimension through a pooling layer to retain the key features to obtain the target time series feature; The Transformer encoder calculates the correlation weights between each feature dimension using the multi-head self-attention mechanism to generate a context-aware feature vector that fuses time series and semantic information; Splice the target time series feature and the context-aware feature vector into a running feature according to the data type.

5. The method for dynamic allocation of cloud platform resources according to claim 4, characterized in that: The method of dynamically adjusting the number of convolution kernels of a one-dimensional convolutional layer to capture patterns in different time windows includes: Determining an adjustment span according to the time scale type and fluctuation frequency type of sub-data, and dynamically adjusting the number of convolution kernels of the one-dimensional convolutional layer according to the adjustment span to capture patterns in different time windows; wherein, the time scale type includes long-period fluctuations, medium-period fluctuations, and short-period fluctuations, and the frequency type includes high frequency, medium frequency, and low frequency.

6. A system for dynamic allocation of cloud platform resources, the system includes a processing device and a storage device, characterized in that: The computer code stored in the storage device is called and executed by the processing device to implement the following steps: The cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine with resource usage data meeting the first over-standard condition appears, it obtains the operation data of the access terminals of other virtual machines respectively. Evaluating, based on the respective operation data corresponding to other virtual machines, a timeliness index used to characterize the expected response speed of the access terminal to the virtual machine, and screening several virtual machines with timeliness indexes lower than the index threshold as coordinated virtual machines. The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.

7. The system for dynamic allocation of cloud platform resources according to claim 6, characterized in that: Before obtaining the operation data of the access terminals of other virtual machines respectively, it further includes: Determining newly added virtual machines that are in an active state and available for the cloud platform to call, calculating the total resources of these newly added virtual machines. If the total resources can make the resource usage data of the first virtual machine not meet the second over-standard condition, no trigger signal is generated; otherwise, a trigger signal is generated; the trigger signal is used to trigger the obtaining of the operation data of the access terminals of other virtual machines respectively.

8. An electronic device, comprising: At least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein: the processor executes the computer program to implement the method according to any one of claims 1-5.

9. A computer storage medium, the computer storage medium stores a computer program, characterized in that: The computer program is executed by the processor to implement the method according to any one of claims 1-5.

10. A computer program product, characterized in that: The computer program product contains computer code, and when the computer code is executed by the processor of the electronic device, it implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Resource scheduling system of AI intelligent computing center

    CN117472587A

  • System and method for resource dynamic allocation and optimal scheduling in cloud computing environment

    CN118838709A

  • Cloud desktop virtual machine hardware resource allocation method and device and medium

    CN118885258A

  • System resource dynamic allocation method and optimization system in cloud network convergence environment

    CN119094612A

  • Cloud computing resource scheduling method and system based on big data

    CN119149231A