A method and system for dynamic allocation of cloud platform resources
By monitoring the virtual machine resource usage data in the cloud platform and filtering and coordinating virtual machines, forward-looking resource allocation is achieved, and the problem of lagging resource adjustment in the existing technology is solved, and the service quality and resource utilization of the cloud platform are improved.
Patent Information
- Application Number
- CN202510660868.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The dynamic allocation method of existing cloud platform resources relies on real-time monitoring and cannot predict future resource demand changes, resulting in resource adjustment lag during sudden business peaks, affecting terminal service quality.
By monitoring the resource usage data of the virtual machine, select coordinated virtual machines with timeliness indicators below the threshold, and when the resource usage data meets the second exceeding the standard, calculate resources will be allocated according to the set monitoring interval to achieve forward-looking resource allocation.
Reduce resource adjustment lag, improve resource utilization, enhance system adaptability, and ensure cloud platform service quality.
Smart Images

Figure CN120179424B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing, and in particular, to a method and system for dynamic allocation of cloud platform resources. Background Art
[0002] At present, with the rapid development of cloud computing technology, cloud platforms have become the core infrastructure for many enterprises and users to process data and run application programs. Currently, the architecture mode widely adopted by cloud platforms to provide data processing services for each terminal is as follows: the cloud platform divides physical servers into multiple virtual machines, each virtual machine manages at least one terminal, and the cloud platform can monitor each virtual machine in real time.
[0003] To ensure the performance of virtual machines, the cloud platform will monitor the resource usage of each virtual machine in real time. When it detects that the resource usage of a certain virtual machine exceeds the preset threshold, it will enable a resource scheduling algorithm to dynamically allocate other idle resources to this virtual machine. Common dynamic allocation algorithms, such as the greedy algorithm, genetic algorithm, simulated annealing algorithm, etc., are all based on the real-time monitored resource usage status for resource allocation.
[0004] However, the above-mentioned resource dynamic allocation method has significant drawbacks. Since it completely relies on real-time monitoring data and cannot predict future resource demand changes, it cannot pre-adjust resource allocation. This makes resource adjustment often lag behind when facing sudden business peaks or resource demand changes, resulting in poor resource adjustment effects. For example, in business scenarios such as e-commerce promotions and live streaming with goods, the number of terminal accesses will surge instantaneously. Due to the lack of pre-adjustment in the existing dynamic allocation method, resource supply is often not timely, resulting in a decline in the performance of virtual machines, which in turn affects the quality of terminal services, causing problems such as poor user experience and business losses.
[0005] Therefore, there is an urgent need for a cloud platform resource dynamic allocation method that can effectively solve the above problems. Summary of the Invention
[0006] In view of the above technical problems, the present invention provides a method, system, electronic device, computer storage medium, and computer program product for dynamic allocation of cloud platform resources.
[0007] The present invention discloses a method for dynamic allocation of cloud platform resources, and the method includes the following steps:
[0008] The cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine whose resource usage data meets the first over-standard condition appears, it obtains the operation data of the access terminals of other virtual machines respectively;
[0009] Based on the respective operating data corresponding to other virtual machines, a timeliness metric is evaluated to represent the expectation of the access terminal's response speed to the virtual machine, and several virtual machines with timeliness metrics lower than the metric threshold are selected as coordinated virtual machines;
[0010] The cloud platform continuously monitors the first virtual machine at a set monitoring interval, and when its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.
[0011] The present invention also discloses a system for dynamic allocation of cloud platform resources, the system includes a processing device and a storage device, and the computer code stored in the storage device is called and executed by the processing device to implement the following steps:
[0012] The cloud platform monitors the resource usage data of each virtual machine. When there is a first virtual machine whose resource usage data meets the first over-standard condition, it obtains the operating data of the access terminals of other virtual machines respectively;
[0013] Based on the respective operating data corresponding to other virtual machines, a timeliness metric is evaluated to represent the expectation of the access terminal's response speed to the virtual machine, and several virtual machines with timeliness metrics lower than the metric threshold are selected as coordinated virtual machines;
[0014] The cloud platform continuously monitors the first virtual machine at a set monitoring interval, and when its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.
[0015] The present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, and the processor executes the computer program to implement the method as described in any one of the previous items.
[0016] The present invention also discloses a computer storage medium, the computer storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in any one of the previous items.
[0017] The present invention also discloses a computer program product, the computer program product contains computer code, and when the computer code is executed by the processor of an electronic device, it implements the method as described in any one of the previous items.
[0018] The beneficial effects of the present invention are at least as follows:
[0019] The present invention first identifies a first virtual machine with a trend of resource tension in advance based on a first over-standard condition, screens out a coordinated virtual machine by evaluating the operation data of the access terminal, and then triggers the allocation of resources according to a second over-standard condition. The present invention realizes forward-looking resource allocation, reduces lag, improves resource utilization based on the needs of the access terminal, enhances system adaptability through multi-dimensional monitoring, and effectively guarantees the service quality of the cloud platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 is a schematic flowchart of a method for dynamically allocating cloud platform resources disclosed in an embodiment of the present invention;
[0022] Figure 2 is a schematic flowchart of using a combination algorithm to extract each operation feature from target operation data disclosed in an embodiment of the present invention;
[0023] Figure 3 is a schematic structural diagram of a system for dynamically allocating cloud platform resources disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.
[0025] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.
[0026] In the field of dynamic allocation of cloud platform resources, the prior art completely relies on real-time monitoring data and cannot predict future resource demand changes, resulting in lag in resource adjustment when facing sudden business peaks. For example, during a major e-commerce promotion, the CPU usage rate of a virtual machine suddenly soars from 30% to 80%. The traditional method can only start resource allocation when the usage rate exceeds a preset threshold (such as 85%). During this period, slow service response may occur due to insufficient resources, affecting the user experience.
[0027] To solve this problem, the present invention proposes a new method for dynamic resource allocation, which is based on the existing architecture model in which physical servers are divided into multiple virtual machines on a cloud platform, and each virtual machine manages at least one terminal and monitors the virtual machines in real time (especially the activation status).
[0028] As Figure 1 shown, an embodiment of the present invention discloses a method for dynamic resource allocation on a cloud platform, and the method includes the following steps:
[0029] S101, the cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine whose resource usage data meets the first over-standard condition appears, obtain the running data of the access terminals of other virtual machines respectively.
[0030] The cloud platform continuously monitors the resource usage data of each virtual machine (such as CPU usage rate, memory occupancy, disk I / O read and write rate, network bandwidth utilization rate, etc.). When the resource usage data of a certain virtual machine meets the first over-standard condition, mark it as the first virtual machine and trigger further operations. For example:
[0031] (1) In an online education cloud platform, virtual machine A responsible for real-time course live broadcast undertakes the live streaming and interactive data processing tasks of multiple classes. During the prime learning period in the evening, a large number of students watch the courses online at the same time and ask questions in real time, resulting in a continuous increase in the CPU usage rate of virtual machine A. When the CPU usage rate exceeds 70% set in the first over-standard condition for 10 consecutive minutes, for example, reaches 75%, at this time, the resource usage data of virtual machine A meets the first over-standard condition. Excessive CPU usage rate will cause problems such as stuttering of the live broadcast screen and delay of interactive messages, affecting the learning experience of students, and further resource allocation is required.
[0032] (2) During the big promotion activity of an e-commerce platform, virtual machine B used for order processing and inventory management has its memory occupancy continuously climbing due to a large number of order placement requests and inventory query operations pouring in within a short period of time. Virtual machine B is initially allocated 8GB of memory. When the memory occupancy exceeds 6.4GB (i.e., 80% of the memory occupancy ratio) set in the first over-standard condition for 15 consecutive minutes and reaches 6.6GB, it indicates that the memory resources of this virtual machine are in short supply. Insufficient memory may lead to a slowdown in order processing speed, and even serious problems such as order loss and inventory data disorder. Therefore, after meeting the first over-standard condition, the subsequent resource allocation process needs to be started.
[0033] (3) In the content storage and distribution system of the short video platform, virtual machine C is mainly responsible for storing the videos uploaded by users and distributing popular videos. When a certain popular video goes viral, a large number of users request to play the video simultaneously, causing the disk I / O read rate of virtual machine C to increase sharply. When the disk I / O read rate exceeds 400 MB / s set in the first over-standard condition for 20 consecutive minutes and reaches 450 MB / s, it indicates that the disk read and write pressure is too high. The excessive disk I / O rate will reduce the video loading speed, resulting in frequent buffering when users watch videos. At this time, the resource usage data of virtual machine C meets the first over-standard condition, and resource adjustment is required.
[0034] (4) In the game cloud platform, virtual machine D provides game data interaction services for numerous online players. When a new game version is launched, a large number of players log in to the game simultaneously for updates and play, causing the network bandwidth utilization rate of virtual machine D to rise rapidly. When the network bandwidth utilization rate exceeds 80% set in the first over-standard condition for 30 consecutive minutes and reaches 85%, it indicates that the network transmission capacity is approaching saturation. Tight network bandwidth will cause excessive game latency, untimely player operation responses, and even disconnections. Therefore, virtual machine D meets the first over-standard condition, and the cloud platform needs to obtain the operation data of other virtual machine terminals and prepare for dynamic resource allocation.
[0035] After the resource usage data of the first virtual machine meets the first over-standard condition, the cloud platform obtains the operation data of the access terminals of other virtual machines respectively to provide a more comprehensive decision-making basis for subsequent resource scheduling. By obtaining the operation data of the access terminals, the cloud platform can more accurately evaluate the resource allocation potential of each virtual machine, so as to screen out appropriate coordinating virtual machines in the subsequent steps, avoid blindly allocating resources and affecting the normal operation of other services, and achieve more reasonable resource allocation.
[0036] S102, based on the respective operation data corresponding to other virtual machines, evaluate the timeliness index used to represent the expectation of the response speed of the access terminal to the virtual machine, and screen out several virtual machines with timeliness indexes lower than the index threshold as coordinating virtual machines.
[0037] Based on the operation data corresponding to each other virtual machine, evaluate the timeliness index. Among them, the timeliness index represents the speed at which each access terminal requires the virtual machine to respond to its data processing and data acquisition requests (which can be represented by the response duration), and this requirement for the response speed is obtained based on the real-time operation data analysis of each terminal. That is, in different operating states, the timeliness index is actually different and needs to be analyzed in real time.
[0038] Under the existing architecture, the mode of each virtual machine management terminal makes the terminal operation data closely associated with the virtual machine. The cloud platform can accurately identify which virtual machines can be used to allocate some resources for the first virtual machine without affecting the quality of its terminal services by screening and coordinating virtual machines based on the computing timeliness index, so as to ensure that the first virtual machine can provide computing services for each access terminal connected to it. Among them, the screening of coordinated virtual machines should follow the principle of sufficient but not excessive, and should not be overly screened to reduce the impact on the normal business operation of other virtual machines.
[0039] S103. The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; among them, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.
[0040] In the foregoing step S102, the cloud platform has screened out each coordinated virtual machine that can be used to provide resource allocation for the first virtual machine, but the first over-standard condition is actually set slightly lower to achieve earlier screening of coordinated virtual machines for the first virtual machine. After that, the cloud platform continues to monitor the first virtual machine. When its resource usage data meets a higher second over-standard condition (that is, more serious over-standard), it controls each coordinated virtual machine to allocate computing resources to the first virtual machine. Among them, the first over-standard condition and the second over-standard condition can be comprehensively determined based on the historical operation data and management requirements of the cloud platform, and will not be elaborated here.
[0041] The cloud platform dynamically adjusts the allocation of virtual machines according to the actual needs of the terminal access type and quantity. For example, the cloud platform divides out a sufficient number of virtual machines at one time, but only activates a part of them based on the actual terminal access type and quantity; as the number of connected terminals increases, the cloud platform will gradually activate some dormant virtual machines, or re-divide and activate some new virtual machines to meet the actual needs. And these newly added (that is, newly activated) virtual machines will make the cloud platform have a greater redundancy when coordinating resources for the first virtual machine. The present invention sets the monitoring interval for adjusting whether or when the first virtual machine reaches the second over-standard condition in this stage of the first over-standard condition and the second over-standard condition according to this redundancy. It can be understood that the monitoring interval is positively correlated with the number of newly added virtual machines.
[0042] The present invention first identifies the first virtual machine with a resource shortage trend in advance with the first over-standard condition, screens out the coordinated virtual machines by evaluating the operation data of the access terminals, and then triggers the allocation of resources according to the second over-standard condition. The present invention realizes forward-looking resource allocation, reduces lag, improves resource utilization rate based on the needs of access terminals, enhances system adaptability through multi-dimensional monitoring, and effectively guarantees the service quality of the cloud platform.
[0043] Optionally, before obtaining the running data of the access terminals of the other virtual machines, the method further includes:
[0044] Determine the newly added virtual machines that are in an active state and available for invocation on the cloud platform, calculate the total resources of these newly added virtual machines, and if the total resources can make the resource usage data of the first virtual machine not meet the second over-standard condition, no trigger signal is generated; otherwise, a trigger signal is generated. The trigger signal is used to trigger the obtaining of the running data of the access terminals of the other virtual machines respectively.
[0045] In this embodiment, during the operation of the cloud platform, the virtual machine allocation is dynamically adjusted according to the terminal access type and quantity. For example, dormant virtual machines are gradually activated or new virtual machines are partitioned, and the active virtual machines without terminal access for the time being also belong to the newly added virtual machines. At this time, the cloud platform first determines the newly added virtual machines that are currently available for invocation and in an active state. For example, before a major e-commerce promotion, the cloud platform pre-activates 10 virtual machines for coping with traffic peaks, and these 10 are the newly added virtual machines. Then, the cloud platform calculates the total resources of these newly added virtual machines (partitioned by the cloud platform when they are generated), including resource data such as the total number of CPU cores, the total memory capacity, the available disk space, and the total network bandwidth. Suppose these 10 newly added virtual machines have a total of 20 CPU cores, 80 GB of memory, 1 TB of disk space, and 10 Gbps of network bandwidth.
[0046] Associate and evaluate the total resources of the newly added virtual machines obtained by calculation with the resource requirements of the first virtual machine (the difference between the predicted resource demand of the first virtual machine and the total resources allocated to the first virtual machine). The basis for the evaluation is to determine whether these newly added resources can make the resource usage data of the first virtual machine not meet the second over-standard condition. If the total resources of the newly added virtual machines are sufficient and can make the resource usage data of the first virtual machine not reach the standard of the second over-standard condition within a foreseeable time range (such as 10 minutes) through resource supplementation, no trigger signal is generated. At this time, there is no need to further obtain the running data of other virtual machine terminals and start the subsequent complex resource allocation process.
[0047] For example, if the second over-standard condition is set that the CPU usage rate exceeds 85%, and according to the calculation, after the 20 CPU cores of the newly added virtual machines are allocated to the first virtual machine, its CPU usage rate can be controlled below 80%, that is, it does not meet the second over-standard condition, and no trigger signal is generated at this time.
[0048] On the contrary, if the total resources of the newly added virtual machines are not sufficient to prevent the first virtual machine from reaching the second over-standard condition, a trigger signal is generated, and subsequent operations such as obtaining the running data of the access terminals of the other virtual machines respectively will be executed to further screen and coordinate the virtual machines and allocate resources to ensure the performance of the first virtual machine.
[0049] Optionally, the timeliness index used to characterize the expected response speed of the terminal to the virtual machine, which is evaluated based on the respective operation data corresponding to other virtual machines, includes:
[0050] Target operation data is intercepted from the operation data of each access terminal corresponding to other virtual machines according to the acquisition duration, and several data types are matched according to the pre-configured tags of the other virtual machines; wherein, the acquisition duration is determined based on the terminal type of the access terminal;
[0051] The operation characteristics corresponding to each data type are extracted from the target operation data using a combination algorithm, and each operation characteristic is spliced into a target operation characteristic;
[0052] The GNN-Transformer evaluator is used to classify and evaluate the target operation characteristics to obtain the timeliness index used to characterize the expected response speed of the terminal to the virtual machine.
[0053] In this embodiment, the cloud platform or the virtual machine continuously records the data interaction records between the virtual machine and each access terminal to form operation data. However, there is a lot of early data in these operation data, which is not helpful for analyzing the current stage's expectation of the response speed of the access terminal. Therefore, it is necessary to determine an appropriate acquisition duration, and intercept a section from the operation data according to this acquisition duration (the interception end point is the current moment, and the interception start point is the historical moment), that is, the target operation data.
[0054] Among them, the acquisition duration is related to the terminal type of the access terminal. For example, for a Web terminal with high-frequency real-time interaction (such as an e-commerce real-time transaction terminal), due to its business characteristics, it has high requirements for instant response, and early data (such as request delay records more than 5 minutes ago) is difficult to reflect the current user experience requirements. Therefore, a 5-minute acquisition duration is set, and only data such as request interval sequences and timeout retry times within the last 5 minutes are intercepted to focus on real-time interaction characteristics. For a low-frequency batch-processing IoT terminal (such as a sensor data upload terminal), its data interaction is periodic and has a relatively high tolerance for response delay. Early data (such as transmission time-consuming records within 30 minutes) can still reflect the current stage's periodic law. Therefore, a 30-minute acquisition duration is set, and data such as batch data transmission success rate and processing time-consuming within this period are intercepted to retain trend characteristics.
[0055] By binding the acquisition duration to the terminal type, the cloud platform can accurately filter out invalid historical data, ensuring that the target operation data used to evaluate the timeliness metric only contains valid information highly relevant to the current terminal's needs, and avoiding evaluation biases caused by the inclusion of outdated data. It should be understood that the terminal type is not limited to the above types, and there are also multiple corresponding acquisition durations. A terminal type - acquisition duration comparison table can be established in advance, which will not be elaborated here.
[0056] Next, after filtering out the target operation data by the acquisition duration, first use a combination algorithm to extract the operation characteristics of the sub - data corresponding to the specified data type from each target operation data. The data types are shown in Table 1 below:
[0057]
[0058] Among them, when the cloud platform generates a virtual machine, it can configure one or more pre - configured tags representing different interaction attributes for it. These tags include business real - time tags, data interaction mode tags, user - experience - sensitive tags, resource allocation strategy tags, etc. (each tag corresponds to multiple sub - tags, such as real - time business tags and non - real - time business tags). Based on the association relationship between these tags and data types, multiple data types corresponding to specific other virtual machines can be determined for which operation characteristics need to be extracted. Specifically as follows:
[0059] 1. Business real - time tags
[0060] Businesses with strong real - time requirements (such as online live streaming, financial transactions): Need to focus on millisecond - level response speed, so key attention is paid to data directly reflecting immediate interaction such as request interval sequences, timeout retry counts, and real - time business tags.
[0061] Non - real - time businesses (such as batch data processing, log archiving): Allow a certain delay, and pay more attention to trend - based metrics (such as batch transmission success rate, processing time), and do not require high - frequency real - time data.
[0062] 2. Data interaction mode tags
[0063] High - frequency burst interaction businesses (such as e - commerce flash sales, game servers): Need to analyze the burst request frequency (QPS peak) and user operation sequences to identify resource pre - emption risks and optimize response priorities.
[0064] Low - frequency periodic businesses (such as IoT sensor data upload): Focus on periodic transmission time and historical feedback data for evaluating long - term stability rather than instantaneous response.
[0065] 3. User - experience - sensitive tags
[0066] Interactive-intensive services (such as online collaboration tools): Users are sensitive to the intervals between consecutive operations (such as editing delays), and it is necessary to monitor the user operation sequence and terminal device attributes (such as network mode) to avoid degradation of the experience caused by device performance or network fluctuations.
[0067] Non-interactive services (such as background data synchronization): Pay more attention to business type tags (non-real-time) and historical feedback data (such as the task failure rate in implicit feedback) to ensure task integrity rather than real-time response.
[0068] 4. Resource Allocation Policy Tags
[0069] High-priority services (such as medical image analysis): It is necessary to judge the impact of response delay on the service through the number of timeout retries and historical feedback data, and preferentially allocate computing resources to ensure timeliness.
[0070] Low-priority services (such as advertising push): Higher delays can be tolerated, and key monitoring is focused on cost-related metrics (such as resource utilization in the request interval sequence) to balance response speed and resource consumption.
[0071] After extracting the running characteristics of each access terminal, these running characteristics are spliced into target running characteristics by means such as weighted splicing and sliding window splicing. Then, a GNN-Transformer evaluator is used to classify and evaluate the target running characteristics to obtain a timeliness index for characterizing the expected response speed of the terminal to the virtual machine.
[0072] Optionally, if the combined algorithm is wavelet transform + CNN + Transformer, the combined algorithm is used to extract the running characteristics corresponding to each data type from the target running data, including:
[0073] Use wavelet transform to perform multi-scale decomposition on the sub-data corresponding to the data type in the target running data, extract the high-frequency component and the low-frequency component, and combine the components after normalization to form preliminary time series characteristics;
[0074] Adopt dynamic adjustment of the convolution kernel number of the one-dimensional convolutional layer to capture the patterns of different time windows, use the one-dimensional convolutional layer of CNN to extract local features from the time series characteristics, and then reduce the dimension through the pooling layer to retain key features to obtain target time series characteristics;
[0075] Use the Transformer encoder to calculate the correlation weights between the feature dimensions in the target time series characteristics by using the multi-head self-attention mechanism, and generate a context-aware feature vector that fuses time series and semantic information;
[0076] Splice the target time series characteristics and the context-aware feature vector into running characteristics according to the data type.
[0077] In this embodiment, since the target operation characteristics obtained in the foregoing steps include multi-type and multi-granularity data, such as time series metric characteristics (request interval sequence, response time fluctuation, number of timeout retries, etc.), device attribute characteristics (terminal type (mobile phone / PC), network mode (5G / Wi-Fi), etc.), and some data also has characteristics such as long-distance dependence. Traditional single algorithms are difficult to capture high-frequency burst characteristics (such as the instantaneous QPS peak in the flash sale scenario), low-frequency trend characteristics (such as the periodic data transmission pattern of IoT terminals), and global correlation relationships (such as the impact of the time difference between "add to cart - settlement" in the user operation sequence on resource requirements) at the same time. Thus, as Figure 2 shown, the combined algorithm composed of wavelet transform + CNN + Transformer in the present invention extracts features of different granularities layer by layer from the target operation characteristics through multi-stage processing, and then obtains operation characteristics that can better reflect the requirements of response timeliness. Specifically:
[0078] First, the sub-data corresponding to each data type in the target operation data is decomposed into high-frequency components (reflecting short-term fluctuations and sudden changes) and low-frequency components (reflecting long-term trends and periodic patterns) through wavelet transform. For example: in the request interval sequence during a major e-commerce promotion, the high-frequency component corresponds to the intensive requests at the moment of the flash sale, and the low-frequency component corresponds to the overall traffic trend during the promotion period. Normalize each component (for example, using Z-score standardization) to eliminate the dimension difference, and then combine them into preliminary time series features containing multi-scale information.
[0079] After wavelet transform processing, signals of different time scales can be separated, avoiding high-frequency noise masking low-frequency trends or low-frequency trends blurring high-frequency details, and providing a cleaner multi-granularity time series input for subsequent CNN and Transformer.
[0080] Then, dynamically adjust the number of convolution kernels in the one-dimensional convolutional layer of CNN, and scan the preliminary time series features through sliding windows of different sizes (such as kernel size = 3, 5, 7) to extract local features, such as continuous high-frequency requests and periodic retry peaks. Use max pooling or average pooling to reduce the feature dimension and retain key features, such as the maximum QPS value and average response delay within a certain time window. Take the key features as the target time series features.
[0081] Next, use a Transformer encoder to perform global dependency modeling on the target time-series features. Utilize the multi-head self-attention mechanism to calculate the correlation weights of each time point and each dimension (such as request interval, number of retries, business type label) in the feature vector, and capture long-distance dependencies (such as the resource requirement correlation between a user's "add to cart" operation and the "checkout" operation 10 minutes later). At the same time, inject position information (such as timestamp) into the time-series features, that is, generate a context-aware feature vector that fuses time-series and semantic information. This context-aware feature vector integrates global time-series dependencies and semantic associations (such as "there is a strong positive correlation between high-frequency request periods and continuous user operations, and resources need to be guaranteed first").
[0082] Finally, concatenate the target time-series features (local features) extracted by the CNN and the context-aware feature vector (global dependency features) generated by the Transformer according to the data type (such as request interval sequence, number of timeout retries) to obtain the running features. After obtaining the running features corresponding to all access terminals, concatenate them end to end to obtain the target running features.
[0083] Optionally, the method of dynamically adjusting the number of convolutional kernels of the one-dimensional convolutional layer to capture patterns in different time windows includes:
[0084] Determine the adjustment span according to the time scale type and fluctuation frequency type of the sub-data, and dynamically adjust the number of convolutional kernels of the one-dimensional convolutional layer according to the adjustment span to achieve capturing patterns in different time windows; wherein, the time scale type includes long-period fluctuations, medium-period fluctuations, short-period fluctuations, and the frequency type includes high frequency, medium frequency, and low frequency.
[0085] In this embodiment, the time scale characteristics and fluctuation frequency characteristics of the sub-data of different data types are different. Refer to the example shown in Table 2 below:
[0086] Table 2
[0087]
[0088] The present invention sets the adjustment span of the number of convolutional kernels (such as the step size ΔK when adjusting from the initial value K0 to K1) to match the time scale type and fluctuation frequency type of the data type. It can be understood that the above Table 2 is only an example for some data types and is not used to limit the data types to only those shown in Table 2.
[0089] For example, for high-frequency short-period data, small-span dynamic adjustment (e.g., ΔK = 2) is adopted to ensure fine-grained changes in the number of convolutional kernels, covering more small-window patterns. For low-frequency long-period data (such as device attributes), large-span dynamic adjustment (e.g., ΔK = 3) is used to reduce redundant calculations and focus on key long-window patterns. For medium-frequency medium-period data (such as business behaviors), medium-span dynamic adjustment (e.g., ΔK = 4) is adopted to balance computational efficiency and feature coverage.
[0090] It should be noted that to avoid feature omission caused by too large a span setting (for example, for low-frequency data, if the span is K0×50%, key windows may be skipped), batch normalization can be combined to mitigate the distribution shift problem caused by the adjustment of the number of kernels.
[0091] As Figure 3 shown, an embodiment of the present invention also discloses a system 100 for dynamic allocation of cloud platform resources. The system includes a processing device 200 and a storage device 300. The computer code stored in the storage device 300 is called and executed by the processing device 200 to implement the following steps:
[0092] The cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine whose resource usage data meets the first over-standard condition appears, the running data of the access terminals of other virtual machines is obtained;
[0093] Based on the respective running data corresponding to other virtual machines, a timeliness index used to characterize the expected response speed of the access terminal to the virtual machine is evaluated, and several virtual machines with timeliness indexes lower than the index threshold are selected as coordinated virtual machines;
[0094] The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition.
[0095] An embodiment of the present invention also discloses an electronic device, including: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. The processor executes the computer program to implement the method as described in the foregoing embodiments.
[0096] An embodiment of the present invention also discloses a computer storage medium. The computer storage medium stores a computer program, and the computer program is executed by a processor to implement the method as described in the foregoing embodiments.
[0097] An embodiment of the present invention also discloses a computer program product, which contains computer code that, when executed by a processor of an electronic device, implements the method described in the foregoing embodiments.
[0098] The above-mentioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0099] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitations are imposed herein.
[0100] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for dynamic allocation of cloud platform resources, characterized in that: The method includes the following steps: The cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine with resource usage data meeting the first over-standard condition appears, the running data of the access terminals of other virtual machines is obtained; Based on the running data corresponding to other virtual machines, a timeliness index for characterizing the expected response speed of the access terminal to the virtual machine is evaluated, and several virtual machines with timeliness indexes lower than the index threshold are selected as coordinated virtual machines; the timeliness index characterizes the speed at which each access terminal requires the virtual machine to respond to its data processing and data acquisition requests; The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition; Evaluating a timeliness index for characterizing the expected response speed of the terminal to the virtual machine based on the running data corresponding to other virtual machines includes: Target running data is intercepted from the running data of the access terminals corresponding to other virtual machines according to the acquisition duration, and several data types are obtained by matching the pre-configured tags of the other virtual machines; wherein, the acquisition duration is determined based on the terminal type of the access terminal; The combination algorithm is used to extract the running characteristics corresponding to each data type from the target running data, and each running characteristic is spliced into a target running characteristic; The GNN-Transformer evaluator is used to classify and evaluate the target running characteristics to obtain a timeliness index for characterizing the expected response speed of the terminal to the virtual machine.
2. The method for dynamically allocating cloud platform resources according to claim 1, wherein: Before obtaining the running data of the access terminals of other virtual machines, the method further includes: Determine the newly added virtual machines in the cloud platform that are in an active state and available for invocation, calculate the total resources of these newly added virtual machines. If the total resources can make the resource usage data of the first virtual machine not meet the second over-standard condition, no trigger signal is generated, otherwise a trigger signal is generated; the trigger signal is used to trigger the acquisition of the running data of the access terminals of other virtual machines.
3. A method for dynamically allocating cloud platform resources according to claim 1, characterized in that: If the combination algorithm is wavelet transform + CNN + Transformer, then using the combination algorithm to extract the running characteristics corresponding to each data type from the target running data includes: Use wavelet transform to perform multi-scale decomposition on the sub-data corresponding to the data type in the target running data, extract the high-frequency component and the low-frequency component, and combine each component after normalization to form a preliminary time series feature; Adopt a dynamic adjustment of the convolution kernel number of the one-dimensional convolutional layer to capture the patterns of different time windows, use the one-dimensional convolutional layer of CNN to extract local features from the time series feature, and then reduce the dimension through the pooling layer to retain the key features to obtain the target time series feature; The Transformer encoder uses the multi-head self-attention mechanism to calculate the correlation weights between the feature dimensions and generate a context-aware feature vector that fuses time series and semantic information. The target timing feature and the context-aware feature vector are concatenated into a running feature according to the data type.
4. A method for dynamically allocating cloud platform resources according to claim 3, characterized in that: The method of dynamically adjusting the number of convolutional kernels of the one-dimensional convolutional layer to capture patterns in different time windows includes: Determining an adjustment span according to the time scale type and fluctuation frequency type of the sub-data, and dynamically adjusting the number of convolutional kernels of the one-dimensional convolutional layer according to the adjustment span to capture patterns in different time windows; wherein, the time scale type includes long-period fluctuations, medium-period fluctuations, and short-period fluctuations, and the frequency type includes high frequency, medium frequency, and low frequency.
5. A system for dynamic allocation of cloud platform resources, the system comprising a processing device and a storage device, characterized in that: The computer code stored in the storage device is called and executed by the processing device to implement the following steps: The cloud platform monitors the resource usage data of each virtual machine. When a first virtual machine whose resource usage data meets the first over-standard condition appears, the running data of the access terminals of the other virtual machines is obtained. Based on the running data corresponding to the other virtual machines, a timeliness index for characterizing the expected response speed of the access terminal to the virtual machine is evaluated, and several virtual machines with timeliness indexes lower than the index threshold are selected as coordinated virtual machines; the timeliness index characterizes the speed at which each access terminal requires the virtual machine to respond to its data processing and data acquisition requests. The cloud platform continuously monitors the first virtual machine at a set monitoring interval. When its resource usage data meets the second over-standard condition, it controls each coordinated virtual machine to allocate computing resources to the first virtual machine; wherein, the monitoring interval is determined based on the number of newly added virtual machines, and the second over-standard condition is higher than the first over-standard condition. Evaluating a timeliness index for characterizing the expected response speed of the terminal to the virtual machine based on the running data corresponding to the other virtual machines, including: Intercepting target running data from the running data of the access terminals of the other virtual machines according to the acquisition duration, and matching several data types according to the pre-configured tags of the other virtual machines; wherein, the acquisition duration is determined based on the terminal type of the access terminal. Using a combination algorithm to extract running features corresponding to each data type from the target running data, and concatenating the running features into a target running feature. Using a GNN-Transformer evaluator to classify and evaluate the target running feature to obtain a timeliness index for characterizing the expected response speed of the terminal to the virtual machine.
6. The system for dynamic allocation of cloud platform resources according to claim 5, wherein: Before obtaining the running data of the access terminals of the other virtual machines, it further includes: Determining newly added virtual machines that are in an active state and available for the cloud platform to call, calculating the total resources of these newly added virtual machines. If the total resources can make the resource usage data of the first virtual machine not meet the second over-standard condition, no trigger signal is generated, otherwise a trigger signal is generated; the trigger signal is used to trigger the acquisition of the running data of the access terminals of the other virtual machines.
7. An electronic device, comprising: At least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, characterized in that: the processor executes the computer program to implement the method according to any one of claims 1-4.
8. A computer storage medium storing a computer program, characterized in that: The computer program is executed by a processor to implement the method according to any one of claims 1-4.
9. A computer program product, characterized in that: The computer program product contains computer code which, when executed by a processor of an electronic device, implements the method according to any one of claims 1-4.
Citation Information
Patent Citations
System and method for resource dynamic allocation and optimal scheduling in cloud computing environment
CN118838709A
Typical adjustable resource optimization method and system for power grid adjustment scene
CN119417136A