Computing power resource prediction method and device based on multi-scale transformer model, and storage medium
Through multi-scale sampling and attention aggregation of transformer models, the problem of insufficient accuracy in traditional computing power network prediction methods is solved, and higher accuracy resource prediction is achieved.
Patent Information
- Application Number
- CN202510504332.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-22
AI Technical Summary
Traditional computing power network resource prediction methods ignore the strong correlation between data on different time scales, resulting in poor prediction accuracy.
The computing resource prediction method based on the multi-scale transformer model is adopted. By sampling the historical state information sequence of the target resource under multiple time scales, the transformer model is used for attention aggregation and feature fusion, and a joint representation is generated to improve prediction accuracy.
Effectively capture the scale correlation characteristics of computing power network resource data, improving the accuracy and accuracy of resource prediction.
Smart Images

Figure CN120523587A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a computing resource prediction method, device, and storage medium based on a multi-scale transformer model. Background Art
[0002] In computing networks, demand for computing resources is often dynamic and uncertain. Different application scenarios have distinct demand patterns for computing resources. Even within the same scenario, resource requirements can vary significantly at different points in time. For example, in cloud computing environments, user request volume fluctuates significantly over time depending on business needs. In edge computing scenarios, due to the limited computing power and bandwidth constraints of devices, precise management and proper scheduling of computing resources are particularly important. To meet these practical needs, computing networks must be able to predict future resource demands in real time, thereby rationally allocating and scheduling computing resources to ensure system stability and efficiency.
[0003] Traditional computing power and network resource forecasting methods typically only consider the status data of different types of computing power resources at a single time scale, ignoring the strong correlations between data at different time scales. For example, due to the cyclical arrangement of workdays and weekends, forecast targets can be inferred not only from recent historical data but also from the computing power resource status at the same time in previous weeks for a more comprehensive forecast. These traditional computing power and network resource forecasting methods, ignoring the correlation and support of multi-scale data, result in poor forecast accuracy.
[0004] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a computing power resource prediction method, device and storage medium based on a multi-scale transformer model, aiming to solve the technical problem of how to improve the accuracy of computing power resource prediction.
[0006] To achieve the above objectives, this application proposes a computing resource prediction method based on a multi-scale transformer model, the method comprising: According to multiple time scales corresponding to the target prediction time point, sampling the historical state information sequence of the target resource at each time scale, and constructing multiple sampling sequences of the target resource at each time scale; The transformer model is used to aggregate the sampling sequences at the same time scale to obtain the first aggregation vector of the sampling sequence at each time scale. The first aggregation vectors of multiple sampling sequences at the same time scale are input into the feature fusion layer to obtain the second aggregation vector that fuses data at different time scales; The second aggregation vector is input into the output layer to obtain the state data of the target resource at the target prediction time point.
[0007] In one embodiment, the step of performing attention aggregation on the sampling sequences at the same time scale using the transformer model to obtain a first aggregation vector of the sampling sequence at each time scale includes: Determine the transformer computation mode corresponding to each time scale and construct a fusion module corresponding to each time scale, wherein each fusion module independently and in parallel performs attention aggregation; The sampling sequence at each time scale is input into the corresponding fusion module respectively, and the sampling sequence at the same time scale is subjected to attention aggregation by the fusion module to obtain a first aggregation vector of the sampling sequence at each time scale.
[0008] In one embodiment, the fusion module uses an interactive self-attention mechanism to perform attention aggregation.
[0009] In one embodiment, the fusion module performs attention aggregation using a multi-head parallel computing method through the interactive self-attention mechanism.
[0010] In one embodiment, before the step of inputting the sampling sequence at each time scale into the corresponding fusion module, the method further includes: According to the time sequence of the sampling sequence, a vector position code is created for each sampling sequence data.
[0011] In one embodiment, before the step of sampling the historical state information sequence of the target resource at each time scale according to the multiple time scales corresponding to the target prediction time point and constructing the multiple sampling sequences of the target resource at each time scale, the step further includes: For each computing power node in the computing power network, obtain historical status data of the target resource at multiple historical time points; The historical status data of the target resource at multiple historical time points are normalized to construct a historical status information sequence of the target resource.
[0012] In one embodiment, the step of sampling a historical state information sequence of a target resource at each time scale according to multiple time scales corresponding to a target prediction time point, and constructing multiple sampling sequences of the target resource at each time scale includes: Get the sampling interval corresponding to each time scale; According to the sampling interval corresponding to each time scale, the historical state information sequence of the target resource is sampled at different sampling starting positions at each time scale to construct multiple sampling sequences of the target resource at each time scale.
[0013] In one embodiment, the step of inputting the second aggregated vector into the output layer to obtain the state data of the target resource at the target prediction time point includes: determining a bias for the output layer; Determine the state data of the target resource at the target prediction time point according to the second aggregation vector, the preset output weight and the bias of the output layer.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a computing power resource prediction device based on a multi-scale transformer model, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the computing power resource prediction method based on the multi-scale transformer model as described above.
[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the computing power resource prediction method based on the multi-scale transformer model as described above are implemented.
[0016] The present application provides a computing power resource prediction method based on a multi-scale transformer model. According to multiple time scales corresponding to the target prediction time point, the historical state information sequence of the target resource is sampled at each time scale to construct multiple sampling sequences of the target resource at each time scale; the sampling sequences at the same time scale are respectively aggregated by the transformer model to obtain the first aggregate vector of the sampling sequence at each time scale; the first aggregate vectors of the multiple sampling sequences at the same time scale are input into the feature fusion layer to obtain the second aggregate vector that fuses data of different time scales; the second aggregate vector is input into the output layer to obtain the state data of the target resource at the target prediction time point.
[0017] This method combines data at different time scales, uses the transformer model to extract features at different time scales, and then integrates the features at different time scales through the feature fusion layer to generate a joint representation. This method can effectively capture the scale-related features of computing network resource data, better understand the feature relationships at different time scales, and improve prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A flowchart illustrating the first embodiment of a computing resource prediction method based on a multi-scale transformer model of this application; Figure 2 A flowchart illustrating the second embodiment of the computing power resource prediction method based on the multi-scale transformer model of this application; Figure 3 A schematic diagram of a prediction model for a computing power resource prediction method based on a multi-scale transformer model provided in Example 2 of the present application; Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the computing power resource prediction method based on the multi-scale transformer model in the embodiment of the present application.
[0021] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0022] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not intended to limit the present application.
[0023] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0024] In computing networks, demand for computing resources is often dynamic and uncertain. Different application scenarios have distinct demand patterns for computing resources. Even within the same scenario, resource requirements can vary significantly at different points in time. For example, in cloud computing environments, user request volume fluctuates significantly over time depending on business needs. In edge computing scenarios, due to the limited computing power and bandwidth constraints of devices, precise management and proper scheduling of computing resources are particularly important. To meet these practical needs, computing networks must be able to predict future resource demands in real time, thereby rationally allocating and scheduling computing resources to ensure system stability and efficiency.
[0025] Traditional computing power and network resource forecasting methods typically only consider the status data of different types of computing power resources at a single time scale, ignoring the strong correlations between data at different time scales. For example, due to the cyclical arrangement of workdays and weekends, forecast targets can be inferred not only from recent historical data but also from the computing power resource status at the same time in previous weeks for a more comprehensive forecast. These traditional computing power and network resource forecasting methods, ignoring the correlation and support of multi-scale data, result in poor forecast accuracy.
[0026] In view of the above problems, this application proposes a computing power resource prediction method based on a multi-scale transformer model. According to the multiple time scales corresponding to the target prediction time point, the historical state information sequence of the target resource is sampled at each time scale to construct multiple sampling sequences of the target resource at each time scale; the sampling sequences at the same time scale are respectively aggregated through the transformer model to obtain the first aggregation vector of the sampling sequence at each time scale; the first aggregation vectors of the multiple sampling sequences at the same time scale are input into the feature fusion layer to obtain the second aggregation vector that fuses data of different time scales; the second aggregation vector is input into the output layer to obtain the state data of the target resource at the target prediction time point.
[0027] This method combines data at different time scales, uses the transformer model to extract features at different time scales, and then integrates the features at different time scales through the feature fusion layer to generate a joint representation. This method can effectively capture the scale-related features of computing network resource data, better understand the feature relationships at different time scales, and improve prediction accuracy.
[0028] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, etc., or an electronic device capable of performing the above functions. The following uses the computing power resource prediction system as an example to illustrate this embodiment and the following embodiments.
[0029] Based on this, the first embodiment proposed in this application provides a computing resource prediction method based on a multi-scale transformer model. Figure 1 In this embodiment, the computing resource prediction method based on the multi-scale transformer model includes steps S10 to S40: Step S10 : sampling the historical state information sequence of the target resource at each time scale according to multiple time scales corresponding to the target prediction time point, and constructing multiple sampling sequences of the target resource at each time scale.
[0030] It should be noted that time scale refers to the division of time into different granularities, such as minutes, hours, days, and weeks. Capturing resource status changes at different time scales can provide a more comprehensive understanding of the dynamic nature of resource demand.
[0031] For example, the minute-scale data can capture short-term fluctuations. In real-world applications, computing resource demand can fluctuate significantly over short periods of time due to unexpected events, surges in user requests, and other factors. Therefore, minute-scale data can provide insights into computing resource fluctuations within a microscopic timeframe, enabling the model to rapidly respond and adjust. The hourly-scale data reflects the patterns of computing resource fluctuations within daily cycles. Many computing network applications are affected by daily work schedules, such as higher loads during office hours and lower loads at night. Therefore, hourly-scale data can help the model identify and capture these intraday fluctuations, providing accurate forecasts of medium-term demand. The weekly-scale data is used to describe overall trends in computing resource changes over longer periods. For example, computing resource demand typically differs significantly between weekdays and weekends, with higher loads on weekdays and lower loads on weekends. Setting a weekly scale allows the model to learn and capture long-term cyclical characteristics, thereby optimizing predictions of future computing demand.
[0032] For example, the target prediction time point and target resource are first determined. Next, multiple time scales are determined, such as minute, hour, and week scales. These scales are used to capture resource state changes at different granularities. The target resource's historical state information sequence X(t) is then retrieved from the system database, where X(t) represents the resource state data at time point t. For each time scale, samples are taken from the historical state information sequence X(t) at the minute, hour, and week scales, respectively, to construct multiple sampling sequences at the minute, hour, and week scales.
[0033] Optionally, step S10 includes steps S11-S12: Step S11: Obtain the sampling interval corresponding to each time scale.
[0034] Step S12: sampling the historical state information sequence of the target resource at each time scale at different sampling start positions according to the sampling interval corresponding to each time scale, and constructing multiple sampling sequences of the target resource at each time scale.
[0035] For example, assume that the sampling intervals corresponding to the minute scale, hour scale, and week scale are all 1. The system performs sampling at the minute scale, hour scale, and week scale according to the sampling interval corresponding to each time scale, and constructs multiple sampling sequences at the minute scale, hour scale, and week scale.
[0036] To better understand the solution provided in this example, this example is further explained in conjunction with specific application scenarios.
[0037] Assume that the target prediction time point is 17:34 on the 10th Sunday of a certain year. Sampling is performed at the minute scale, hour scale, and weekly scale, respectively, to construct multiple sampling sequences at the minute scale, hour scale, and weekly scale: Minute scale: computing power data at 5:33 PM on the 10th week, computing power data at 5:32 PM on the 10th week, computing power data at 5:31 PM on the 10th week, computing power data at 5:30 PM on the 10th week... Hourly scale: computing power data at 16:34 pm on the 10th week, computing power data at 15:34 pm on the 10th week, computing power data at 14:34 pm on the 10th week, computing power data at 13:34 pm on the 10th week... Weekly scale: computing power data at 3:34 PM on the 9th Sunday, computing power data at 3:34 PM on the 8th Sunday, computing power data at 3:34 PM on the 7th Sunday, computing power data at 3:34 PM on the 6th Sunday... After sampling according to the above sampling method, multiple sampling sequences of computing power data at each time scale can be obtained:
[0038] in, , , Sampling sequences at minute scale, hour scale and week scale respectively; [ , ,..., ],[ , ,..., ],[ , ,..., ] are the computing power data selected for sampling at minute, hour and week scales respectively; , , are the input window sizes corresponding to minute scale, hour scale, and week scale respectively.
[0039] In step S20, the sampling sequences at the same time scale are respectively subjected to attention aggregation by the transformer model to obtain the first aggregation vector of the sampling sequence at each time scale.
[0040] It should be noted that the transformer model is a deep learning model based on the attention mechanism, which is suitable for capturing long-term dependencies in sequences. It can extract and aggregate features of the input sequence and generate vectors that can represent the global features of the sequence.
[0041] Optionally, step S20 includes steps S21 and S22: Step S21, determine the transformer calculation mode corresponding to each time scale, and construct a fusion module corresponding to each time scale, wherein each fusion module independently and in parallel performs attention aggregation.
[0042] In step S22, the sampling sequence at each time scale is input into the corresponding fusion module respectively, and the fusion module performs attention aggregation on the sampling sequence at the same time scale to obtain a first aggregation vector of the sampling sequence at each time scale.
[0043] For example, for each time scale, the sampling sequence at that time scale is input into the transformer model. The transformer model consists of multiple parallel transformer fusion modules, each of which independently processes the sampling sequence at the corresponding time scale to obtain the first aggregate vector of the sampling sequence at each time scale:
[0044] in, , , Represent the transformer calculation modes corresponding to minute scale, hour scale and week scale respectively. , , are the first aggregation vectors of the sampling sequences at minute scale, hour scale, and week scale, respectively.
[0045] Optionally, before inputting the sampling sequences at each time scale into the corresponding fusion module, a vector position encoding is created for each sampling sequence data according to the time order of each sampling sequence, so that the transformer model can capture the sequential characteristics of the sequence when performing calculations.
[0046] For example, the sampling sequence corresponding to the minute scale is =[ , ,..., ] as an example, the sampling sequence Each element in , ,..., Represented as a vector , where i represents the position of the element in the sequence (starting from 0). Use the sine and / or cosine functions to generate the vector position encoding and compare it with the vector Add. The length of the position encoding vector is equal to the vector Specifically, assuming the vector The total length is D, that is, from 0 to D-1. For an element at position i and dimension d, if d is an even number, determine its vector position encoding .in, is a hyperparameter that controls the frequency of position encoding. Once the vector position encoding is generated , add the vector position code to the corresponding vector Specifically, for each position i and dimension d, the vector The dth dimension of will be updated as: .
[0047] Optionally, the above fusion module adopts an interactive self-attention mechanism to perform attention aggregation.
[0048] It's important to note that the self-attention mechanism generates new representations by calculating correlations between elements within a sequence. The interactive self-attention mechanism, on the other hand, builds on self-attention by adding the ability to account for interactions between multiple sequences. This means it dynamically allocates attention weights across multiple input sequences, capturing dependencies between them. The interactive self-attention mechanism calculates attention weights for each component. For example, for the self-attention component, it calculates attention weights for elements within the sequence; for the interactive component, it calculates attention weights between different sequences. It then takes a weighted sum of these attention weights across the sequence elements to generate a new sequence representation. The interactive self-attention mechanism not only focuses on information within a single sequence but also leverages information from other sequences to enhance the representation of the current sequence.
[0049] In this embodiment, an interactive self-attention mechanism is used to introduce interaction between multiple parallel fusion modules, allowing them to interact, share information, and jointly optimize attention weights. This mechanism allows different modules to not only focus on the characteristics of their own sequences when processing their respective input sequences, but also refer to the characteristics of other sequences, thereby enhancing the expressive power of their own sequences.
[0050] Exemplarily, each fusion module independently processes the sampling sequence at the corresponding time scale and generates its own first aggregation vector. In the attention calculation process of each module, the features of other modules are introduced as a reference. For example, for fusion module A that processes minute-scale data, when calculating the attention weight, it not only uses the input features of fusion module A, but also combines the features of fusion module B that processes hour-scale data and fusion module C that processes weekly-scale data. The interactive self-attention mechanism enables the transformer model to better understand the feature relationship of computing power data at different time scales, thereby improving the accuracy of prediction.
[0051] Optionally, the fusion module performs attention aggregation using a multi-head parallel computing method through an interactive self-attention mechanism.
[0052] For example, the transformer calculation mode of each fusion module can be expressed as follows:
[0053]
[0054]
[0055] MultiHead represents a multi-head attention mechanism, which is used to capture various relationships between different positions in the input sequence. Q, K, and V represent the query, key, and value matrices, respectively, which are the embedding representations of the input sequence. represents the output of the i-th attention head. , , are the projection weight matrices of the query, key, and value of the i-th attention head, respectively, used to linearly transform the input vector into different spaces. Concat means concatenating the outputs of multiple attention heads to form a comprehensive feature representation. is the multi-head fusion weight matrix, which is used to map the concatenated feature vector to the final output space. Attention represents the single-head attention mechanism. , , are the query, key, and value matrices of the h-th attention head, respectively. The Softmax function is used to convert the attention scores into a probability distribution. Compute the dot product between the query and the key to get the attention score matrix. The attention score is scaled, where is the dimension of the key, used to stabilize the gradient. is a matrix of values used to generate the final attention output.
[0056] It's understandable that by using multiple independent attention heads, each fusion module can simultaneously capture multiple relationships between different positions in the input sequence. For example, one head might focus on local features, while another might focus on global features. Each attention head can learn different feature patterns, and through splicing and fusion, the final output can more comprehensively represent the characteristics of the input sequence.
[0057] Step S30: input the first aggregation vectors of the sampling sequences at the same time scale into the feature fusion layer to obtain a second aggregation vector that fuses data at different time scales.
[0058] Exemplarily, the feature fusion layer can be expressed as follows:
[0059] in, is the second aggregation vector, is the bias of the feature fusion layer. , , Represent the fusion weights corresponding to minute scale, hour scale and weekly scale respectively. Represents the activation function, which usually uses a nonlinear function to introduce nonlinearity so that the model can learn more complex feature relationships.
[0060] Step S40: input the second aggregated vector into the output layer to obtain the state data of the target resource at the target prediction time point.
[0061] Optionally, step S40 includes steps S41-S42: Step S41, determining the bias of the output layer.
[0062] Step S42: determining the state data of the target resource at the target prediction time point according to the second aggregation vector, the preset output weight and the bias of the output layer.
[0063] For example, the output layer can be expressed as follows:
[0064] in, Indicates that the target resource is at the target prediction time point Status data, is the preset output weight corresponding to the output layer, is the bias of the output layer.
[0065] In this embodiment, by sampling the historical state information sequence of the target resource at multiple time scales and constructing the sampling sequence at each time scale, the dynamic nature of the computing power resource demand is fully captured. The multi-head attention mechanism of the transformer model is used to extract and aggregate the features of the sampling sequence at the same time scale to generate a first aggregate vector that can characterize the global features of the sequence. By introducing information interaction between multiple parallel fusion modules through the interactive self-attention mechanism, the feature expression capability is enhanced and the prediction accuracy is improved. The feature fusion layer integrates the first aggregate vectors of different time scales to generate a second aggregate vector of joint representation, further optimizing the feature expression. Finally, the fused feature vector is mapped to the state data of the target resource at the target prediction time point through the output layer, achieving high-precision computing power resource prediction.
[0066] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment can be referred to the above introduction and will not be described in detail later. Figure 2 Before step S10, the computing power resource prediction method based on the multi-scale transformer model further includes steps S60 to S70: Step S60: For each computing power node in the computing power network, obtain historical status data of the target resource at multiple historical time points.
[0067] Step S70 : normalizing the historical status data of the target resource at multiple historical time points to construct a historical status information sequence of the target resource.
[0068] A computing node is a computing unit in a computing network, such as a server, virtual machine, or container. The target resources corresponding to each computing node can be either computing resources or storage resources. Computing resources refer to the core hardware resources used by each computing node in the computing network to perform computing tasks, primarily including computing units such as central processing units (CPUs) and graphics processing units (GPUs). Storage resources refer to the hardware resources used to store and manage data in the computing network, including memory and disk storage. These storage resources are responsible for temporary and long-term data storage, as well as data read and write operations, directly impacting the system's computing efficiency and response speed.
[0069] For example, first identify the data source and retrieve historical status data for each node in the computing network from a database or data warehouse. Use query statements to obtain the target resource's status data at multiple historical points in time. This data is normalized to a fixed range, such as 0 to 1, to facilitate better data processing by the model. The normalized data is then arranged in chronological order to form a historical status information sequence.
[0070] For example, to help understand the implementation process of the computing power resource prediction method based on the multi-scale transformer model obtained by combining this embodiment with the above embodiments, please refer to Figure 3 , Figure 3 A schematic diagram of a prediction model for implementing the computing power resource prediction method based on a multi-scale transformer model in an embodiment of the present application is provided. Specifically, the prediction model includes an input layer, a transformer layer, a feature fusion layer, and an output layer, and the transformer layer includes multiple fusion modules. First, multiple sampling sequences of the target resource at various time scales are input into the input layer, and attention aggregation is performed in the transformer layer to obtain a first aggregation vector. The first aggregation vector is input into the feature fusion layer to obtain a second aggregation vector. The second aggregation vector is then input into the output layer to obtain the state data of the target resource at the target prediction time point.
[0071] Optionally, before performing the prediction, the historical status information sequence of each computing power node is obtained from the database or data warehouse, and the historical status information sequence of each computing power node is divided into a training data set and a test data set, where a portion of the training data set should be further divided as a verification data set for the prediction model.
[0072] The prediction model is trained on a training dataset using the mini-batch gradient descent method. For example, the following steps are performed: First, the prediction model parameters are randomly initialized and continuously updated during the training process. The entire training dataset is then divided into several mini-batches, each containing a portion of the data samples, such as 8, 16, 32, or 64 samples. The training process then iterates over these mini-batches. In each iteration, the prediction model extracts a mini-batch from the training dataset and uses it to calculate the error, or loss, of the current prediction model. Based on the loss of this mini-batch, the prediction model parameter adjustment direction, or gradient, is calculated. These gradients are then used to update the prediction model parameters, optimizing the prediction model performance towards reduced error. This iterative process is repeated until all mini-batches have been used, completing one round of training on the entire training dataset, known as an epoch. To help the prediction model better learn the data characteristics, this training process typically involves multiple rounds, or epochs, with all the data used in each round and the prediction model parameters continuously adjusted.
[0073] During training, the prediction model uses a validation dataset to monitor its performance to avoid overfitting and, if necessary, adjust hyperparameters such as the learning pace or learning rate. Training continues until the prediction model reaches a satisfactory level of performance or stops improving significantly. After this process is complete, the prediction model is validated using a test dataset to achieve prediction accuracy that meets the requirements of the contextual task. Metrics such as root mean square error (RMSE) can be used to evaluate the prediction model. Finally, the prediction model is deployed online to enable real-time prediction of computing power and network resources.
[0074] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the computing power resource prediction method based on the multi-scale transformer model of the present application. More forms of simple transformations based on this technical concept are all within the scope of protection of the present application.
[0075] The present application provides a computing power resource prediction device based on a multi-scale transformer model. The computing power resource prediction device based on the multi-scale transformer model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the computing power resource prediction method based on the multi-scale transformer model in the above-mentioned embodiment one.
[0076] Reference below Figure 4 , which shows a schematic diagram of the structure of a computing power resource prediction device based on a multi-scale transformer model suitable for implementing the embodiments of the present application. The computing power resource prediction device based on a multi-scale transformer model in the embodiments of the present application may include, but is not limited to, mobile terminals such as laptops and tablet computers (PADs) and fixed terminals such as desktop computers. Figure 4 The computing power resource prediction device based on the multi-scale transformer model shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0077] like Figure 4As shown, the computing power resource prediction device based on the multi-scale transformer model may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the computing power resource prediction device based on the multi-scale transformer model. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the computing power resource prediction device based on the multi-scale transformer model to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a computing power resource prediction device based on the multi-scale transformer model with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have instead.
[0078] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0079] The computing power resource prediction device based on the multi-scale transformer model provided in this application adopts the computing power resource prediction method based on the multi-scale transformer model in the above-mentioned embodiment, which can solve the technical problem of how to improve the accuracy of computing power resource prediction. Compared with the prior art, the beneficial effects of the computing power resource prediction device based on the multi-scale transformer model provided in this application are the same as the beneficial effects of the computing power resource prediction method based on the multi-scale transformer model provided in the above-mentioned embodiment, and the other technical features of the computing power resource prediction device based on the multi-scale transformer model are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0080] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0081] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0082] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the computing power resource prediction method based on the multi-scale transformer model in the above-mentioned embodiment.
[0083] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0084] The above-mentioned computer-readable storage medium can be included in the computing power resource prediction device based on the multi-scale transformer model; or it can exist independently without being assembled into the computing power resource prediction device based on the multi-scale transformer model.
[0085] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the computing power resource prediction device based on the multi-scale transformer model, the computing power resource prediction device based on the multi-scale transformer model can be written in one or more programming languages or a combination thereof to write computer program code for performing the operations of the present application. The programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computer, partially on the user computer, or as a separate software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect via the Internet).
[0086] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0087] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0088] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for predicting computing power resources based on a multi-scale transformer model. This computer-readable storage medium can address the technical problem of improving the accuracy of computing power resource predictions. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for predicting computing power resources based on a multi-scale transformer model provided in the aforementioned embodiments, and are not further elaborated here.
[0089] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A computing resource prediction method based on a multi-scale transformer model, characterized in that: The method comprises: According to multiple time scales corresponding to the target prediction time point, sampling the historical state information sequence of the target resource at each time scale, and constructing multiple sampling sequences of the target resource at each time scale; The transformer model is used to aggregate the sampling sequences at the same time scale to obtain the first aggregation vector of the sampling sequence at each time scale. The first aggregation vectors of multiple sampling sequences at the same time scale are input into the feature fusion layer to obtain the second aggregation vector that fuses data at different time scales; The second aggregation vector is input into the output layer to obtain the state data of the target resource at the target prediction time point.
2. The computing power resource prediction method based on the multi-scale transformer model according to claim 1, characterized in that: The step of performing attention aggregation on the sampling sequences at the same time scale by using the transformer model to obtain the first aggregation vector of the sampling sequence at each time scale includes: Determine the transformer computation mode corresponding to each time scale and construct a fusion module corresponding to each time scale, wherein each fusion module independently and in parallel performs attention aggregation; The sampling sequence at each time scale is input into the corresponding fusion module respectively, and the sampling sequence at the same time scale is subjected to attention aggregation by the fusion module to obtain a first aggregation vector of the sampling sequence at each time scale.
3. The computing power resource prediction method based on the multi-scale transformer model according to claim 2, characterized in that: The fusion module adopts an interactive self-attention mechanism to perform attention aggregation.
4. The computing power resource prediction method based on the multi-scale transformer model according to claim 3 is characterized in that: The fusion module performs attention aggregation using a multi-head parallel computing method through the interactive self-attention mechanism.
5. The computing power resource prediction method based on the multi-scale transformer model according to any one of claims 2 to 4, characterized in that: Before the step of inputting the sampling sequence at each time scale into the corresponding fusion module, the method further includes: According to the time sequence of the sampling sequence, a vector position code is created for each sampling sequence data.
6. The computing power resource prediction method based on the multi-scale transformer model according to claim 1, characterized in that: Before the step of sampling the historical state information sequence of the target resource at each time scale according to the multiple time scales corresponding to the target prediction time point and constructing the multiple sampling sequences of the target resource at each time scale, the method further includes: For each computing power node in the computing power network, obtain historical status data of the target resource at multiple historical time points; The historical status data of the target resource at multiple historical time points are normalized to construct a historical status information sequence of the target resource.
7. The computing power resource prediction method based on the multi-scale transformer model according to claim 1, characterized in that: The step of sampling the historical state information sequence of the target resource at each time scale according to the multiple time scales corresponding to the target prediction time point, and constructing the multiple sampling sequences of the target resource at each time scale includes: Get the sampling interval corresponding to each time scale; According to the sampling interval corresponding to each time scale, the historical state information sequence of the target resource is sampled at different sampling starting positions at each time scale to construct multiple sampling sequences of the target resource at each time scale.
8. The computing power resource prediction method based on the multi-scale transformer model according to claim 1, characterized in that: The step of inputting the second aggregated vector into the output layer to obtain the state data of the target resource at the target prediction time point includes: determining a bias for the output layer; Determine the state data of the target resource at the target prediction time point according to the second aggregation vector, the preset output weight and the bias of the output layer.
9. A computing resource prediction device based on a multi-scale transformer model, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the computing power resource prediction method based on a multi-scale transformer model as described in any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the computing power resource prediction method based on the multi-scale transformer model as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Resource demand prediction model training, demand prediction and resource scheduling method and system
CN116627630A
Multi-granularity sampling-based computing power network multi-dimensional resource joint prediction method and system
CN118555216A
Data prediction method, resource data prediction method and computing device
CN119848753A
Cited By
Water quality prediction method and system based on multi-scale feature fusion
CN120873994A