Model training method, device and program product

By processing and training the resource usage information of the cloud platform and generating a target prediction model, the problem that cloud platform is difficult to accurately predict user needs in resource allocation is solved, and the prediction accuracy and resource utilization efficiency of resource usage are improved.

CN120050245APending Publication Date: 2025-05-27GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411999740.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When cloud platforms allocate resources, it is difficult for them to accurately predict the user's resource needs, resulting in extended response time or service interruption, or waste of resources.

Method used

By obtaining resource usage information of the cloud platform, processing and generating target fragment information, including resource usage characteristics and timestamp characteristics, it is used to train pre-trained models and generate target prediction models to predict future resource usage.

Benefits of technology

It improves the accuracy of predicting cloud platform resource usage, helps cloud platform allocate resources more accurately, and avoids resource waste and service interruptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050245A_ABST
    Figure CN120050245A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, equipment and a program product, and belongs to the technical field of computers. The method comprises: obtaining first resource usage information of a cloud platform, the first resource usage information comprising a plurality of resource usage and a plurality of timestamps, the plurality of resource usage being in one-to-one correspondence with the plurality of timestamps; the first resource usage information is processed, n pieces of target fragment information are obtained, and each piece of target fragment information in the n pieces of target fragment information comprises a resource usage feature and a timestamp feature; and training a pre-training model according to the n pieces of target fragment information to obtain a target prediction model. According to the method, the pre-training model is trained through the timestamp feature and the resource usage feature of the cloud platform, the target prediction model obtained through training can better understand the dependency between time and resource usage in the cloud platform, and therefore the prediction accuracy of the target prediction model for the resource usage of the cloud platform can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a model training method, device, and program product. Background Art

[0002] Cloud platforms have been widely adopted because they can flexibly provide computing services, storage services, etc. for users. When a user uses the services of a cloud platform, the cloud platform will allocate resources for the user to use. However, if the allocated resources are less than the resources actually required by the user, it may lead to an extended or even interrupted service response time; if the allocated resources are greater than the resources actually required by the user, it will meaninglessly increase the burden on the cloud platform. Therefore, how to accurately allocate resources for users is an issue that needs to be focused on currently. Summary of the Invention

[0003] This application provides a model training method, device, and program product, which can improve the prediction accuracy of the trained model. The technical solutions are as follows:

[0004] In a first aspect, a model training method is provided. The method includes:

[0005] Obtain first resource usage information of a cloud platform. The first resource usage information includes multiple resource usages and multiple timestamps. The multiple resource usages correspond one-to-one to the multiple timestamps. Each resource usage in the multiple resource usages is the resource usage of the cloud platform at the moment indicated by the corresponding timestamp. Each moment indicated by the multiple timestamps is within a first time period;

[0006] Process the first resource usage information to obtain n target segment information. Each target segment information in the n target segment information includes a resource usage feature and a timestamp feature. n is an integer greater than or equal to 2;

[0007] Train a pre-trained model according to the n target segment information to obtain a target prediction model, where the target prediction model is used to predict the resource usage of the cloud platform in a second time period.

[0008] In this application, training the pre-training module according to n target segment information including the resource usage feature and timestamp feature of the cloud platform can enable the obtained target prediction model to better understand the dependence between resource usage and time in the cloud platform. In this way, the prediction accuracy of the target prediction model for the resource usage of the cloud platform can be improved.

[0009] Optionally, the processing the first resource usage information to obtain n target segment information includes:

[0010] Normalize each of the multiple resource usages to obtain the resource usage characteristics of each resource usage;

[0011] Extract the timestamp characteristics of each timestamp among the multiple timestamps;

[0012] Generate time series information based on the resource usage characteristics of the multiple resource usages and the timestamp characteristics of the multiple timestamps, where the time series information includes multiple target characteristics, and each target characteristic among the multiple target characteristics includes a timestamp characteristic and a corresponding resource usage characteristic;

[0013] Obtain the n target segment information based on the time series information.

[0014] Optionally, the extracting the timestamp characteristics of each timestamp among the multiple timestamps includes:

[0015] For any one timestamp among the multiple timestamps, perform at least one of the following multiple operations on the timestamp to obtain the one-dimensional or multi-dimensional characteristics of the timestamp as the timestamp characteristic:

[0016] Extract the characteristic of the hour in the timestamp;

[0017] Extract the characteristic of the week number in the timestamp;

[0018] Extract the characteristic of the minute in the timestamp;

[0019] Extract the characteristic of the date in the timestamp.

[0020] Optionally, the extracting the characteristic of the hour in the timestamp includes: processing the hour in the timestamp according to a preset sine function to obtain the first characteristic of the hour in the timestamp, and processing the hour in the timestamp according to a preset cosine function to obtain the second characteristic of the hour in the timestamp;

[0021] and / or,

[0022] The extracting the characteristic of the week number in the timestamp includes: normalizing the week number in the timestamp to obtain the characteristic of the week number in the timestamp;

[0023] and / or,

[0024] The extracting the characteristic of the minute in the timestamp includes: determining a target minute number according to the hour and minute in the timestamp, where the target minute number indicates the total number of minutes elapsed on the day; dividing the target minute number by a preset minute number to obtain the characteristic of the minute in the timestamp, and the preset minute number is the total number of minutes in a day;

[0025] and / or

[0026] The feature of extracting the date from the timestamp includes: if the date in the timestamp is a holiday, taking the first value as the feature of the date in the timestamp; if the date in the timestamp is not a holiday, taking the second value as the feature of the date in the timestamp.

[0027] Optionally, generating the time series information according to the resource usage features of the multiple resource usages and the timestamp features of the multiple timestamps includes:

[0028] Concatenating the timestamp feature of each timestamp in the multiple timestamps with the resource usage feature of the corresponding resource usage to obtain multiple first target features;

[0029] If the number of the multiple first target features is not an integer multiple of the preset number, generating at least one second target feature so that the total number of the multiple first target features and the at least one second target feature is an integer multiple of the preset number, and the time series information includes the multiple first target features and the at least one second target feature.

[0030] Optionally, obtaining the n target segment information according to the time series information includes:

[0031] Dividing the time series information into n time segment information, and each time segment information in the n time segment information includes a preset number of the target features;

[0032] For any one of the n time segment information, generating position embedding information according to the positions of the respective target features in the time segment information;

[0033] Generating the target segment information according to the time segment information and the position embedding information.

[0034] Optionally, training the pre-trained model according to the n target segment information includes:

[0035] Let i be equal to 1, and taking the i-th target segment information in the n target segment information as the input data in the i-th training sample;

[0036] Taking the (i + 1)-th target segment information in the n target segment information as the sample label in the i-th training sample;

[0037] Input the input data in the \(i\)-th training sample into the pre-trained model to obtain the output data of the pre-trained model, and adjust the parameters in the pre-trained model according to the output data and the sample label in the \(i\)-th training sample;

[0038] Determine whether \(i\) is equal to \(n - 1\);

[0039] If \(i\) is not equal to \(n - 1\), then let \(i=i + 1\), use the output data as the input data in the \(i\)-th training sample, and re-execute the step of using the \((i + 1)\)-th target segment information in the \(n\) target segment information as the sample label in the \(i\)-th training sample and subsequent steps until \(i\) is equal to \(n - 1\).

[0040] Optionally, after training the pre-trained model according to the \(n\) target segment information to obtain a target prediction model, it further includes:

[0041] Obtain the second resource usage information of the cloud platform, where the second resource usage information is the resource usage information of the cloud platform in the third time period, and the third time period is the time period before the second time period;

[0042] Predict the resource usage of the cloud platform in the second time period according to the second resource usage information and the target prediction model.

[0043] In a second aspect, a model training device is provided, and the device includes:

[0044] A first acquisition module, configured to acquire the first resource usage information of the cloud platform, where the first resource usage information includes multiple resource usages and multiple timestamps, the multiple resource usages correspond to the multiple timestamps one by one, each resource usage in the multiple resource usages is the resource usage of the cloud platform at the moment indicated by the corresponding timestamp, and each moment indicated by the multiple timestamps is within the first time period;

[0045] A processing module, configured to process the first resource usage information to obtain \(n\) target segment information, where each target segment information in the \(n\) target segment information includes a resource usage feature and a timestamp feature, and \(n\) is an integer greater than or equal to 2;

[0046] A training module, configured to train a pre-trained model according to the \(n\) target segment information to obtain a target prediction model, where the target prediction model is used to predict the resource usage of the cloud platform in the second time period.

[0047] In a third aspect, a computer device is provided. The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the model training method described in the first aspect above is implemented.

[0048] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the model training method described in the first aspect above is implemented.

[0049] In a fifth aspect, a computer program product is provided. When the computer program product runs on a computer device, the computer device is caused to execute the model training method described in the first aspect above.

[0050] It can be understood that for the beneficial effects of the second, third, fourth, and fifth aspects above, reference can be made to the relevant descriptions in the first aspect above, and details are not repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 is a schematic structural diagram of a pre-trained model provided by an embodiment of the present application;

[0053] Figure 2 is a schematic structural diagram of another pre-trained model provided by an embodiment of the present application;

[0054] Figure 3 is a schematic diagram of the data processing process of a pre-trained model provided by an embodiment of the present application;

[0055] Figure 4 is a flowchart of a model training method provided by an embodiment of the present application;

[0056] Figure 5 is a schematic diagram of a model training method provided by an embodiment of the present application;

[0057] Figure 6 is a flowchart of a resource usage prediction process provided by an embodiment of the present application;

[0058] Figure 7 is a schematic structural diagram of a model training device provided by an embodiment of the present application;

[0059] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific embodiments

[0060] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe in detail the embodiments of the present application with reference to the accompanying drawings.

[0061] It should be understood that the "multiple" mentioned in the present application refers to two or more. In the description of the present application, unless otherwise specified, " / " means "or", for example, A / B can represent A or B; the "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone these three situations. In addition, in order to clearly describe the technical solutions of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit to be different.

[0062] The statement "an embodiment" or "some embodiments" described in the present application means that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of the present application. Thus, the statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" and the like that appear in different parts of the present application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. In addition, the terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0063] The application scenarios of the embodiments of the present application will be described below.

[0064] The cloud platform provides services such as computing and storage through the data center, enabling users to flexibly utilize resources. Currently, the cloud platform accounts for 4.5% of the global energy consumption. However, the resource utilization rate of the cloud platform is very low, usually less than 30%. Therefore, the cloud platform has the problem of high energy consumption and low output. In the era when cloud services are widely adopted, accurately predicting the resource usage of the cloud platform is crucial.

[0065] When predicting the resource usage of a cloud platform, if the predicted value is lower than the actual value, resource underestimation occurs. This situation may cause the cloud platform to be unable to quickly allocate sufficient resources during sudden high loads, thereby affecting system performance and stability. Additionally, resource underestimation may also lead to an extended service response time or even interruption, ultimately resulting in a poor user experience and damaging the reputation of the cloud service provider. If the predicted value is higher than the actual value, resource overestimation occurs. This situation may cause the cloud platform to be overly large in a large-scale cluster, with an increase in the types and quantities of resources, and the management complexity increasing exponentially. Additionally, resource overestimation will also lead to an increase in user rent, affecting the user experience.

[0066] Due to the highly volatile and diverse patterns of resource usage in the cloud platform, accurately predicting the resource usage of the cloud platform is extremely challenging. In related technologies, time is not considered when training a neural network model for resource usage prediction, and time is one of the important features in the cloud service scenario. Therefore, the prediction accuracy of the trained neural network model is not high.

[0067] For this reason, the embodiments of the present application provide a model training method, which can train a pre-trained model based on the timestamps and corresponding resource usages of a user in a certain time period to obtain a target prediction model with relatively high prediction accuracy. When predicting the resource usage required by the user, the resource usage in the previous time period of the user can be used to predict the resource usage in the next time period through the target prediction model, and the prediction accuracy is relatively high. Thus, it can not only ensure the normal operation of the cloud service to a certain extent, but also avoid increasing the burden on the cloud platform meaninglessly to a certain extent.

[0068] The pre-trained model provided by the embodiments of the present application will be described below.

[0069] Figure 1 is a schematic diagram of a pre-trained model provided by the embodiments of the present application. Refer to Figure 1 In, the pre-trained model 10 may include a decoding module 101 and a projection module 102.

[0070] The pre-trained model 10 is a pre-trained neural network model. Exemplarily, the pre-trained model 10 may be a time series model trained based on self-supervised learning. The pre-trained model 10 is used to predict the resource usage in a certain time period. Optionally, the pre-trained model 10 may be a transformer model. Of course, the pre-trained model may also be other models that can predict the resource usage in a certain time period. The embodiments of the present application do not limit this.

[0071] Those skilled in the art can understand that Figure 1 is only an example of the pre-trained model 10 and does not constitute a limitation on the pre-trained model 10. The pre-trained model 10 may include more or fewer modules than those shown in the figure.

[0072] The input of the pre-training module 10 can be a sequence of a cloud platform in a certain time period, and this sequence can include resource usage characteristics and timestamp characteristics of this time period.

[0073] Exemplarily, the timestamp can be a unix timestamp, etc., and the embodiments of the present application do not limit this.

[0074] Exemplarily, the timestamp can be represented using the Timestamp type of pandas.

[0075] Exemplarily, the resource usage can include one or more of the central processing unit (CPU) time, input / output (I / O) time, bandwidth, etc. of the cloud platform, and the embodiments of the present application do not limit this.

[0076] The decoding module 101 is used to generate an output sequence. The decoding module 101 uses the input sequence as a context representation and predicts the next output based on this representation and the already generated output. The decoding module 101 realizes the conversion from input to output by generating each element of the output sequence one by one.

[0077] Exemplarily, the pre-training model 10 can be a Decoder-Only model. Refer to Figure 2 , the decoding module 101 can be a decoder (Decoder) in the transformer model. The decoder can be stacked by multiple encoders (Encoder). Exemplarily, each encoder can include an attention layer, a feed-forward neural network (FNN), and layer normalization, etc.

[0078] The attention layer can focus on different positions in the input sequence and calculate the dependencies between them. The attention layer can adopt a multi-head attention mechanism.

[0079] The FNN is used to further process the output of the attention layer.

[0080] Layer normalization is used to normalize the input.

[0081] The projection module 102 is used to convert the output of the decoding module 101 into a required specific dimension. The projection module 102 can project the high-dimensional hidden representation of the model into a low-dimensional space. In this way, the original input can be converted into a more compact and more semantically informative representation.

[0082] Exemplarily, the projection module 102 may include a projection head.

[0083] In some embodiments, the projection module 102 may further include an inverse normalization module. The inverse normalization module may convert the normalized data back to its original numerical range or distribution. By inverse normalization, the true characteristics of the data before normalization can be restored.

[0084] In some embodiments, as Figure 3 shown, the pre-trained model 10 may further include an autoregression module, which is used to gradually generate the next output based on the output of the projection module 102.

[0085] Exemplarily, as Figure 3 shown, a sequence is input to the pre-trained model 10, and the sequence includes data 1, data 2, data 3, data 4, and data 5. The decoding module 101 may predict data 2' according to data 1, predict data 3' according to data 1 and data 2, predict data 4' according to data 1, data 2, and data 3, predict data 5' according to data 1, data 2, data 3, and data 4, predict data 6' according to data 1, data 2, data 3, data 4, and data 5, and then output data 2', data 3', data 4', data 5', and data 6' to the projection module 102. The projection module 102 may project data 2', data 3', data 4', data 5', and data 6', and output the projected data 2', data 3', data 4', data 5', and data 6' to the autoregression module. The autoregression module may predict the next-step data based on data 2', data 3', data 4', data 5', and data 6', and output data 3", data 4", data 5", data 6", etc.

[0086] Optionally, the sequences of each cloud platform at each time period may be used as training samples to train the pre-trained model 10.

[0087] After the pre-trained model 10 is trained, it can be used to predict the resource usage of a certain cloud platform at certain time periods. In the embodiments of the present application, during the use of the pre-trained model 10, in order to improve the prediction accuracy of the pre-trained model 10, the pre-trained model 10 may be fine-tuned according to the historical resource usage of the cloud platform at some time periods, so as to improve the prediction accuracy of the pre-trained model 10 for the cloud platform.

[0088] The model training method provided by the embodiments of the present application will be explained in detail below.

[0089] Figure 4 is a flowchart of a model training method provided by the embodiments of the present application. Refer toFigure 4 , the method includes the following steps:

[0090] Step 401: The computer device obtains the first resource usage information of the cloud platform. The first resource usage information includes multiple resource usages and multiple timestamps. The multiple resource usages correspond to the multiple timestamps one by one. Each resource usage in the multiple resource usages is the resource usage of the cloud platform at the moment indicated by the corresponding timestamp. Each moment indicated by the multiple timestamps is within the first time period.

[0091] The first resource usage information is the historical resource usage information of the cloud platform. The first resource usage information includes the resource usage of the cloud platform at each moment within the first time period and the timestamps corresponding to each moment within the multiple moments.

[0092] The first time period can be set in advance. Exemplarily, the duration of the first time period can be set to 1 hour, 2 hours, 3 hours, etc. The embodiments of the present application do not limit this.

[0093] In this case, the multiple timestamps in the first resource usage information have a time sequence, and the multiple resource usages corresponding to the multiple timestamps also have a time sequence. That is to say, the first resource usage information is actually a time series. Here, the time series refers to a series of monitoring results of resource usage that changes over time.

[0094] Optionally, the computer device can obtain the resource usage of the cloud platform and the corresponding timestamp every preset duration within the first time period to obtain the first resource usage information. In this case, each timestamp in the multiple timestamps in the first resource usage information is separated from the adjacent timestamp by a preset duration.

[0095] The preset duration can be set in advance. Exemplarily, the preset duration can be set to 20S (seconds), 30S, 40S, etc. The embodiments of the present application do not limit this.

[0096] Exemplarily, assuming that the duration of the first time period is 2 hours, the computer device can start from 14:00 and obtain the resource usage of the cloud platform and the corresponding timestamp every 30S until the time reaches 16:00. Then, the resource usage of the cloud platform and the corresponding timestamp obtained from 14:00 to 16:00 are the first resource usage information. In this case, the first resource usage information is obtained by the computer device by collecting the real-time resource usage of the cloud platform and the corresponding timestamp, where the sampling frequency is 30S.

[0097] Step 402: The computer device processes the first resource usage information to obtain n target segment information. Each piece of target segment information in the n pieces of target segment information includes a resource usage feature and a timestamp feature, where n is an integer greater than or equal to 2.

[0098] Both the resource usage feature and the timestamp feature included in the target segment information have a time sequence, that is, the target segment information is a time series.

[0099] The resource usage feature is a feature obtained by extracting features from the resource usage.

[0100] The timestamp feature is a feature obtained by extracting features from the timestamp.

[0101] Since the target segment information includes the resource usage feature and the timestamp feature of the cloud platform, when the pre-trained model is trained later, it is trained according to the resource usage feature and the timestamp feature of the cloud platform. Therefore, not only can the dependence between the resource usage feature and the timestamp feature be strengthened, but also the prediction performance of the trained model on the cloud platform can be improved.

[0102] In some embodiments, the operation of step 402 may include the following steps (1) to (4):

[0103] Step (1): The computer device performs standardization processing on each of the multiple resource usages to obtain the resource usage feature of each resource usage.

[0104] Exemplarily, the sum of the multiple resource usages can be divided by the total number of the multiple resource usages to obtain an average value, and then the standard deviation of the multiple resource usages is calculated; for any one of the multiple resource usages, the resource usage is subtracted from the average value and then divided by the standard deviation to obtain the resource usage feature of the resource usage.

[0105] By performing standardization processing on each of the multiple resource usages, the scale difference of each feature of each resource usage in the multiple resource usages can be reduced. Standardized data usually helps the model learn features and rules more efficiently, thereby improving the training effect and generalization ability of the model.

[0106] Optionally, step (1) can be implemented by a data normalizer (StandardScaler).

[0107] Step (2): The computer device extracts the timestamp feature of each of the multiple timestamps.

[0108] Exemplarily, the timestamp feature of a timestamp may include one or more of the features of the hour in the timestamp, the feature of the week number, the feature of the minute, the feature of the date, etc.

[0109] In some embodiments, the operation in step (2) may be: for any one of the multiple timestamps, perform at least one of the following multiple operations on the timestamp to obtain a one-dimensional feature or a multi-dimensional feature of the timestamp as the timestamp feature of the timestamp: extract the feature of the hour in the timestamp; extract the feature of the week number in the timestamp; extract the feature of the minute in the timestamp; extract the feature of the date in the timestamp.

[0110] Exemplarily, the operation of the computer device to extract the feature of the hour in the timestamp may be: process the hour in the timestamp according to a preset sine function to obtain a first feature of the hour in the timestamp, and process the hour in the timestamp according to a preset cosine function to obtain a second feature of the hour in the timestamp.

[0111] The preset sine function can be set in advance. Exemplarily, the preset sine function can be set to sin(2πh / 24), where / represents the division sign, h is the hour in the timestamp, and π is the pi.

[0112] The preset cosine function can be set in advance. Exemplarily, the preset cosine function can be set to cos(2πh / 24), where / represents the division sign, h is the hour in the timestamp, and π is the pi.

[0113] Since the values obtained by the preset sine function when h is equal to 12 and when h is equal to 24 are the same, which may cause the first feature extracted not to accurately represent the feature of the hour in the timestamp, the second feature can be continuously extracted to represent the feature of the hour in the timestamp by the first feature and the second feature together. In this way, the accuracy of the extracted feature of the hour can be improved.

[0114] Exemplarily, the operation of the computer device to extract the feature of the week number in the timestamp may be: perform normalization processing on the week number in the timestamp to obtain the feature of the week number in the timestamp.

[0115] Specifically, the specific week number in the timestamp can be determined. The value range of the week number is [0, 6], where 0 represents Monday and 6 represents Sunday. Then, the determined week number is divided by 6 to obtain the feature of the week number in the timestamp.

[0116] Exemplarily, the operation of the computer device to extract the feature of the minute in the timestamp may be: determine a target minute number according to the hour and minute in the timestamp. The target minute number indicates the total number of minutes elapsed on the day; divide the target minute number by a preset minute number to obtain the feature of the minute in the timestamp, and the preset minute number is the total number of minutes in a day.

[0117] For example, the preset number of minutes is 1440.

[0118] For example, if the hour in the timestamp is 2 and the minute is 30, then the target number of minutes is 2×60 + 30 = 150 minutes. Dividing 150 minutes by 1440, the feature of the minute in the timestamp is 0.1, and this 0.1 represents 1 / 10 of the total time of the day that has passed.

[0119] For example, the operation for the computer device to extract the feature of the date in the timestamp can be: if the date in the timestamp is a holiday, then use the first value as the feature of the date in the timestamp; if the date in the timestamp is not a holiday, then use the second value as the feature of the date in the timestamp.

[0120] The first value can be preset. For example, the first value can be set to 1.

[0121] The second value can be preset. For example, the second value can be set to 0.

[0122] By extracting one or more features from the hour, week number, minute, and date in the timestamp, the accuracy of the obtained timestamp features can be improved.

[0123] For example, the computer device can extract information such as the date, hour, minute, and week number in the timestamp through the apply function and the lambda function.

[0124] Step (3): The computer device generates time series information based on the resource usage characteristics of the multiple resource usages and the timestamp characteristics of the multiple timestamps. The time series information includes multiple target features, and each target feature in the multiple target features includes a timestamp feature and a corresponding resource usage feature.

[0125] In some embodiments, the operation for the computer device to generate time series information based on the resource usage characteristics of the multiple resource usages and the timestamp characteristics of the multiple timestamps can be: the computer device splices the timestamp feature of each timestamp in the multiple timestamps with the resource usage feature of the corresponding resource usage to obtain multiple first target features; if the number of the multiple first target features is an integer multiple of the preset number, then use the multiple first target features as the time series information; if the number of the multiple first target features is not an integer multiple of the preset number, then generate at least one second target feature so that the total number of the multiple first target features and the at least one second target feature is an integer multiple of the preset number, and the time series information includes the multiple first target features and the at least one second target feature.

[0126] It should be noted that in the embodiment of the present application, when training the pre-trained model, the number of features in the input data of the pre-trained model can be controlled to be a preset number. In this way, the training efficiency of the pre-trained model can be improved.

[0127] A first target feature is a feature obtained by concatenating the timestamp feature of a timestamp and the resource usage feature corresponding to the resource usage at that timestamp.

[0128] A second target feature is a feature generated by the computer device. Exemplarily, the second target feature can be a feature generated by the computer device in a blank padding (Padding) manner, that is, the second target feature is empty, or the second target feature can be a feature obtained by the computer device copying the last first target feature among the multiple first target features. Of course, the second target feature can also be a feature generated by the computer device in other ways, and the embodiments of the present application do not limit this.

[0129] Exemplarily, the operation of the computer device to generate the at least one second target feature can be implemented through the nn.ReplicationPad1d layer. The nn.ReplicationPad1d layer is used to perform replication padding on the input data.

[0130] The preset number can be set in advance. Exemplarily, the preset number can be 96, 97, 98, etc., and the embodiments of the present application do not limit this.

[0131] Step (4): The computer device obtains the n target segment information according to the time series information.

[0132] After obtaining the n target segment information, the computer device can train the pre-trained model.

[0133] In some embodiments, the operation of step (4) can be: the computer device divides the time series information into n time segment information, and each time segment information in the n time segment information includes a preset number of target features; for any one of the n time segment information, the computer device generates position embedding information according to the positions of the respective target features in the time segment information; the computer device generates target segment information according to the time segment information and the position embedding information.

[0134] The position embedding information is used to indicate the time order of the respective target features in the time segment information.

[0135] Exemplarily, the target feature can be a feature vector. In this case, the time segment information can be a feature matrix, and the position embedding information is a matrix with the same dimension as the feature matrix, where each element represents the position encoding of the corresponding target feature. In this situation, the computer device can add these two matrices, and the resulting matrix is the target segment information.

[0136] Exemplarily, the computer device can use the unfold function to split the time series information into n segments of a preset size (patch_len, i.e., the preset number), that is, obtain n pieces of time segment information.

[0137] Exemplarily, the computer device can perform position embedding on each piece of time segment information in the n pieces of time segment information through the enc_embedding layer to obtain n pieces of target segment information.

[0138] By dividing the time series information into n pieces of time segment information, it is possible to avoid memory overflow caused by inputting overly large data at one time, thereby effectively reducing memory occupancy and computational complexity during the training process. Moreover, by performing position embedding on each target feature in the time segment information, it can help the model understand the order information of the target features in the time series, so that the model can effectively capture the position relationship of the target features in the context, and thus better learn the time dependence during the training process, improving the training efficiency and prediction effect of the model.

[0139] Optionally, after the computer device obtains the n pieces of target segment information, it can regularize the n pieces of target segment information. In this way, it is possible to prevent the trained model from overfitting.

[0140] Exemplarily, the computer device can regularize the n pieces of target segment information through the dropout regularization technique.

[0141] Step 403: The computer device trains the pre-trained model according to the n pieces of target segment information to obtain a target prediction model, and the target prediction model is used to predict the resource usage of the cloud platform in the second time period.

[0142] The second time period is the time period that needs to be predicted.

[0143] Since the target segment information includes the timestamp feature of the timestamp and the resource usage feature of the resource usage corresponding to the timestamp, training the pre-trained model according to the n pieces of target segment information can enable the pre-trained model to better understand the dependence between time and resource usage in the cloud platform. In this way, the accuracy of the target prediction model in predicting the resource usage of the cloud platform in the future time period is relatively high.

[0144] In some embodiments, the operations in step 403 may include the following steps a to d:

[0145] Step a: The computer device sets i equal to 1 and uses the i-th target segment information among the n target segment information as the input data in the i-th training sample.

[0146] Step b: The computer device uses the (i + 1)-th target segment information among the n target segment information as the sample label in the i-th training sample.

[0147] Step c: The computer device inputs the input data in the i-th training sample into the pre-trained model to obtain the output data of the pre-trained model. According to the output data and the sample label in the i-th training sample, the parameters in the pre-trained model are adjusted.

[0148] Exemplarily, the operation of the computer device to adjust the parameters in the pre-trained model according to the output data and the sample label in the i-th training sample may be: determining the loss value between the output data and the sample label in the i-th training sample through a loss function, and adjusting the parameters in the pre-trained model according to the loss value.

[0149] The operation of the computer device to adjust the parameters in the pre-trained model according to the loss value may refer to related technologies, and the embodiments of the present application do not elaborate on this in detail.

[0150] For example, the computer device may use the formula to adjust any parameter in the pre-trained model. Among them, is the adjusted parameter. w is the parameter before adjustment. α is the learning rate, and α can be set in advance. For example, α can be 0.001, 0.000001, etc. The embodiments of the present application do not make a unique limitation on this. dw is the partial derivative of the loss function with respect to w, which can be obtained according to the loss value.

[0151] Step d: The computer device determines whether i is equal to n - 1; if i is not equal to n - 1, then it sets i = i + 1, uses the output data as the input data in the i-th training sample, and re-executes the above step c and subsequent step d until i is equal to n - 1.

[0152] If i is not equal to n - 1, it means that the model training has not been completed. Therefore, the computer device can set i = i + 1 to continue training the model with the next training sample; if i is equal to n - 1, it means that the current model training has been completed, and at this time, the pre-trained model is the target prediction model.

[0153] In the embodiments of the present application, starting from the second training sample, the output data (i.e., prediction data) of the previous pre-trained model is used as the input data in the training sample and input into the pre-trained model. The pre-trained model then makes predictions based on this prediction data. Subsequently, the parameters in the pre-trained model are adjusted according to the sample labels of the output data sample this time, that is, the difference between the output data obtained by making predictions based on this prediction data and the sample labels is reduced. In this way, the prediction accuracy of the trained target prediction model can be further improved.

[0154] In some embodiments, the computer device can also use an early stopping strategy to determine the timing of stopping training. Specifically, after the above step c, the computer device can determine whether to stop training based on the loss value between the output data this time and the sample label in the i-th training sample. For example, if the loss value is less than or equal to a preset threshold, training can be stopped to obtain the target prediction model. If the loss value is greater than the preset threshold, the above step d can be executed.

[0155] For ease of understanding, the following Figure 5 gives an exemplary illustration of the above model training method:

[0156] Figure 5 is a schematic diagram of a model training method provided by an embodiment of the present application.

[0157] See Figure 5 , the computer device obtains the time series of resource usage of the cloud platform, and this time series of resource usage is the above-mentioned first resource usage information. Subsequently, the resource usage in the time series of resource usage is normalized to obtain resource usage features, and the timestamps in the time series of resource usage are encoded to obtain timestamp features, such as five-dimensional encoding in the dimensions of hours, week numbers, minutes, and dates. Subsequently, the timestamp features and resource usage features are concatenated into target features, and then feature padding is performed based on a preset quantity to obtain time series information. Subsequently, the time series information is segmented into multiple time segment information, and then the position encoding of each time segment information is obtained, and position embedding is performed based on the position encoding to obtain multiple target segment information. Subsequently, the pre-trained module including a decoding module and a projection module can be trained based on the multiple target segment information to obtain the target prediction model.

[0158] In the embodiments of the present application, as Figure 5As shown, the preprocessing module not only realizes data standardization, but also performs five-dimensional encoding on timestamps, thus significantly enhancing the data processing efficiency and the utilization rate of time information, and avoiding the processing problems caused by inconsistent data formats or scales. On the basis of ensuring the consistency of data dimensions, the input module enriches the local features of the data and effectively reduces the overfitting risk of the model through strategies such as slicing processing and positional encoding, providing a more robust input for the model. The decoding module uses a multi-layer stacked structure and normalization processing, which not only deepens the extraction of data features, but also ensures the stability and accuracy of the model when dealing with long-distance dependence relationships. The projection module, through precise projection head design and optional inverse normalization processing, makes the output of the model not only conform to the expected format, but also be consistent with the original data scale, facilitating the intuitive understanding and evaluation of the results. The synergistic effect of these modules jointly improves the overall performance and prediction accuracy of the model.

[0159] After obtaining the target prediction model through the above process, the resource usage of the cloud platform can be predicted through the target prediction model subsequently. The resource usage prediction process is described below.

[0160] Figure 6 It is a flowchart of a resource usage prediction process provided by an embodiment of the present application. This method is applied to a computer device, and the computer device can be the same device as the computer device that trains the target prediction model, or can be a different device. See Figure 6 , and this process includes the following steps:

[0161] Step 601: The computer device obtains the second resource usage information of the cloud platform.

[0162] The second resource usage information is the resource usage information of the cloud platform in the third time period. The second resource usage information includes multiple timestamps and multiple resource usages, and the multiple timestamps correspond to the multiple resource usages one by one. Each timestamp in the multiple timestamps indicates a moment within the third time period.

[0163] The third time period is the time period before the second time period. For example, the third time period can be the time period before and adjacent to the second time period. The duration of the third time period can be the same as or different from the duration of the second time period.

[0164] The second time period is the time period for which the resource usage needs to be predicted.

[0165] Step 602: The computer device predicts the resource usage of the cloud platform in the second time period according to the second resource usage information and the target prediction model.

[0166] Exemplarily, the computer device can perform normalization processing on each of the multiple resource usages in the second resource usage information to obtain the resource usage characteristics of each resource usage; perform feature extraction on each of the multiple timestamps in the second resource usage information to obtain the timestamp characteristics of each timestamp; splice the timestamp characteristics of each timestamp in the multiple timestamps with the resource usage characteristics of the corresponding resource usage to obtain multiple target features; generate position embedding information according to the positions of the respective target features in the multiple target features, and generate input data according to the multiple target features and the position embedding information. Input the input data into the target prediction model to obtain the prediction data output by the target prediction model, and the prediction data can indicate the resource usage of the cloud platform in the second time period.

[0167] It should be noted that the predicted time range of the output of the target prediction model in the embodiments of the present application is adjustable, that is, each time the target prediction model is used for prediction, it can be adjusted so that the target prediction model can output the resource usage within the required time range. For example, after inputting the resource usage characteristics and timestamp characteristics of 1 hour into the target prediction model, it can be controlled that the target prediction model outputs the resource usage of the next 1 hour, or it can be controlled that the target prediction model outputs the resource usage of the next 2 hours or the resource usage of other time ranges. In this way, not only can the prediction accuracy be guaranteed, but also the prediction flexibility can be improved.

[0168] In the embodiments of the present application, the resource usage of the cloud platform in certain time periods can be accurately predicted through the target prediction model. Based on this, resource allocation can be carried out more accurately, thereby reducing the resource cost and manpower maintenance cost of the cloud platform, further reducing the risk of cloud service interruption, and reducing the risk of node overload during peak hours.

[0169] In the embodiments of the present application, the computer device obtains the first resource usage information of the cloud platform. The first resource usage information includes multiple resource usages and multiple timestamps. The multiple resource usages correspond to the multiple timestamps one by one. Each of the multiple resource usages is the resource usage of the cloud platform at the moment indicated by the corresponding timestamp. Each of the multiple timestamps indicates a moment within the first time period. After that, the first resource usage information is processed to obtain n target segment information. Since each of the n target segment information includes the resource usage characteristics and timestamp characteristics in the cloud platform, training the pre-trained model according to the n target segment information can enable the trained target prediction model to better understand the dependence between the resource usage and time in the cloud platform. In this way, the prediction accuracy of the target prediction model for the resource usage of the cloud platform can be improved.

[0170] Figure 7It is a schematic structural diagram of a model training device provided by an embodiment of the present application. This device can be implemented by software, hardware, or a combination of both as part or all of a computer device, and this computer device can be the computer device shown below Figure 8 as shown. Refer to Figure 7 , this device includes: a first acquisition module 701, a processing module 702, and a training module 702.

[0171] The first acquisition module 701 is used to acquire first resource usage information of a cloud platform. The first resource usage information includes multiple resource usages and multiple timestamps. The multiple resource usages are in one-to-one correspondence with the multiple timestamps. Each resource usage in the multiple resource usages is the resource usage of the cloud platform at the moment indicated by the corresponding timestamp, and the moment indicated by each timestamp in the multiple timestamps is within a first time period;

[0172] The processing module 702 is used to process the first resource usage information to obtain n target segment information. Each target segment information in the n target segment information includes a resource usage feature and a timestamp feature, and n is an integer greater than or equal to 2;

[0173] The training module 703 is used to train a pre-trained model according to the n target segment information to obtain a target prediction model, and the target prediction model is used to predict the resource usage of the cloud platform in a second time period.

[0174] Optionally, the processing module 702 is used for:

[0175] Performing normalization processing on each resource usage in the multiple resource usages to obtain a resource usage feature of each resource usage;

[0176] Extracting the timestamp feature of each timestamp in the multiple timestamps;

[0177] Generating time series information according to the resource usage features of the multiple resource usages and the timestamp features of the multiple timestamps. The time series information includes multiple target features, and each target feature in the multiple target features includes a timestamp feature and a corresponding resource usage feature;

[0178] Obtaining the n target segment information according to the time series information.

[0179] Optionally, the processing module 702 is used for:

[0180] For any one timestamp in the multiple timestamps, performing at least one of the following multiple operations on the timestamp to obtain a one-dimensional feature or a multi-dimensional feature of the timestamp as the timestamp feature:

[0181] Extracting the feature of the hour in the timestamp;

[0182] Extract the feature of the week number in this timestamp;

[0183] Extract the feature of the minute in this timestamp;

[0184] Extract the feature of the date in this timestamp.

[0185] Optionally, the processing module 702 is used to:

[0186] Process the hour in this timestamp according to a preset sine function to obtain the first feature of the hour in this timestamp, and process the hour in this timestamp according to a preset cosine function to obtain the second feature of the hour in this timestamp.

[0187] Optionally, the processing module 702 is used to:

[0188] Perform normalization processing on the week number in this timestamp to obtain the feature of the week number in this timestamp.

[0189] Optionally, the processing module 702 is used to:

[0190] Determine the target minute number according to the hour and minute in this timestamp, where the target minute number indicates the total number of minutes elapsed on the current day; divide the target minute number by the preset minute number to obtain the feature of the minute in this timestamp, and the preset minute number is the total number of minutes in a day.

[0191] Optionally, the processing module 702 is used to:

[0192] If the date in this timestamp is a holiday, use the first value as the feature of the date in this timestamp; if the date in this timestamp is not a holiday, use the second value as the feature of the date in this timestamp.

[0193] Optionally, the processing module 702 is used to:

[0194] Concatenate the timestamp features of each timestamp in the multiple timestamps with the resource usage features of the corresponding resource usage to obtain multiple first target features;

[0195] If the number of the multiple first target features is not an integer multiple of the preset number, generate at least one second target feature so that the total number of the multiple first target features and the at least one second target feature is an integer multiple of the preset number, and the time series information includes the multiple first target features and the at least one second target feature.

[0196] Optionally, the processing module 702 is used to:

[0197] Divide the time series information into n pieces of time segment information, where each piece of time segment information in the n pieces of time segment information includes a preset number of target features;

[0198] For any one of the n pieces of time segment information, generate position embedding information according to the positions of the respective target features in the time segment information;

[0199] Generate target segment information according to the time segment information and the position embedding information.

[0200] Optionally, the training module 703 is used for:

[0201] Let i be equal to 1, and use the i-th target segment information in the n target segment information as the input data in the i-th training sample;

[0202] Use the (i + 1)-th target segment information in the n target segment information as the sample label in the i-th training sample;

[0203] Input the input data in the i-th training sample into the pre-trained model to obtain the output data of the pre-trained model, and adjust the parameters in the pre-trained model according to the output data and the sample label in the i-th training sample;

[0204] Judge whether i is equal to n - 1;

[0205] If i is not equal to n - 1, then let i = i + 1, use the output data as the input data in the i-th training sample, and re-execute the step of using the (i + 1)-th target segment information in the n target segment information as the sample label in the i-th training sample and subsequent steps until i is equal to n - 1.

[0206] Optionally, the device further includes:

[0207] A second acquisition module, configured to acquire second resource usage information of the cloud platform, where the second resource usage information is the resource usage information of the cloud platform in a third time period, and the third time period is the time period between the second time periods;

[0208] A prediction module, configured to predict the resource usage of the cloud platform in the second time period according to the second resource usage information and the target prediction model.

[0209] In the embodiment of the present application, the device obtains the first resource usage information of the cloud platform. The first resource usage information includes multiple resource usages and multiple timestamps. The multiple resource usages correspond to the multiple timestamps one by one. Each resource usage in the multiple resource usages is the resource usage of the cloud platform at the moment indicated by the corresponding timestamp. The moment indicated by each timestamp in the multiple timestamps is within the first time period. After that, the first resource usage information is processed to obtain n target segment information. Since each target segment information in the n target segment information includes the resource usage characteristics and timestamp characteristics in the cloud platform, training the pre-trained model according to the n target segment information can enable the trained target prediction model to better understand the dependence between the resource usage and time in the cloud platform. In this way, the prediction accuracy of the target prediction model for the resource usage of the cloud platform can be improved.

[0210] It should be noted that when training the model by the model training device provided in the above embodiment, only the division of the above functional modules is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0211] Each functional module in the above embodiment can be integrated in a processing unit, or each functional module exists physically as a separate processing unit, or two or more functional modules are integrated in a processing unit. The above processing unit can be implemented in the form of hardware or software. In addition, the specific names of the functional modules are only for the convenience of distinguishing each other and do not limit the protection scope of the embodiments of the present application.

[0212] The model training device provided in the above embodiment and the embodiment of the model training method belong to the same concept. For the specific working process and the technical effects brought by the functional modules in the above embodiment, reference can be made to the method embodiment part, which will not be elaborated here.

[0213] Figure 8 It is a schematic structural diagram of a computer device provided in an embodiment of the present application. As Figure 8 shown, the computer device 8 includes: a processor 80, a memory 81, and a computer program 82 stored in the memory 81 and executable on the processor 80. When the processor 80 executes the computer program 82, the steps in the model training method in the above embodiment are implemented.

[0214] The computer device 8 can be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device 8 can be a desktop computer, a laptop computer, a network server, a personal digital assistant, a mobile phone, a tablet computer. The embodiments of the present application do not limit the type of the computer device 8. Those skilled in the art can understand that Figure 8 merely examples of the computer device 8, which do not constitute a limitation to the computer device 8, may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0215] The processor 80 can be a central processing unit (CPU), and the processor 80 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0216] In some embodiments, the memory 81 can be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 can also be an external storage device of the computer device 8, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 8. Further, the memory 81 can also include both the internal storage unit and the external storage device of the computer device 8. The memory 81 is used to store an operating system, application programs, a boot loader, data, and other programs, etc. The memory 81 can also be used to temporarily store data that has been output or will be output.

[0217] The embodiments of the present application also provide a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in any of the above method embodiments are implemented.

[0218] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.

[0219] The embodiments of the present application provide a computer program product, and when it runs on a computer, the computer is enabled to execute the steps in the above method embodiments.

[0220] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the photographing device / terminal device, recording medium, computer memory, ROM (read-only memory), RAM (random access memory), CD-ROM (compact disc read-only memory), magnetic tape, floppy disk, and optical data storage device, etc. The computer-readable storage medium mentioned in the present application can be a non-volatile storage medium, in other words, a non-transitory storage medium.

[0221] It should be understood that all or part of the steps to implement the above embodiments can be achieved by software, hardware, firmware or any combination thereof. When implemented using software, it can be fully or partially implemented in the form of a computer program product. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above computer-readable storage medium.

[0222] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0223] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0224] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer device and method can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0225] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0226] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0227] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A model training method, characterized in that: The method comprises: Acquire first resource usage information of the cloud platform, the first resource usage information comprising multiple resource usages and multiple timestamps, the multiple resource usages corresponding to the multiple timestamps one-to-one, each resource usage in the multiple resource usages being the resource usage of the cloud platform at a moment indicated by a corresponding timestamp, and the moment indicated by each timestamp in the multiple timestamps is within a first time period; Processing the first resource usage information to obtain n target fragment information, each of the n target fragment information includes a resource usage feature and a timestamp feature, and n is an integer greater than or equal to 2; The pre-trained model is trained according to the n target segment information to obtain a target prediction model, where the target prediction model is used to predict the resource usage of the cloud platform in the second time period.

2. The method according to claim 1, characterized in that The processing of the first resource usage information to obtain n target fragment information includes: Performing standardization processing on each resource usage of the multiple resource usages to obtain a resource usage feature of each resource usage; Extracting a timestamp feature of each timestamp in the multiple timestamps; Generate time series information according to the resource usage characteristics of the multiple resource usages and the timestamp characteristics of the multiple timestamps, the time series information including multiple target characteristics, each of the multiple target characteristics including a timestamp characteristic and a corresponding resource usage characteristic; The n target segment information is acquired according to the time series information.

3. The method according to claim 2, characterized in that The extracting the timestamp feature of each timestamp from the multiple timestamps includes: For any one of the multiple timestamps, at least one of the following operations is performed on the timestamp to obtain a one-dimensional feature or a multi-dimensional feature of the timestamp as the timestamp feature: Extracting hour features from the timestamp; Extracting features of the week number in the timestamp; Extracting minute features from the timestamp; Extract the date features in the timestamp.

4. The method according to claim 3, characterized in that The extracting the hour feature in the timestamp includes: processing the hour in the timestamp according to a preset sine function to obtain a first feature of the hour in the timestamp, and processing the hour in the timestamp according to a preset cosine function to obtain a second feature of the hour in the timestamp; and / or, The extracting the feature of the week number in the timestamp includes: normalizing the week number in the timestamp to obtain the feature of the week number in the timestamp; and / or, The extracting the minute feature in the timestamp includes: determining a target number of minutes according to the hour and minute in the timestamp, the target number of minutes indicating the total number of minutes that have passed on that day; dividing the target number of minutes by a preset number of minutes to obtain the minute feature in the timestamp, the preset number of minutes being the total number of minutes in a day; and / or, The extracting the feature of the date in the timestamp includes: if the date in the timestamp is a holiday, using the first value as the feature of the date in the timestamp; if the date in the timestamp is not a holiday, using the second value as the feature of the date in the timestamp.

5. The method according to claim 2, characterized in that The generating time series information according to the resource usage characteristics of the multiple resource usages and the timestamp characteristics of the multiple timestamps includes: splicing the timestamp feature of each timestamp in the multiple timestamps with the resource usage feature of the corresponding resource usage to obtain multiple first target features; If the number of the multiple first target features is not an integer multiple of a preset number, at least one second target feature is generated so that the total number of the multiple first target features and the at least one second target feature is an integer multiple of the preset number, and the time series information includes the multiple first target features and the at least one second target feature.

6. The method according to claim 2, characterized in that The acquiring the n target segment information according to the time series information includes: Dividing the time series information into n time segment information, each of the n time segment information includes a preset number of the target features; For any one of the n time segment information, generating position embedding information according to the position of each target feature in the time segment information in the time segment information; The target segment information is generated according to the time segment information and the position embedding information.

7. The method according to claim 1, characterized in that The training of the pre-trained model according to the n target segment information includes: Let i be equal to 1, and use the i-th target segment information among the n target segment information as input data in the i-th training sample; Using the i+1th target segment information among the n target segment information as the sample label in the i-th training sample; Inputting the input data in the i-th training sample into the pre-training model to obtain output data of the pre-training model, and adjusting the parameters in the pre-training model according to the output data and the sample labels in the i-th training sample; Determine whether i is equal to n-1; If i is not equal to n-1, then set i=i+1, use the output data as the input data in the i-th training sample, and re-execute the step of using the i+1-th target segment information among the n target segment information as the sample label in the i-th training sample and subsequent steps until i is equal to n-1.

8. The method according to any one of claims 1 to 7, characterized in that: After the pre-training model is trained according to the n target segment information to obtain the target prediction model, the method further includes: Acquire second resource usage information of the cloud platform, where the second resource usage information is resource usage information of the cloud platform in a third time period, and the third time period is a time period before the second time period; The resource usage of the cloud platform in the second time period is predicted based on the second resource usage information and the target prediction model.

9. A computer device, characterized in that: The computer device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by the processor.

10. A computer program product, characterized in that When the computer program product is executed on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 8.