Information processing method and apparatus, device, storage medium, and program product
By combining the periodic prediction meta-model and the periodic determination model, the Stacking method is used to integrate the model, which solves the problem of insufficient generalization capabilities of the cloud disk duration prediction model, and achieves higher prediction accuracy.
Patent Information
- Application Number
- PCT/CN2024/113714
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-08-21
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the cloud disk duration prediction model has weak generalization ability, resulting in low prediction accuracy.
The periodic prediction meta-model is adopted, and the user attribute data and resource usage data are input to at least two periodic prediction models for direct prediction, the resource lifetime information is spliced, and the periodic determination model is used for indirect prediction, and the model is fusion combined with the Stacking method to improve the generalization ability of the model.
It improves the prediction accuracy of cloud disk survival time, enhances the generalization ability of the model, and ensures accurate prediction under different users and resources.
Smart Images

Figure CN2024113714_03072025_PF_FP_ABST
Abstract
Description
Information processing method, device, equipment, storage medium and program product
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311860180.6 and application name “Information Processing Method and Device,” the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0002] The embodiments of the present disclosure relate to the field of computer technology, and in particular to an information processing method, apparatus, device, storage medium, and program product. Background Art
[0003] With the continuous development of the Internet, network infrastructure has also been rapidly iterating and improving, bringing with it the "digital" transformation of projects in various industries. In the process of project processing, storage resources such as cloud disks and network disks are indispensable. Service providers provide cloud disks, network disks and other devices with different durations based on user requirements for storage resources in terms of time, storage space, etc. Since the creation of cloud disks takes a certain amount of time, if the cloud disk duration can be predicted based on user usage data of the cloud disk and the cloud disk can be created in advance, the user experience will be improved. In the existing technology, the prediction of cloud disk duration is usually based on tree models and deep learning classifiers. However, the generalization ability of this prediction model is weak, and the prediction accuracy is not high. Therefore, a more effective method is urgently needed to solve the above problems.
[0004] Summary of the Invention
[0005] In view of this, embodiments of the present disclosure provide an information processing method. One or more embodiments of the present disclosure also relate to an information processing apparatus, a model training method, a model training apparatus, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0006] According to a first aspect of an embodiment of the present disclosure, there is provided an information processing method, including:
[0007] Obtaining user attribute data of resource users and resource usage data for target resources;
[0008] The user attribute data and the resource usage data are input into a cycle prediction metamodel for direct prediction to obtain at least two pieces of resource duration information, the at least two pieces of resource duration information are spliced together, and the spliced results are indirectly predicted to obtain the target resource duration information output by the cycle prediction metamodel for allocating the target resource to the resource user.
[0009] According to a second aspect of an embodiment of the present disclosure, there is provided an information processing apparatus, including:
[0010] an acquisition module configured to acquire user attribute data of a resource user and resource usage data for a target resource;
[0011] The prediction module is configured to input the user attribute data and the resource usage data into a period prediction metamodel for direct prediction, obtain at least two pieces of resource duration information, splice the at least two pieces of resource duration information, and indirectly predict the splicing result to obtain the target resource duration information output by the period prediction metamodel for allocating the target resource to the resource user.
[0012] According to a third aspect of an embodiment of the present disclosure, another information processing method is provided, which is applied to a cloud-side device, including:
[0013] Obtain user attribute data of resource users associated with the end-side device, as well as resource usage data for the target resource;
[0014] Inputting the user attribute data and the resource usage data into a cycle prediction metamodel for direct prediction to obtain at least two pieces of resource duration information, concatenating the at least two pieces of resource duration information, and performing indirect prediction on the concatenated result to obtain target resource duration information for allocating the target resource to the resource user, output by the cycle prediction metamodel;
[0015] The target resource is allocated to the resource user according to the target resource lifetime information, and a resource allocation result is sent to the terminal side device.
[0016] According to a fourth aspect of an embodiment of the present disclosure, another information processing apparatus is provided, which is applied to a cloud-side device, including:
[0017] an acquisition module configured to acquire user attribute data of a resource user associated with the terminal-side device and resource usage data for a target resource;
[0018] The prediction module is configured to input the user attribute data and the resource usage data into a period prediction meta-model for direct prediction, obtain at least two pieces of resource duration information, splice the at least two pieces of resource duration information, and perform indirect prediction on the splicing result to obtain the target resource duration information output by the period prediction meta-model for allocating the target resource to the resource user.
[0019] The allocation module is configured to allocate the target resource to the resource user according to the target resource lifetime information and send the resource allocation result to the terminal side device.
[0020] According to a fifth aspect of an embodiment of the present disclosure, a model training method is provided, including:
[0021] Training at least two initial cycle prediction models based on target samples and sample labels corresponding to the target samples, obtaining at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determining at least two cycle prediction models that meet a first training stop condition based on the training results;
[0022] Training an initial period determination model based on the sample label and the at least two pieces of period prediction information, and determining a period determination model that meets a second training stop condition according to the training result;
[0023] Among them, the at least two cycle prediction models are used to directly predict the resource duration information of the resource user for the target resource, the cycle determination model is used to indirectly predict the resource duration information, and the input of the cycle determination model is the integrated result of the output of the at least two cycle prediction models.
[0024] According to a sixth aspect of an embodiment of the present disclosure, a model training device is provided, comprising:
[0025] a first training module configured to train at least two initial cycle prediction models based on target samples and sample labels corresponding to the target samples, obtain at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determine at least two cycle prediction models that meet a first training stop condition based on the training results;
[0026] The second training module is configured to train the initial cycle determination model based on the sample labels and the at least two cycle prediction information, and determine the cycle determination model that meets the second training stop condition according to the training results; wherein, the at least two cycle prediction models are used to directly predict the resource duration cycle information of the resource user for the target resource, and the cycle determination model is used to indirectly predict the resource duration cycle information, and the input of the cycle determination model is the integrated result of the output of the at least two cycle prediction models.
[0027] According to a seventh aspect of an embodiment of the present disclosure, another model training method is provided, which is applied to a cloud-side device, including:
[0028] The target sample submitted by the receiving end device and the sample label corresponding to the target sample;
[0029] Training at least two initial cycle prediction models based on the target sample and the sample label to obtain at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determining at least two cycle prediction models that meet a first training stop condition based on the training results;
[0030] Training an initial period determination model based on the sample label and the at least two pieces of period prediction information, and determining a period determination model that meets a second training stop condition according to the training result;
[0031] Sending prediction model parameters of the at least two period prediction models and determination model parameters of the period determination model to the end-side device;
[0032] Among them, the at least two cycle prediction models are used to directly predict the resource duration information of the resource user for the target resource, the cycle determination model is used to indirectly predict the resource duration information, and the input of the cycle determination model is the integrated result of the output of the at least two cycle prediction models.
[0033] According to an eighth aspect of an embodiment of the present disclosure, another model training apparatus is provided, which is applied to a cloud-side device, including:
[0034] A receiving module, configured to receive a target sample and a sample label corresponding to the target sample submitted by an end-side device;
[0035] a first training module configured to train at least two initial cycle prediction models based on the target sample and the sample label, obtain at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determine at least two cycle prediction models that meet a first training stop condition based on the training results;
[0036] a second training module configured to train the initial period determination model based on the sample label and the at least two pieces of period prediction information, and determine a period determination model that meets a second training stop condition according to the training result;
[0037] A sending module is configured to send the prediction model parameters of the at least two cycle prediction models and the determination model parameters of the cycle determination model to the end-side device; wherein the at least two cycle prediction models are used to directly predict the resource duration information of the resource user for the target resource, the cycle determination model is used to indirectly predict the resource duration information, and the input of the cycle determination model is the integrated result of the output of the at least two cycle prediction models.
[0038] According to a ninth aspect of an embodiment of the present disclosure, there is provided a computing device, including:
[0039] memory and processor;
[0040] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.
[0041] According to a tenth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions, and the instructions implement the steps of the above method when executed by a processor.
[0042] According to an eleventh aspect of the embodiments of the present disclosure, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.
[0043] An embodiment of the present disclosure obtains user attribute data of a resource user and resource usage data for a target resource; inputs the user attribute data and resource usage data into a period prediction metamodel for direct prediction to obtain at least two pieces of resource duration information; splices the at least two pieces of resource duration information; and indirectly predicts the spliced result to obtain target resource duration information for allocating target resources to the resource user output by the period prediction metamodel; and performs a secondary prediction on the resource duration information of the resource user based on the direct prediction of the resource duration information of the resource user based on the period prediction metamodel, thereby improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] FIG1 is a schematic diagram of a processing process of an information processing method provided by an embodiment of the present disclosure;
[0045] FIG2 is a flow chart of an information processing method provided by an embodiment of the present disclosure;
[0046] FIG3 is a schematic diagram of a model of an information processing method provided by an embodiment of the present disclosure;
[0047] FIG4 is a flowchart of a processing process of an information processing method provided by an embodiment of the present disclosure;
[0048] FIG5 is a flow chart of information processing of an information processing method provided by one embodiment of the present disclosure;
[0049] FIG6 is a schematic structural diagram of an information processing device provided by an embodiment of the present disclosure;
[0050] FIG7 is a flowchart of another information processing method provided by an embodiment of the present disclosure;
[0051] FIG8 is a schematic structural diagram of another information processing device provided by an embodiment of the present disclosure;
[0052] FIG9 is a flow chart of a model training method provided by one embodiment of the present disclosure;
[0053] FIG10 is a schematic structural diagram of a model training device provided by one embodiment of the present disclosure;
[0054] FIG11 is a flowchart of another model training method provided by one embodiment of the present disclosure;
[0055] FIG12 is a schematic structural diagram of another model training device provided by one embodiment of the present disclosure;
[0056] FIG13 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0057] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.
[0058] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0059] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0060] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0061] First, the terms involved in one or more embodiments of the present disclosure are explained.
[0062] Cloud disk: A virtual block storage device provided to users.
[0063] Cloud disk lifecycle: The time it takes for a user to create a cloud disk and release it.
[0064] Gradient Boosting: Gradient boosting is an ensemble learning method that gradually improves the accuracy of the model by optimizing the residual error at each round using gradient descent. Common Gradient Boosting models include XGBoost, LightGBM, and CatBoost.
[0065] One-hot encoding: One-hot encoding, also known as single-bit encoding, uses an N-bit state register to encode N states. Each state has its own independent register bit, and at any given time, only one of the bits is valid. That is, only one bit is 1, and the rest are zero. One-hot encoding uses 0s and 1s to represent some parameters, using an N-bit state register to encode N states.
[0066] Stacking: Fusion model, a model fusion method that stacks multiple basic models.
[0067] Cross-validation involves repeatedly using data, splitting the resulting sample data into different training and test sets. The training set is used to train the model, and the test set is used to evaluate the model's predictions. This allows for multiple different training and test sets, and a sample in one training set may become a sample in the next test set. This is known as "cross-validation."
[0068] Figure 1 is a schematic diagram of the processing process of an information processing method provided by an embodiment of the present disclosure. When predicting resource duration information for a resource user, the prediction process is shown in Figure 1. In the cloud disk resource prediction scenario, the resource user is determined, and the resource user is the cloud resource user. The user attribute data of the resource user and the resource usage data of the resource user using the target resource are obtained, and resource usage characteristics are constructed based on the user attribute data of the resource user and the resource usage data for the target resource. The resource usage characteristics are respectively input into at least two cycle prediction models for direct prediction, so as to realize a prediction of the resource user based on the resource usage characteristics, that is, a preliminary prediction, and obtain the resource duration information output by at least two cycle prediction models. The resource duration information is the cloud disk cycle prediction result. At least two pieces of resource duration information are spliced, and the spliced resource duration is input into the cycle determination model for indirect prediction to obtain the target resource duration information for allocating target resources to the resource user.
[0069] At least two cycle prediction models and the cycle determination model are trained using the same target samples. Therefore, there is a correlation between the at least two cycle prediction models and the cycle determination model. The existence of the cycle determination model can improve the model's generalization ability and prediction accuracy. Based on the prediction of resource duration information for resource users using the cycle prediction model, the cycle determination model is then used to perform a secondary prediction of resource duration information for resource users based on the prediction results of the cycle prediction model, thereby improving prediction accuracy.
[0070] In the present disclosure, an information processing method is provided. The present disclosure also relates to an information processing device, a model training method, a model training device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.
[0071] Referring to FIG. 2 , FIG. 2 shows a flow chart of an information processing method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
[0072] Step 202: Obtain user attribute data of the resource user and resource usage data for the target resource.
[0073] Specifically, resource users refer to users who have resource usage needs, where resources can be storage resources such as cloud disks, computing resources, and other resources that can be provided to users. Users can be resource providers of e-commerce, games, social software and other projects; accordingly, user attribute data includes but is not limited to user geographical characteristics, user level and user tags (e-commerce, games and social software, etc.); resource usage data includes but is not limited to resource type information, resource quantity, payment type, and resource attribute information corresponding to resources used by resource users, project type, and virtual machine attribute information.
[0074] In practical applications, user attribute data and resource usage data can be obtained by reading from the management and control system. If the resource is a cloud disk, user data, virtual machine data, and cloud disk creation and release data can be obtained from the cloud disk management and control system. If the cloud disk is still running when the data is obtained, the current time is the release time of the cloud disk. The cloud disk lifecycle is calculated based on the creation and release times, which is used for subsequent cloud disk lifecycle prediction.
[0075] Step 204: Input the user attribute data and the resource usage data into the cycle prediction metamodel for direct prediction to obtain at least two pieces of resource duration information, splice the at least two pieces of resource duration information, and perform indirect prediction on the splicing result to obtain the target resource duration information output by the cycle prediction metamodel for allocating the target resource to the resource user.
[0076] Specifically, after obtaining the user attribute data of the resource user and the resource usage data for the target resource as mentioned above, the user attribute data and the resource usage data can be input into the cycle prediction metamodel for direct prediction to obtain at least two resource duration information, splice the at least two resource duration information, and indirectly predict the splicing result to obtain the target resource duration information for allocating the target resource to the resource user output by the cycle prediction metamodel, wherein the cycle prediction metamodel is a model combination, and the cycle prediction metamodel includes at least two cycle predictions and a cycle determination model for predicting the user attribute data and resource usage data; the resource duration information is the output information of the cycle prediction model after the user attribute data and resource usage data are input into the cycle prediction model, which represents the prediction result of the cycle prediction model for classifying and predicting the user attribute data and resource usage data; splicing refers to stacking at least two resource duration information in a column manner; the splicing result is the splicing information obtained after stacking at least two resource duration information in a column manner; the target resource duration information is the prediction result obtained by the cycle determination model based on the splicing result.
[0077] Based on this, after obtaining the resource user's user attribute data and resource usage data for the target resource, the user attribute data and resource usage data are input into the cycle prediction meta-model for direct prediction. This combined user attribute data and resource usage data predict the user's resource usage cycle, resulting in at least two pieces of resource duration information. These at least two pieces of resource duration information are then concatenated to obtain a concatenated result. This concatenated result is then subjected to a further prediction, i.e., an indirect prediction, to obtain the target resource duration information output by the cycle prediction meta-model. The target resource duration information represents the target resource allocated to the resource user.
[0078] Furthermore, different models can be used for direct prediction and indirect prediction respectively, thereby improving the generalization ability of the model. The specific implementation is as follows:
[0079] The user attribute data and the resource usage data are respectively input into at least two cycle prediction models included in the cycle prediction meta-model for direct prediction, and resource duration information respectively output by the at least two cycle prediction models is obtained; at least two pieces of resource duration information are spliced, and the splicing result is input into the cycle determination model in the cycle prediction meta-model for indirect prediction, and target resource duration information for allocating the target resource to the resource user is obtained, wherein the at least two cycle prediction models and the cycle determination model use target samples for model training.
[0080] Specifically, the cycle prediction model can be a tree structure model, which uses gradient boosting to implement integrated learning based on classifiers corresponding to multiple tree structures. The cycle prediction model can be any tree model such as XGBoost, LightGBM and CatBoost; at least two cycle prediction models and cycle determination models use target samples for model training, wherein the cycle determination model is used to re-predict the prediction results of at least two cycle prediction models, and the training of the cycle determination model can be implemented by K-fold cross-validation; the target sample refers to the sample data used for training at least two cycle prediction models and cycle determination models.
[0081] Based on this, at least two cycle prediction models are determined, and user attribute data and resource usage data are respectively input into the at least two cycle prediction models for direct prediction, obtaining resource duration information output by the at least two cycle prediction models, thereby achieving preliminary prediction of user attribute data and resource usage data through multiple cycle prediction models. The at least two resource duration information are spliced in the order of each information column, and the at least two resource duration information are spliced into one resource duration information. This resource duration information is input into the cycle determination model as the splicing result for indirect prediction, obtaining target resource duration information. The target resource duration information is used to allocate target resources to resource users. The at least two cycle prediction models and the cycle determination model are trained using target samples. During the model training process, the at least two cycle prediction models are used as base models, and the cycle determination model is used as a meta-model for model fusion.
[0082] In practical applications, stacking is employed when training cycle prediction models. User attribute data, resource usage data, and their corresponding labels are fed directly into at least two cycle prediction models for prediction. Resource lifetime information is then output by the at least two cycle prediction models. This information is then used to construct new sample data for training the cycle determination model. Using model fusion, the at least two cycle prediction models serve as base models, and the cycle determination model serves as a meta-model. The meta-model is trained to combine the prediction results of multiple base models. This prevents overfitting of the base models.
[0083] In summary, based on the prediction of resource duration information for resource users using the period prediction model, the period determination model is then used to perform a secondary prediction of resource duration information for resource users based on the prediction results of the period prediction model, thereby improving prediction accuracy. By combining the period prediction model and the period determination model, the model's generalization ability is improved.
[0084] Furthermore, the training of the period determination model includes: inputting the training samples in the target sample into at least two period prediction models respectively to obtain period prediction information output by at least two period prediction models respectively; splicing at least two pieces of period prediction information to obtain period prediction splicing information; constructing a target sample pair based on the period prediction splicing information and the sample label of the target sample, and writing the target sample pair into a target sample pair set; training the initial period determination model based on the target sample pair set until a period determination model that meets the first training stop condition is obtained.
[0085] Specifically, the period prediction concatenation information refers to information obtained by stacking at least two period prediction information pieces in a column-by-column manner; a target sample pair refers to a sample pair consisting of the period prediction concatenation information as a sample and the sample label of the target sample as a label; and the target sample pair set is used to store the target sample pairs. The first training termination condition can be when the number of training times reaches a prediction threshold, or when, after training the period determination model, the trained period determination model is validated based on a validation set and reaches a preset accuracy prediction. This embodiment does not impose any restrictions on the first training termination condition.
[0086] Based on this, training samples from the target samples are used as training samples for each cycle prediction model and are input into at least two cycle prediction models, respectively, to obtain cycle prediction information output by the at least two cycle prediction models. The at least two pieces of cycle prediction information are concatenated in columns to obtain concatenated cycle prediction information. Target sample pairs are constructed using the concatenated cycle prediction information as samples and the sample labels of the target samples as labels. The target sample pairs are written into a target sample pair set. The initial cycle determination model is trained based on the target sample pairs in the target sample pair set until a cycle determination model that meets a first training stop condition is obtained.
[0087] For example, in the scenario of cloud disk lifecycle prediction, multiple training samples are constructed based on the user's cloud disk usage information, as well as the user's region, level, and tags, to form the target sample. As shown in Figure 3(a), n sets of training samples (feature 1-label 1, feature 2-label 2, and feature n-label n) are input into cycle prediction model 1, cycle prediction model 2, and cycle prediction model 3, respectively. The cycle prediction information output by cycle prediction model 1, cycle prediction model 2, and cycle prediction model 3, respectively, is obtained, namely, predicted value y1, predicted value y2, and predicted value y3. The predicted values y1, y2, and y3 are then concatenated in the order of the information columns. The predicted value 1 in predicted value y1 is concatenated with the predicted value 1 in predicted value y2 and the predicted value 1 in predicted value y3, and so on, until the concatenated cycle prediction information, namely, feature x', is obtained. The period prediction splicing information is used as a new training sample, and the original sample label y of the target sample is used as the new sample label to form a target sample pair. The initial period determination model is trained based on the target sample pair until the training is completed to obtain the period determination model.
[0088] To sum up, during the model training process, at least two initial period prediction models are trained based on the target samples, and then the initial period determination model is trained based on the sample labels corresponding to the target samples and the period prediction information obtained by predicting the target samples by at least two initial period prediction models, thereby improving the generalization ability of the model. After the model training is completed, the prediction accuracy can be improved when performing the prediction task.
[0089] Furthermore, considering that the samples in the target sample pair set are obtained by predicting the target samples based on at least two period prediction models, if a general single model training method is used to train the target sample pair set, the generalization ability of the model will be weak. Therefore, the initial period determination model can be further trained through cross-validation. The specific implementation is as follows:
[0090] Determine the number of periodic determination sub-models based on the training task corresponding to the initial periodic determination model; divide the target sample pair set according to the number of the periodic determination sub-models to obtain a target sample pair subset; use the target sample pair subset to train the periodic determination sub-model corresponding to the training task until a target periodic determination sub-model that meets the second training stop condition is obtained; and form a periodic determination model that meets the first training stop condition based on the target periodic determination sub-model that meets the second training stop condition.
[0091] Specifically, the training task refers to the task of training the initial cycle determination model, including the training task of training the initial cycle determination model using a cross-validation method; the cycle determination sub-model is a verification model divided for the initial cycle determination model, which is used to correspond to different target sample subsets; the second training stop condition can be that the number of training times reaches the prediction threshold, or it can be that after the cycle determination model is trained, the trained cycle determination model is verified based on the verification set and a preset accuracy prediction is achieved. This embodiment does not impose any limitation on the first training stop condition. The target cycle determination sub-model is the training result obtained by training the subset based on the target sample.
[0092] Based on this, a training task corresponding to the initial cycle determination model is determined and executed. During the execution of the training task, the number of cycle determination sub-models corresponding to the initial cycle determination model is determined. The target sample pair set is divided according to the number of cycle determination sub-models to obtain at least one target sample pair subset that matches the number of cycle determination sub-models. The target sample pair subsets are used to train the cycle determination sub-models corresponding to the training task until a target cycle determination sub-model that meets the second training stop condition is obtained. Parameter fusion is performed based on the model parameters of the target cycle determination sub-model that meets the second training stop condition to obtain a cycle determination model that meets the first training stop condition.
[0093] Continuing with the previous example, we use 4-fold cross-validation to train the cycle determination model. After obtaining the target sample pair set consisting of feature x' and label y, we divide the target sample pair set consisting of feature x' and label y into four target sample pair subsets. We then train the initial cycle determination sub-model corresponding to the initial cycle determination model based on each target sample pair subset until a cycle determination model that meets the first training stop condition is obtained.
[0094] In summary, the initial cycle determination model is further trained through cross-validation, thereby improving the generalization ability of the model on the basis of at least two cycle prediction models, and also improving the accuracy of subsequent model predictions.
[0095] Furthermore, the training of any period determination sub-model includes: dividing the target sample pair subset corresponding to the initial period determination sub-model into a training data group and a verification data group; training the initial period determination sub-model based on the training data group to obtain an intermediate period determination sub-model; verifying the intermediate period determination sub-model based on the verification data group, and if the intermediate period determination sub-model meets the verification conditions, using the intermediate period determination sub-model as the target period determination sub-model.
[0096] Specifically, the data in the target sample subset are divided into a training data group and a verification data group according to the function of the data. The training data group is used to train the initial cycle determination sub-model, and the verification data group is used to verify the intermediate cycle determination sub-model obtained after training, and to verify the prediction ability and prediction accuracy of the intermediate cycle determination sub-model; the verification condition can be that after the verification data in the verification data set is input into the intermediate cycle determination sub-model for prediction, the matching degree between the output result of the intermediate cycle determination sub-model and the label data in the verification data set reaches a preset matching degree threshold.
[0097] Based on this, the subset of target sample pairs corresponding to the initial cycle determination sub-model is divided into a training data set and a validation data set according to the data partitioning ratio. The training data set is used to train the initial cycle determination sub-model. After the initial cycle determination sub-model is trained, the trained initial cycle determination sub-model is trained and validated based on the validation data set. The initial cycle determination sub-model is trained based on the training data set to obtain an intermediate cycle determination sub-model. The intermediate cycle determination sub-model is then validated based on the validation data set. If the intermediate cycle determination sub-model meets the validation criteria, the intermediate cycle determination sub-model is used as the target cycle determination sub-model.
[0098] Continuing with the previous example, as shown in Figure 3(b), after obtaining a target sample pair set consisting of feature x' and label y, we divide this set into four target sample pair subsets consisting of feature x' and label y. The target sample pair subsets contain training features 1-label 1 through training features n1-label n1, as well as validation features 1 and label 1. The initial period determination submodel is trained based on the training features in the target sample pair subsets. After training, it is validated based on the validation features, obtaining training and validation results, and completing the training of the initial period determination submodel.
[0099] During the initial training of the cycle determination model, a 4-fold cross-validation approach was used. The target sample pair set consisting of feature x' and label y was divided into four target sample pair subsets. Each target sample pair subset was then divided into a training data set and a validation data set. Cycle determination model 1 was trained on each of the target sample pair subsets, resulting in corresponding training and validation results. The validation results corresponding to the four target sample pair subsets were then integrated to obtain the validation results corresponding to the cycle determination sub-model.
[0100] To sum up, by dividing the target sample pair subset into a training data group and a verification data group, and then using the training data group when training the initial cycle determination sub-model, when the initial cycle determination sub-model training is completed and the intermediate cycle determination sub-model is obtained, the prediction ability of the intermediate cycle determination sub-model can be verified based on the verification data group, thereby ensuring the training effect of the initial cycle determination sub-model.
[0101] Furthermore, the training of any cycle prediction model includes: selecting training samples and sample labels from the target samples and inputting them into the initial cycle prediction model, obtaining initial cycle prediction information output by the initial cycle prediction model, adjusting the parameters of the initial cycle prediction model based on the initial cycle prediction information and the sample labels until a cycle prediction model that meets the third training stop condition is obtained, or adjusting the parameters of the cycle prediction model based on the cycle prediction information and the sample labels until a target cycle prediction model that meets the fourth training stop condition is obtained, and using the target cycle prediction model as the cycle prediction model.
[0102] Specifically, the initial cycle prediction model refers to a cycle prediction model to be trained, which is used to directly predict resource duration cycle information for resource users; the initial cycle prediction information is the prediction information output by the initial cycle prediction model after the training sample is input into the initial cycle prediction model for prediction; the third training stop condition can be that the number of training times reaches the prediction threshold, or that after the cycle determination model is trained, the trained cycle determination model is verified based on the validation set, and the preset accuracy prediction is achieved. Correspondingly, the fourth training stop condition can also be that the number of training times reaches the prediction threshold, or that after the cycle determination model is trained, the trained cycle determination model is verified based on the validation set, and the preset accuracy prediction is achieved. This embodiment does not impose any restrictions on the third training stop condition and the fourth training stop condition.
[0103] Based on this, training samples and sample labels are selected from the target samples and input into the initial cycle prediction model. The initial cycle prediction model can make predictions for the training samples in combination with the sample labels to obtain initial cycle prediction information output by the initial cycle prediction model. The parameters of the initial cycle prediction model are adjusted based on the initial cycle prediction information and sample labels until a cycle prediction model that meets the third training stop condition is obtained. In addition, after the training samples are input into the initial cycle prediction model, the parameters of the cycle prediction model can be adjusted based on the cycle prediction information and sample labels output by the initial cycle prediction model until a target cycle prediction model that meets the fourth training stop condition is obtained. The target cycle prediction model is the cycle prediction model.
[0104] In summary, the parameters of the initial cycle prediction model are adjusted based on the initial cycle prediction information and sample labels, or the parameters of the cycle prediction model are adjusted based on the cycle prediction information and sample labels output by the initial cycle prediction model, thereby achieving diversification of cycle prediction model training.
[0105] Furthermore, considering that the user data and historical resource data of historical users obtained cannot be directly used for model training due to missing values and outliers, it is necessary to clean the user data and historical resource data of historical users before obtaining the target samples for model training. The specific implementation is as follows:
[0106] Acquire user data and historical resource data of historical users; perform data cleaning on the user data and the historical resource data respectively to obtain user-cleaned data corresponding to the user data and historical resource-cleaned data corresponding to the historical resource data; perform data integration on the user-cleaned data and the historical resource-cleaned data to obtain a target sample.
[0107] Specifically, historical users refer to resource users corresponding to target resources; user data refers to data such as the user level, user region, user project type, and payment type for target resources of historical users; historical resource data includes but is not limited to virtual machine data, target resource type, payment type, resource size, and other data; data cleaning operations include but are not limited to data operations such as filling missing values, removing outliers and duplicate values; data integration includes integration of user cleansing data and historical resource cleansing data and feature conversion.
[0108] Based on this, we obtain the user data and historical resource data of historical users. We perform data cleaning operations on these data, such as filling missing values and removing outliers and duplicates, to obtain user-cleaned data corresponding to the user data and historical resource-cleaned data corresponding to the historical resource data. After integrating the user-cleaned data and the historical resource-cleaned data, we perform feature transformation and convert the integrated user-cleaned data and historical resource-cleaned data into vector representations to form the target sample.
[0109] To sum up, by performing data cleaning on the user data and historical resource data of historical users, we can process abnormal values such as missing values and difference values in the user data and historical resource data of historical users, obtain more balanced and standardized target samples, and thus improve the accuracy of subsequent model training.
[0110] Furthermore, considering that the training of the initial period determination model is based on at least two target sample subsets, the initial period determination model corresponds to at least two initial period determination sub-models. After the at least two initial period determination sub-models are trained separately, the training results of the at least two initial period determination sub-models need to be integrated to obtain the period determination model. The specific implementation is as follows:
[0111] Determine at least two target cycle determination sub-models that meet the second training stop condition; perform parameter fusion on model parameters corresponding to the at least two target cycle determination sub-models to obtain target model parameters; and construct a cycle determination model that meets the first training stop condition based on the target model parameters.
[0112] Specifically, the model parameters refer to the model parameters obtained after training the target period determination sub-model based on the target sample pair subset and adjusting the initial model parameters of the target period determination sub-model; the target model parameters are the model parameters used to construct the period determination model obtained by fusing the model parameters corresponding to at least two target period determination sub-models.
[0113] Based on this, at least two target period determination sub-models that meet the second training stop condition are determined. Model parameters corresponding to the at least two target period determination sub-models are determined, and parameter fusion is performed on the at least two model parameters to obtain target model parameters. The target model parameters are used as the result of training the initial period determination model, and a period determination model that meets the first training stop condition is constructed based on the target model parameters.
[0114] In summary, the model parameters corresponding to at least two target period determination sub-models are fused to obtain the target model parameters, and then the period determination model is constructed based on the target model parameters. The period determination model is trained by cross-validation to avoid model overfitting.
[0115] Furthermore, after determining the target resource life cycle information corresponding to the resource user based on the user attribute data of the resource user and the resource usage data of the target resource, the target resource can be allocated to the resource user when the resource user has resource usage needs. The specific implementation is as follows:
[0116] Allocate the target resource to the resource user according to the target resource lifetime information, and send the resource allocation result to the resource user; when the resource usage time of the target resource meets the target resource lifetime information, release the target resource allocated to the resource user.
[0117] Specifically, the resource allocation result refers to the attribute information about the target resource, such as the resource type, cost information, and resource quantity allocated to the resource user; the resource usage time is the length of time the resource user uses the target resource; and the target resource lifetime information represents the time range information from the time the target resource is allocated to the resource user until the release time of the target resource.
[0118] Based on this, the target resource corresponding to the target resource duration is allocated to the resource user according to the target resource duration corresponding to the target resource duration information, and the resource allocation results, including resource type, cost information, and resource quantity, are sent to the resource user. If the resource usage time of the target resource matches the target resource duration corresponding to the target resource duration information, the target resource allocated to the resource user is released.
[0119] Continuing with the previous example, if the target resource's lifecycle information determines that the cloud disk (target resource) has a lifecycle of one month, a cloud disk with a lifecycle of one month is created and made available to the resource user. Information such as the cloud disk fee, cloud disk lifecycle, and cloud disk payment type is sent to the resource user. When the resource user has used the cloud disk for one month, the cloud disk is released, achieving resource recycling.
[0120] To sum up, after determining the target resource lifetime information corresponding to the resource user based on the resource user's user attribute data and the resource usage data for the target resource, the target resource is allocated to the resource user when the resource user has resource usage needs, and the target resource is subsequently released to achieve the purpose of resource recovery, thereby improving the utilization rate of the target resource and avoiding resource waste.
[0121] An embodiment of the present disclosure obtains user attribute data of a resource user and resource usage data for a target resource; inputs the user attribute data and resource usage data into a period prediction metamodel for direct prediction to obtain at least two pieces of resource duration information; splices the at least two pieces of resource duration information; and indirectly predicts the spliced result to obtain target resource duration information for allocating target resources to the resource user output by the period prediction metamodel; and performs a secondary prediction on the resource duration information of the resource user based on the direct prediction of the resource duration information of the resource user based on the period prediction metamodel, thereby improving the prediction accuracy.
[0122] The following further illustrates the information processing method provided by the present disclosure using the application of the information processing method in cloud disk lifecycle prediction as an example, in conjunction with FIG4 . FIG4 shows a flowchart of the processing process of an information processing method provided by an embodiment of the present disclosure, which specifically includes the following steps.
[0123] Step 402: Obtain user data, cloud disk attribute data, and virtual machine attribute data of historical users.
[0124] In practical applications, user data includes, but is not limited to, the proportion of cloud disks with different payment types, the average lifetime, minimum lifetime, and maximum lifetime of pay-as-you-go cloud disks, user region information, user level information, and user project tag information. Cloud disk attribute data includes, but is not limited to, cloud disk payment type, cloud disk type, and cloud disk size. Virtual machine attribute data includes, but is not limited to, specification family, instance type, payment type, and project type.
[0125] Step 404: clean the user data, cloud disk attribute data, and virtual machine attribute data respectively to obtain cleaned user data, cloud disk attribute data, and virtual machine attribute data.
[0126] Perform data cleaning on historical user data, cloud disk attribute data, and virtual machine attribute data. Data cleaning includes but is not limited to filling in missing values, removing outliers and duplicate values.
[0127] Step 406: Integrate the cleaned user data, cloud disk attribute data, and virtual machine attribute data to obtain a target sample.
[0128] After cleaning, the data is processed using encoding methods such as Onehot encoding. Columns whose value distributions do not conform to a Gaussian distribution are logarithmized. Columns with excessively large value ranges are normalized to align their dimensions, which helps improve accuracy. The processed data is then assembled into a dataset, which is divided into training, validation, and test sets for model training, tuning, and evaluation. For imbalanced datasets, methods such as undersampling and oversampling are used to balance the data distribution, ensuring that a consistent number of samples are present within each lifecycle.
[0129] Step 408: Divide the target samples to obtain at least three target sample groups.
[0130] The target sample contains both training samples and labeled data. K-fold cross-training is used to perform a 4-fold split on the target sample. This split creates four datasets, three of which serve as training sets and one as validation set, serving as the formal training data. K-fold cross-training can improve the model's generalization capabilities and prevent overfitting.
[0131] Step 410: Input the training sets in the at least three target sample groups into at least two prediction sub-models respectively. After training the at least two prediction sub-models, input the validation sets in the at least three target sample groups into the at least two trained prediction sub-models to obtain validation results.
[0132] The prediction sub-model uses a tree-based boosting model as the main classifier. Three classifiers, XGBoost, LightGBM, and CatBoost, are selected for model training.
[0133] Step 412: Integrate the verification results and the sample labels of the verification data in the verification set to obtain a preliminary prediction result.
[0134] The verification results output by the three classifiers, XGBoost, LightGBM, and CatBoost, are stacked in columns to form new sample data, that is, preliminary prediction results.
[0135] Step 414: Use the preliminary prediction results to train the result fusion model until a trained result fusion model is obtained.
[0136] The preliminary prediction results are input into the result fusion model for training.
[0137] Step 416: Jointly deploy the at least two trained prediction sub-models and the trained result fusion model.
[0138] After model training is complete, a machine learning platform is selected for deployment. Upper-layer applications implement online reasoning by calling interfaces.
[0139] Step 418: Execute the prediction task based on the deployed prediction sub-model and the deployed result fusion model.
[0140] In practical applications, an online service for cloud disk lifecycle prediction can be provided by training a lifecycle prediction model consisting of at least two prediction sub-models and a fusion model. The specific implementation process is shown in Figure 5. The online service for cloud disk lifecycle prediction includes both offline and real-time links. The offline link performs model training, evaluation, and storage; the real-time link handles model retrieval and real-time inference based on real-time input.
[0141] During the model training phase, after data cleaning of user tags, virtual machine information, and cloud disk information, the lifecycle prediction model to be trained is input offline for model training. The trained lifecycle prediction model is then evaluated and training continues based on the feedback. Model training continues until a lifecycle prediction model that meets the prediction requirements is obtained. After the lifecycle prediction model passes the evaluation, it is deployed. During the real-time inference phase, the deployed lifecycle prediction model provides online services. The model performs real-time inference based on the user's real-time input, obtains the corresponding output of the lifecycle prediction model, and feeds it back to the user, completing the online service.
[0142] In summary, by analyzing historical user usage patterns, we can identify the usage duration characteristics of different user resources. By introducing user tags and levels, we can establish a correlation between the cloud disk lifecycle and the user. By using the stacking method to fuse models, we can improve generalization performance and prediction accuracy without sacrificing generalization performance.
[0143] Corresponding to the above method embodiments, the present disclosure also provides an information processing device embodiment. FIG6 shows a schematic diagram of the structure of an information processing device provided by an embodiment of the present disclosure. As shown in FIG6, the device includes:
[0144] An acquisition module 602 is configured to acquire user attribute data of a resource user and resource usage data for a target resource;
[0145] The prediction module 604 is configured to input the user attribute data and the resource usage data into a period prediction meta-model for direct prediction, obtain at least two pieces of resource duration information, splice the at least two pieces of resource duration information, and perform indirect prediction on the splicing result to obtain the target resource duration information output by the period prediction meta-model for allocating the target resource to the resource user.
[0146] In an optional embodiment, the indirect prediction module 604 is further configured to:
[0147] Inputting the user attribute data and the resource usage data into at least two period prediction models included in the period prediction meta-model respectively for direct prediction, and obtaining resource duration information outputted by the at least two period prediction models respectively;
[0148] At least two pieces of resource duration information are spliced, and the splicing result is input into the cycle determination model in the cycle prediction meta-model for indirect prediction to obtain the target resource duration information for allocating the target resource to the resource user, wherein the at least two cycle prediction models and the cycle determination model use target samples for model training.
[0149] In an optional embodiment, the indirect prediction module 604 is further configured to:
[0150] Inputting the training samples in the target sample into at least two cycle prediction models respectively to obtain cycle prediction information output by the at least two cycle prediction models respectively;
[0151] splicing at least two pieces of period prediction information to obtain period prediction splicing information;
[0152] Constructing a target sample pair according to the period prediction splicing information and the sample label of the target sample, and writing the target sample pair into a target sample pair set;
[0153] The initial period determination model is trained based on the target sample pair set until a period determination model that meets a first training stop condition is obtained.
[0154] In an optional embodiment, the indirect prediction module 604 is further configured to:
[0155] Determine the number of periodic sub-models based on the training task corresponding to the initial periodic determination model;
[0156] Dividing the target sample pair set into a number of sub-models determined according to the period to obtain a target sample pair subset;
[0157] Training the period determination sub-model corresponding to the training task using the target sample pair subset until a target period determination sub-model that meets a second training stop condition is obtained;
[0158] Based on the target period determination sub-model that meets the second training stop condition, a period determination model that meets the first training stop condition is composed.
[0159] In an optional embodiment, the indirect prediction module 604 is further configured to:
[0160] Divide the target sample pair subset corresponding to the sub-model determined by the initial cycle into a training data group and a validation data group;
[0161] Training the initial period determination sub-model based on the training data group to obtain an intermediate period determination sub-model;
[0162] The intermediate period determination sub-model is verified based on the verification data group, and when the intermediate period determination sub-model meets the verification conditions, the intermediate period determination sub-model is used as the target period determination sub-model.
[0163] In an optional embodiment, the indirect prediction module 604 is further configured to:
[0164] Selecting training samples and sample labels from the target samples and inputting them into an initial cycle prediction model, obtaining initial cycle prediction information output by the initial cycle prediction model, and adjusting parameters of the initial cycle prediction model based on the initial cycle prediction information and the sample labels until a cycle prediction model that satisfies a third training stop condition is obtained, or
[0165] The cycle prediction model is adjusted based on the cycle prediction information and the sample labels until a target cycle prediction model that meets a fourth training stop condition is obtained, and the target cycle prediction model is used as the cycle prediction model.
[0166] In an optional embodiment, the indirect prediction module 604 is further configured to:
[0167] Obtain user data and historical resource data of historical users;
[0168] Performing data cleansing on the user data and the historical resource data respectively to obtain user cleansing data corresponding to the user data and historical resource cleansing data corresponding to the historical resource data;
[0169] The user cleansed data and the historical resource cleansed data are integrated to obtain a target sample.
[0170] In an optional embodiment, the indirect prediction module 604 is further configured to:
[0171] Determining at least two target cycles that satisfy a second training stop condition to determine a sub-model;
[0172] Performing parameter fusion on model parameters corresponding to at least two target period determination sub-models to obtain target model parameters;
[0173] A period determination model that satisfies a first training stop condition is constructed based on the target model parameters.
[0174] In an optional embodiment, the indirect prediction module 604 is further configured to:
[0175] Allocate the target resource to the resource user according to the target resource lifetime information, and send a resource allocation result to the resource user;
[0176] When the resource usage time of the target resource satisfies the target resource lifetime information, the target resource allocated to the resource user is released.
[0177] In summary, an embodiment of the present disclosure obtains user attribute data of a resource user and resource usage data for a target resource; inputs the user attribute data and resource usage data into a period prediction metamodel for direct prediction, obtains at least two pieces of resource duration information, splices the at least two pieces of resource duration information, and indirectly predicts the spliced result to obtain the target resource duration information for allocating the target resource to the resource user output by the period prediction metamodel; on the basis of directly predicting the resource duration information of the resource user based on the period prediction metamodel, a secondary prediction is made on the resource duration information of the resource user, thereby improving the prediction accuracy.
[0178] The above is a schematic diagram of an information processing device according to this embodiment. It should be noted that the technical solution of the information processing device and the technical solution of the above-mentioned information processing method are based on the same concept. For details not described in detail in the technical solution of the information processing device, please refer to the description of the technical solution of the above-mentioned information processing method.
[0179] 7 , which shows a flow chart of another information processing method provided according to an embodiment of the present disclosure. The information processing method is applied to a cloud-side device and specifically includes the following steps:
[0180] Step 702: Obtain user attribute data of the resource user associated with the terminal device, and resource usage data for the target resource;
[0181] Step 704: Input the user attribute data and the resource usage data into a cycle prediction metamodel for direct prediction to obtain at least two pieces of resource duration information; concatenate the at least two pieces of resource duration information; and perform indirect prediction on the concatenated result to obtain target resource duration information for allocating the target resource to the resource user, as output by the cycle prediction metamodel.
[0182] Step 706: Allocate the target resource to the resource user according to the target resource lifetime information, and send the resource allocation result to the terminal device.
[0183] In practical applications, information processing can be performed collaboratively by cloud-side devices and end-side devices. The cloud-side device obtains user attribute data of the resource user associated with the end-side device, as well as resource usage data for the target resource. The user attribute data and resource usage data are input into at least two period prediction models included in the period prediction meta-model for direct prediction, and resource duration information output by the at least two period prediction models is obtained, thereby completing a primary prediction of the user attribute data and resource usage data. After splicing the at least two resource duration information pieces, the splicing result is input into the period determination model included in the period prediction meta-model for indirect prediction, and target resource duration information for allocating target resources to the resource user is obtained, thereby achieving a secondary prediction based on the prediction results of the at least two period prediction models. Target resources are allocated to resource users according to the target resource duration information, and the resource allocation results are sent to the end-side device. It should be noted that the at least two period prediction models and the period determination model use target samples for model training, so that the at least two period prediction models and the period determination model learn based on the same target samples during training, thereby improving the generalization ability of the model.
[0184] In summary, based on the prediction of resource duration information for resource users using the period prediction model, the period determination model is then used to perform a secondary prediction of resource duration information for resource users based on the prediction results of the period prediction model, thereby improving prediction accuracy. By combining the period prediction model and the period determination model, the model's generalization ability is improved.
[0185] Corresponding to the above method embodiments, the present disclosure also provides an information processing device embodiment. FIG8 shows a schematic diagram of the structure of another information processing device provided by one embodiment of the present disclosure. The information processing device, applied to a cloud-side device, as shown in FIG8, includes:
[0186] An acquisition module 802 is configured to acquire user attribute data of a resource user associated with the terminal device, and resource usage data for a target resource;
[0187] The prediction module 804 is configured to input the user attribute data and the resource usage data into a period prediction meta-model for direct prediction to obtain at least two pieces of resource duration information, concatenate the at least two pieces of resource duration information, and perform indirect prediction on the concatenated result to obtain target resource duration information for allocating the target resource to the resource user, as output by the period prediction meta-model;
[0188] The sending module 806 is configured to allocate the target resource to the resource user according to the target resource lifetime information, and send the resource allocation result to the terminal side device.
[0189] In summary, based on the prediction of resource duration information for resource users based on the periodic prediction model, the periodic determination model is used to perform a secondary prediction of resource duration information for resource users based on the prediction results of the periodic prediction model, thereby improving the prediction accuracy. By combining the periodic prediction model and the periodic determination model, the generalization ability of the model is improved. The above is a schematic scheme of another information processing device of this embodiment. It should be noted that the technical scheme of the information processing device and the technical scheme of the above-mentioned information processing method belong to the same concept. For details not described in detail in the technical scheme of the information processing device, please refer to the description of the technical scheme of the above-mentioned information processing method.
[0190] Referring to FIG9 , FIG9 shows a flowchart of a model training method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
[0191] Step 902: training at least two initial cycle prediction models based on target samples and sample labels corresponding to the target samples, obtaining at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determining at least two cycle prediction models that meet a first training stop condition based on the training results;
[0192] Step 904: Train the initial cycle determination model based on the sample label and the at least two cycle prediction information, and determine the cycle determination model that meets the second training stop condition according to the training result; wherein, the at least two cycle prediction models are used to directly predict the resource duration cycle information of the resource user for the target resource, and the cycle determination model is used to indirectly predict the resource duration cycle information, and the input of the cycle determination model is the integrated result of the output of the at least two cycle prediction models.
[0193] Specifically, the initial cycle prediction model is a machine learning model to be trained, which is used to directly predict the resource duration cycle information of resource users; correspondingly, the initial cycle determination model is a machine learning model to be trained, which is used to indirectly predict the resource duration cycle information of resource users by combining the prediction results of the initial cycle prediction model, that is, to make a secondary prediction based on the prediction made by the initial cycle prediction model, so as to achieve the purpose of making a further prediction based on the preliminary prediction.
[0194] Based on this, a target sample for model training is obtained, and at least two initial cycle prediction models are trained based on the target sample to obtain at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, as well as at least two cycle prediction models that meet a first training stop condition. A new sample is formed based on the sample label corresponding to the target sample and the at least two pieces of cycle prediction information, and the initial cycle determination model is trained based on the new sample to obtain a cycle determination model that meets the training stop condition. The at least two cycle prediction models are used for direct prediction, and the cycle determination model is used for indirect prediction based on the prediction results of the at least two cycle prediction models.
[0195] Furthermore, after the initial cycle prediction model training is completed to obtain the cycle prediction model, and the initial cycle determination model training is completed to obtain the cycle determination model, the cycle prediction model and the cycle determination model can be jointly deployed to provide prediction services, which is specifically implemented as follows:
[0196] The at least two period prediction models and the period determination model are jointly deployed; and a prediction task is performed based on the at least two deployed period prediction models and the deployed period determination model.
[0197] Specifically, the prediction task may be the time period information of the cloud disk when creating resources such as a cloud disk; the prediction may be made based on the user's historical usage data of the cloud disk.
[0198] Based on this, after obtaining at least two cycle prediction models and cycle determination models, the at least two cycle prediction models and cycle determination models can be jointly deployed and deployed to the machine learning platform, and prediction tasks can be performed based on the at least two deployed cycle prediction models and the deployed cycle determination models.
[0199] For example, after training at least two cycle prediction models and a cycle determination model, these models can be jointly deployed on a machine learning platform. Upon receiving a prediction task for a cloud disk lifecycle, the cloud disk lifecycle can be predicted and a cloud disk can be created based on the prediction results.
[0200] To sum up, during the model training process, at least two initial period prediction models are trained based on the target samples, and then the initial period determination model is trained based on the sample labels corresponding to the target samples and the output information of at least two initial period prediction models for the target samples, thereby improving the generalization ability of the model. After the model training is completed, the prediction accuracy can be improved when performing the prediction task.
[0201] Corresponding to the above method embodiment, the present disclosure also provides an embodiment of a model training device. FIG10 shows a schematic diagram of the structure of a model training device provided by an embodiment of the present disclosure. As shown in FIG10 , the device includes:
[0202] A first training module 1002 is configured to train at least two initial cycle prediction models based on target samples and sample labels corresponding to the target samples, obtain at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determine at least two cycle prediction models that meet a first training stop condition based on the training results;
[0203] The second training module 1004 is configured to train the initial cycle determination model based on the sample label and the at least two cycle prediction information, and determine the cycle determination model that meets the second training stop condition according to the training results; wherein, the at least two cycle prediction models are used to directly predict the resource duration cycle information of the resource user for the target resource, and the cycle determination model is used to indirectly predict the resource duration cycle information, and the input of the cycle determination model is the integrated result of the output of the at least two cycle prediction models.
[0204] To sum up, during the model training process, at least two initial period prediction models are trained based on the target samples, and then the initial period determination model is trained based on the sample labels corresponding to the target samples and the output information of at least two initial period prediction models for the target samples, thereby improving the generalization ability of the model. After the model training is completed, the prediction accuracy can be improved when performing the prediction task.
[0205] The above is a schematic scheme of a model training device of this embodiment. It should be noted that the technical scheme of the model training device and the technical scheme of the above-mentioned model training method are of the same concept. For details not described in detail in the technical scheme of the model training device, please refer to the description of the technical scheme of the above-mentioned model training method.
[0206] Referring to FIG11 , FIG11 shows a flow chart of another model training method provided according to an embodiment of the present disclosure. The model training method is applied to a cloud-side device and specifically includes the following steps:
[0207] Step 1102: Receive a target sample submitted by a terminal device and a sample label corresponding to the target sample;
[0208] Step 1104: training at least two initial cycle prediction models based on the target sample and the sample label, obtaining at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determining at least two cycle prediction models that meet a first training stop condition based on the training results;
[0209] Step 1106: training the initial period determination model based on the sample label and the at least two pieces of period prediction information, and determining a period determination model that satisfies a second training stop condition according to the training result;
[0210] Step 1108: Send the prediction model parameters of the at least two cycle prediction models and the determination model parameters of the cycle determination model to the end-side device; wherein, the at least two cycle prediction models are used to directly predict the resource duration information of the resource user for the target resource, and the cycle determination model is used to indirectly predict the resource duration information, and the input of the cycle determination model is the integrated result of the output of the at least two cycle prediction models.
[0211] In actual applications, when the end-side device has a need for model training, it submits the target sample and the sample label corresponding to the target sample to the cloud-side device. The cloud-side device trains at least two initial cycle prediction models based on the target sample and the sample label corresponding to the target sample, and trains the initial cycle determination model based on at least two pieces of cycle prediction information output by the at least two initial cycle prediction models and the sample label, until at least two cycle prediction models and a cycle determination model are obtained. The prediction model parameters of the at least two cycle prediction models and the determination model parameters of the cycle determination model are sent to the end-side device, and the end-side device constructs at least two cycle prediction models and a cycle determination model based on the obtained prediction model parameters of the at least two cycle prediction models and the determination model parameters of the cycle determination model.
[0212] To sum up, in the process of model training on the cloud-side device, at least two initial period prediction models are trained based on the target samples and sample labels provided by the terminal-side device, and then the initial period determination model is trained based on the sample labels and the output information of the target samples of at least two initial period prediction models, thereby improving the generalization ability of the model. After the model training is completed, the prediction accuracy can be improved when executing the prediction task.
[0213] Corresponding to the above method embodiment, the present disclosure also provides an embodiment of a model training device. FIG12 shows a schematic diagram of the structure of another model training device provided by one embodiment of the present disclosure. As shown in FIG12 , the device includes:
[0214] The receiving module 1202 is configured to receive a target sample and a sample label corresponding to the target sample submitted by an end-side device;
[0215] A first training module 1204 is configured to train at least two initial cycle prediction models based on the target sample and the sample label, obtain at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determine at least two cycle prediction models that meet a first training stop condition based on the training results;
[0216] A second training module 1206 is configured to train the initial period determination model based on the sample label and the at least two pieces of period prediction information, and determine a period determination model that meets a second training stop condition according to the training result;
[0217] The sending module 1208 is configured to send the prediction model parameters of the at least two cycle prediction models and the determination model parameters of the cycle determination model to the end-side device; wherein, the at least two cycle prediction models are used to directly predict the resource duration information of the resource user for the target resource, and the cycle determination model is used to indirectly predict the resource duration information, and the input of the cycle determination model is the integrated result of the output of the at least two cycle prediction models.
[0218] To sum up, in the process of model training on the cloud-side device, at least two initial period prediction models are trained based on the target samples and sample labels provided by the terminal-side device, and then the initial period determination model is trained based on the sample labels and the output information of the target samples of at least two initial period prediction models, thereby improving the generalization ability of the model. After the model training is completed, the prediction accuracy can be improved when executing the prediction task.
[0219] The above is a schematic scheme of another model training device of this embodiment. It should be noted that the technical scheme of the model training device and the technical scheme of the above-mentioned model training method are of the same concept. For details not described in detail in the technical scheme of the model training device, please refer to the description of the technical scheme of the above-mentioned model training method.
[0220] Figure 13 shows a block diagram of a computing device 1300 according to one embodiment of the present disclosure. Components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 via a bus 1330, and a database 1350 is used to store data.
[0221] The computing device 1300 also includes an access device 1340 that enables the computing device 1300 to communicate via one or more networks 1360. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0222] In one embodiment of the present disclosure, the aforementioned components of the computing device 1300 and other components not shown in FIG13 may also be connected to each other, for example, via a bus. It should be understood that the block diagram of the computing device structure shown in FIG13 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.
[0223] Computing device 1300 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1300 may also be a mobile or stationary server.
[0224] The processor 1320 is configured to execute the following computer-executable instructions, which implement the steps of the above method when executed by the processor.
[0225] The above is a schematic solution of a computing device of this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above method.
[0226] An embodiment of the present disclosure further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above method when executed by a processor.
[0227] The above is a schematic solution of a computer-readable storage medium of this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the above method belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above method.
[0228] An embodiment of the present disclosure further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.
[0229] The above is an illustrative solution of a computer program of this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above method belong to the same concept, and any details not described in detail in the technical solution of the computer program can be referred to the description of the technical solution of the above method.
[0230] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0231] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0232] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.
[0233] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0234] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. An information processing method, comprising: Obtaining user attribute data of a resource user and resource usage data for a target resource; Inputting the user attribute data and the resource usage data into a periodic prediction meta-model for direct prediction to obtain at least two pieces of resource survival period information, splicing the at least two pieces of resource survival period information, and performing indirect prediction on the splicing result to obtain the target resource survival period information for allocating the target resource to the resource user output by the periodic prediction meta-model.
2. The information processing method according to claim 1, wherein the inputting the user attribute data and the resource usage data into a periodic prediction meta-model for direct prediction to obtain at least two pieces of resource survival period information, splicing the at least two pieces of resource survival period information, and performing indirect prediction on the splicing result comprises: Respectively inputting the user attribute data and the resource usage data into at least two periodic prediction models included in the periodic prediction meta-model for direct prediction to obtain resource survival period information respectively output by the at least two periodic prediction models; Splicing the at least two pieces of resource survival period information, and inputting the splicing result into a period determination model in the periodic prediction meta-model for indirect prediction to obtain the target resource survival period information for allocating the target resource to the resource user, wherein the at least two periodic prediction models and the period determination model are trained using a target sample.
3. The information processing method according to claim 2, wherein the training of the period determination model comprises: Respectively inputting the training samples in the target sample into at least two periodic prediction models to obtain periodic prediction information respectively output by the at least two periodic prediction models; Splicing the at least two pieces of periodic prediction information to obtain periodic prediction splicing information; Constructing a target sample pair according to the periodic prediction splicing information and the sample label of the target sample, and writing the target sample pair into a target sample pair set; Training an initial period determination model based on the target sample pair set until a period determination model that meets the first training stop condition is obtained.
4. The information processing method according to claim 3, wherein the training of the initial period determination model based on the target sample pair set until a period determination model that meets the first training stop condition is obtained comprises: Determining the number of period determination sub-models according to the training task corresponding to the initial period determination model; Dividing the target sample pair set according to the number of the period determination sub-models to obtain target sample pair subsets; Training the period determination sub-model corresponding to the training task using the target sample pair subsets until a target period determination sub-model that meets the second training stop condition is obtained; Composing a period determination model that meets the first training stop condition based on the target period determination sub-models that meet the second training stop condition.
5. The information processing method according to claim 4, wherein the training of any one period determination sub-model comprises: Dividing the target sample pair subset corresponding to the initial period determination sub-model into a training data group and a validation data group; Train the initial cycle determination sub-model based on the training data set to obtain an intermediate cycle determination sub-model; Validate the intermediate cycle determination sub-model based on the validation data set. When the intermediate cycle determination sub-model meets the validation conditions, use the intermediate cycle determination sub-model as the target cycle determination sub-model.
6. The information processing method according to claim 3, wherein the training of any cycle prediction model includes: Select a training sample and a sample label from the target samples and input them into the initial cycle prediction model to obtain the initial cycle prediction information output by the initial cycle prediction model. Adjust the parameters of the initial cycle prediction model based on the initial cycle prediction information and the sample label until a cycle prediction model that meets the third training stop condition is obtained, or Adjust the parameters of the cycle prediction model based on the cycle prediction information and the sample label until a target cycle prediction model that meets the fourth training stop condition is obtained, and use the target cycle prediction model as the cycle prediction model.
7. The information processing method according to any one of claims 1-6, wherein the determination of the target samples includes: Obtain the user data and historical resource data of historical users; Perform data cleaning on the user data and the historical resource data respectively to obtain the user cleaning data corresponding to the user data and the historical resource cleaning data corresponding to the historical resource data; Integrate the user cleaning data and the historical resource cleaning data to obtain the target samples.
8. The information processing method according to claim 4 or 5, wherein the composition of the cycle determination model that meets the first training stop condition based on the target cycle determination sub-model that meets the second training stop condition includes: Determine at least two target cycle determination sub-models that meet the second training stop condition; Perform parameter fusion on the model parameters corresponding to at least two target cycle determination sub-models to obtain target model parameters; Construct a cycle determination model that meets the first training stop condition based on the target model parameters.
9. The information processing method according to claim 2, the method further includes: Allocate the target resources to the resource users according to the target resource survival cycle information, and send the resource allocation result to the resource users; Release the target resources allocated to the resource users when the resource usage time of the target resources meets the target resource survival cycle information.
10. An information processing method applied to a cloud-side device, including: Obtain the user attribute data of the resource users associated with the end-side device and the resource usage data for the target resources; Input the user attribute data and the resource usage data into the cycle prediction meta-model for direct prediction to obtain at least two resource survival cycle information. Concatenate the at least two resource survival cycle information, and perform indirect prediction on the concatenated result to obtain the target resource survival cycle information for allocating the target resources to the resource users output by the cycle prediction meta-model; Allocate the target resource to the resource user according to the target resource survival cycle information, and send the resource allocation result to the terminal device.
11. A model training method, comprising: Training at least two initial cycle prediction models based on a target sample and a sample label corresponding to the target sample to obtain at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determining at least two cycle prediction models that meet a first training stop condition according to the training result; Training an initial cycle determination model based on the sample label and the at least two pieces of cycle prediction information, and determining a cycle determination model that meets a second training stop condition according to the training result; Wherein, the at least two cycle prediction models are used to directly predict the resource survival cycle information of a resource user for a target resource, the cycle determination model is used to indirectly predict the resource survival cycle information, and the input of the cycle determination model is the integration result output by the at least two cycle prediction models.
12. A model training method, applied to a cloud device, comprising: Receiving a target sample and a sample label corresponding to the target sample submitted by a terminal device; Training at least two initial cycle prediction models based on the target sample and the sample label to obtain at least two pieces of cycle prediction information output by the at least two initial cycle prediction models, and determining at least two cycle prediction models that meet a first training stop condition according to the training result; Training an initial cycle determination model based on the sample label and the at least two pieces of cycle prediction information, and determining a cycle determination model that meets a second training stop condition according to the training result; Sending the prediction model parameters of the at least two cycle prediction models and the determination model parameters of the cycle determination model to the terminal device; Wherein, the at least two cycle prediction models are used to directly predict the resource survival cycle information of a resource user for a target resource, the cycle determination model is used to indirectly predict the resource survival cycle information, and the input of the cycle determination model is the integration result output by the at least two cycle prediction models.
13. An information processing device, comprising: An acquisition module configured to acquire user attribute data of a resource user and resource usage data for a target resource; A prediction module configured to input the user attribute data and the resource usage data into a cycle prediction meta-model for direct prediction to obtain at least two pieces of resource survival cycle information, splice the at least two pieces of resource survival cycle information, and perform indirect prediction on the splicing result to obtain target resource survival cycle information for allocating the target resource to the resource user output by the cycle prediction meta-model.
14. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
16. A computer program product, comprising a computer program, wherein, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Model integrating method and equipment
CN109903840A
Prediction method and device, training method and device, server and a mediumPrediction method, training method and device, server and medium
CN111340244A
Demand prediction method and device for cloud resources
CN114186717A
Business resource capacity dynamic allocation method and device, equipment and storage medium
CN114666224A
Resource lifecycle optimization in disaggregated data centers
US20200099592A1