Computing resource prediction method for large sample behavior simulation
Through deep learning models, the resource configuration network is built and the computing resource requirements for dynamically configuring large-sample simulations has been solved, which has solved the problems of low resource utilization and poor simulation efficiency caused by traditional static configuration, and realized the intelligent allocation and efficient utilization of computing resources.
Patent Information
- Application Number
- CN202510726926.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Large sample simulation jobs have unreasonable static configuration in hardware resource configuration, resulting in low resource utilization or poor simulation efficiency.
The resource configuration network is built using a deep learning model, and through feature extraction and training, the computing resource requirements of each simulation job are dynamically configured to achieve intelligent allocation.
The computing resource utilization rate and simulation processing efficiency of large-sample simulation are improved, resource waste is reduced, and the overall performance of the simulation system is improved.
Smart Images

Figure CN120256135A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer simulation technology, and particularly to a computing resource prediction method for large-sample behavior simulation. Background Art
[0002] Computer simulation is a means of simulating the objective laws of the real world through computer technology. With the development of intelligent algorithms, the variable space of simulation conditions has expanded rapidly, and the combination of simulation jobs to be run has exploded. Large-sample simulation is a necessary means to explore the vast simulation space.
[0003] Large-sample simulation runs require a large amount of underlying hardware resources such as the CPU (Central Processing Unit) and memory. Large-sample simulation exhibits the characteristics of job-level parallelism, and the resource requirements vary greatly among different simulation jobs. Static configuration of computing resources for large-sample simulation jobs through traditional methods will result in low running efficiency or poor utilization rate of computing resources for large-sample simulation. Summary of the Invention
[0004] Based on this, it is necessary to provide a computing resource prediction method for large-sample behavior simulation aiming at the above technical problems, which aims to distinguish the resource requirement differences among different simulation jobs based on a deep learning model, realize dynamic adjustment of computing resource allocation, and improve simulation efficiency and resource utilization rate.
[0005] In some embodiments, the present invention provides a computing resource prediction method for large-sample behavior simulation, including: Obtaining a simulation job sample to be processed, where the simulation job sample includes sample data of one or more behaviors; Performing feature extraction on the simulation job sample to obtain data features corresponding to the simulation job sample; Inputting the data features into a pre-trained resource configuration network to obtain computing resource parameters predicted and output by the resource configuration network, where the computing resource parameters represent the computing resources required to process the simulation job sample; Configuring computing resources based on the computing resource parameters, and using the computing resources to process the simulation job sample to obtain corresponding simulation results.
[0006] In some embodiments, performing feature extraction on the simulation job sample to obtain data features corresponding to the simulation job sample includes: Performing feature extraction on the sample data of each behavior in the simulation job sample to obtain the start time, end time, and behavior attribute corresponding to each behavior, where the behavior attribute represents the type of the behavior; Construct a feature vector corresponding to each behavior based on the start time, the end time, and the behavior attribute of each behavior; Construct data features corresponding to the simulation job sample based on the feature vectors of each behavior in the simulation job sample.
[0007] In some embodiments, input the data features into a pre-trained resource allocation network to obtain the computing resource parameters predicted and output by the resource allocation network, including: Input the data features into a pre-trained resource allocation network. Each node in the input layer of the resource allocation network extracts the latent vector of the data features and inputs the latent vector into the fully connected layer of the resource allocation network; Each node in the fully connected layer obtains a calculation result based on the latent vectors of all nodes in the input layer and inputs the calculation result into the activation function layer of the resource allocation network; The activation function layer performs non-linear activation on the calculation result of the fully connected layer and inputs the activation result into the output layer of the resource allocation network; The output layer outputs the corresponding computing resource parameters according to the activation result of the fully connected layer.
[0008] In some embodiments, the training process of the resource allocation network includes: Obtain a training sample set, where the training sample set includes multiple simulation job samples; Preprocess each simulation job sample to obtain the data label corresponding to the simulation job sample; Extract features from the simulation job sample to obtain the sample data features corresponding to the simulation job sample data; Input the sample data features into the resource allocation network to be trained to obtain the prediction result output by the resource allocation network; Based on the difference between the prediction result and the data label, adjust the network parameters of the resource allocation network until the convergence condition is met to obtain the trained resource allocation network.
[0009] In some embodiments, preprocessing each simulation job sample to obtain the data label corresponding to the simulation job sample includes: Perform simulation processing on each simulation job sample to obtain the target running duration and target computing resources corresponding to each simulation job sample; Based on the target running duration and the target computing resources, obtain the data label of the simulation job sample.
[0010] In some embodiments, the prediction result includes the predicted computing resources and the predicted running duration corresponding to the simulation job sample; based on the difference between the prediction result and the data label, adjusting the network parameters of the resource allocation network until the convergence condition is met, obtaining the trained resource allocation network, including: Construct a mean squared error loss function, and based on the mean squared error loss function, calculate the difference between the predicted computing resources and the target computing resources, and the difference between the predicted running duration and the target running duration; Adjust the network parameters of the resource allocation network by minimizing the loss function until the convergence condition is met, obtaining the trained resource allocation network.
[0011] In some embodiments, the configuring the computing resources based on the computing resource parameters and using the computing resources to process the simulation job sample to obtain the corresponding simulation result includes: Configure the resource manager of the simulation system based on the computing resource parameters; The resource manager invokes the computing resources corresponding to the computing resource parameters to process the simulation job sample, obtaining the corresponding simulation result.
[0012] In some embodiments, the computing resources include one or more of CPU resources, IO resources, GPU resources, and memory resources.
[0013] The computing resource prediction method of the present invention includes extracting data features from the simulation job sample to be processed, inputting the data features into a pre-trained resource allocation network to predict the corresponding computing resource parameters, and configuring and processing the computing resources based on the computing resource parameters to obtain the simulation result. In the embodiments of the present invention, for the large sample behavior simulation scenario, in the case of limited hardware resources, by using the resource allocation network to dynamically configure the required computing resources for each behavior simulation job, the intelligent allocation of computing resources is realized, which can fully improve the utilization rate of computing resources on the basis of ensuring the simulation processing efficiency. Description of the Drawings
[0014] Figure 1 Is the simulation job flow chart of the simulation system in the related art; Figure 2 Is the simulation job flow chart of the simulation system according to some embodiments of the present invention; Figure 3 Is the structural diagram of the resource allocation network according to some embodiments of the present invention; Figure 4 Is the training flow chart of the resource allocation network according to some embodiments of the present invention; Figure 5It is a training flowchart of a resource allocation network according to some embodiments of the present invention; Figure 6 It is a training flowchart of a resource allocation network according to some embodiments of the present invention; Figure 7 It is a flowchart of a computing resource prediction method according to some embodiments of the present invention; Figure 8 It is a flowchart of a computing resource prediction method according to some embodiments of the present invention. Detailed embodiments
[0015] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0016] Computer simulation is a virtual simulation technology that simulates the objective laws of the real world through computer technology. Virtual simulation has the characteristics of flexibility and variability, allowing users to explore objective laws in a virtual space at low cost by flexibly setting simulation conditions. Simulation technology is widely used in fields such as military deduction, industrial manufacturing, and traffic planning.
[0017] With the development of intelligent algorithms, the variable space of the simulation conditions of computer simulation has expanded rapidly, and the combination of simulation jobs to be run has exploded. Large-scale simulation is a necessary means to explore the vast simulation space. Large-scale behavioral simulation refers to taking behavior as the object, and a large number of simulation jobs are executed concurrently to explore the problem space spanned by behavior variables. Large-scale behavioral simulation runs require a large amount of underlying hardware resources such as CPU (Central Processing Unit) and memory, and present new characteristics compared to traditional high-performance computing.
[0018] For example, traditional high-performance computing simulation jobs usually belong to data-level parallelism, that is, a simulation job requires the cooperation of multiple computing nodes to complete. The resource requirements for the same type of simulation jobs are relatively fixed, so a single static configuration can be adopted for multiple simulation jobs.
[0019] However, large-scale simulation belongs to job-level parallelism, that is, the resource requirements of each simulation job depend on the specific behavior, so the resource requirements of different simulation jobs are different. Under the traditional hardware resource management model, the resources for simulation jobs need to be configured in advance. If the traditional single static configuration is adopted for large-scale simulation, the resource requirement differences between different simulation jobs cannot be distinguished. If there is a lot of redundancy in the statically configured computing resources, the resource utilization rate of the computing cluster will be low. On the contrary, if the statically configured computing resources are less, the running efficiency of large-scale behavioral simulation will be reduced.
[0020] Figure 1 The operation process of a traditional simulation system is shown, see Figure 1 As shown, before the simulation starts, first, one or more simulation job samples are generated through a sample generation module. A simulation job sample is various simulation conditions and data configured for a single simulation job.
[0021] For example, taking the urban traffic simulation scenario as an example, a simulation job sample can include sample data corresponding to a series of simulation conditions such as the number of vehicles, the starting position and destination of each vehicle, the road traffic signal pattern and change period, the time step, and the vehicle movement information. And by combining different simulation condition values, a large number of simulation job samples can be generated, that is, the large-sample simulation job scenario described in the present invention.
[0022] Continue to refer to Figure 1 As shown, the resource configuration module needs to configure corresponding computing resources for the simulation job samples. In the traditional solution, generally, the operator statically configures the computing resources in the simulation system based on empirical values. For example, the staff can input parameters of hardware resources such as the number of CPU cores and memory required to execute the simulation job in the simulation system. Then, the behavior simulation module can call the computing resources to execute the simulation task according to the set computing resource parameters and obtain the corresponding simulation results.
[0023] It can be seen that in the simulation system of the traditional solution, the hardware resource parameters for the simulation job need to be manually configured in advance, which has high limitations in the large-sample simulation scenario. For example, in an example scenario, the sample generation module generates thousands or even tens of thousands of simulation job samples by combining simulation conditions. If the computing resource parameters are manually configured for these simulation job samples one by one, the efficiency is very low, and the simulation efficiency and resource utilization rate are highly dependent on the operator's experience. Even for the same simulation scenario, the simulation effects produced by the same staff at different times may also be different.
[0024] If the same computing resource parameters are configured for all simulation job samples, since the job contents of different simulation job samples are different, the required computing resources are also different. For some simulation job samples, there may be a lot of redundant configured computing resources, resulting in a low resource utilization rate. For some simulation job samples, there may be less configured computing resources, resulting in a poor simulation job efficiency.
[0025] Based on the defects existing in the above related technologies, the present invention provides a computing resource prediction method for large-sample behavior simulation. In the large-sample behavior simulation scenario, a deep learning network model is used to dynamically configure the required computing resources for each behavior simulation job, realizing the intelligent allocation of computing resources and improving the large-sample behavior simulation efficiency and the utilization rate of the computing cluster resources.
[0026] Figure 2 shows the operation process of the simulation system in some embodiments of the present invention. Refer to Figure 2 As shown, in the embodiments of the present invention, by constructing a resource configuration network, the resource configuration network is used to intelligently predict the required computing resources for each simulation job sample in a large number of simulation job samples, so that the simulation system can respectively configure the corresponding computing resources for each simulation job sample according to the computing resources predicted by the resource configuration network, realizing the dynamic configuration of resources in a large-sample simulation scenario.
[0027] In the embodiments of the present invention, the resource configuration network is a network model based on a deep neural network (DNN, Deep Neural Network). The deep neural network DNN is a neural network architecture with the ability to process large-scale data and can better handle large-sample simulation scenarios. The advantage of deep learning is that it can learn complex features and patterns from data, so it is a suitable choice for the prediction task of computing resource requirements.
[0028] Figure 3 shows the network structure of the resource configuration network in an exemplary embodiment of the present invention. Refer to Figure 3 As shown, the resource configuration network from the top layer to the bottom layer is successively an input layer, a fully connected layer, an activation function layer, and an output layer. The role of the input layer is to extract latent features of the simulation job sample, so as to extract the high-dimensional information contained in the data and obtain the feature vector corresponding to the job sample. In some embodiments, the input layer includes multiple nodes, each node corresponding to the feature of one dimension of the input data, and the input layer can be a network layer composed of one or more convolutional layers. The role of the fully connected layer is to extract and combine the feature vectors of the input layer. Each node of it is connected to all nodes of the input layer, so as to establish a full connection relationship between the input data and each node. The activation function layer refers to the activation layer of the resource configuration network, which is generally a non-linear activation function. Using the ReLU function, introducing non-linear transformation helps the network learn complex data. The output layer is used to output the final result according to the result of the activation function layer, that is, the computing resource parameter described in the present invention.
[0029] It can be understood that for Figure 3 the resource configuration network shown, it needs to be optimized through network training to improve the network performance. Therefore, the resource configuration network can be divided into two stages, namely the training stage and the application stage. The following will explain these two stages respectively.
[0030] 1. Training stage In the embodiments of the present invention, the network structure of the resource configuration network can be referred toFigure 3 As described in the embodiments, the network training process of the resource allocation network will be described below in conjunction with Figure 4 the following.
[0031] As Figure 4 shown, in some embodiments, the training process of the resource allocation network according to the examples of the present invention includes: S410. Obtain a training sample set.
[0032] The training sample set refers to a data set used for network training of the resource allocation network. The training sample set generally includes a large number of simulation job samples. As described above, each simulation job sample includes sample data of one or more behaviors. For example, in one example, 500 simulation job samples can be generated by a sample generation module to form a training sample set.
[0033] S420. Preprocess each simulation job sample to obtain a data label corresponding to the simulation job sample.
[0034] First, in the data preparation stage, a corresponding data label (Lable) needs to be generated for each simulation job sample. The data label refers to the true value (GroundTruth) corresponding to the simulation job sample. Specifically, in the embodiments of the present invention, the data label corresponding to each simulation job sample can be understood as: the optimal computing resource parameters required to execute the simulation job sample.
[0035] For example, in one example, taking the number of CPU cores as an example of computing resources, for different simulation job samples, if the number of behaviors they include is different, then the number of CPU cores required to execute the simulation job will also be different. Taking a simulation job sample as an example, if too many CPU cores are called when executing the simulation job sample, it will cause redundant computing resources and reduce the utilization rate of computing resources. On the contrary, if too few CPU cores are called when executing the simulation job sample, the execution efficiency of the simulation job sample will be reduced. Based on this, the goal of computing resource allocation is to find an optimal or relatively optimal number of CPU cores, which can not only meet the high-efficiency execution of the simulation job sample, but also make full use of computing resources and avoid redundant waste of computing resources.
[0036] In the embodiments of the present invention, the data label corresponding to each simulation job sample in the training sample set can be marked through data preprocessing, and the data label represents the optimal computing resource parameters required to execute the simulation job sample. In the embodiments of the present invention, the optimal computing resource parameters included in the data label are defined as the target computing resources.
[0037] It should be noted that during the behavioral simulation operation, the running duration of the simulation operation is also a relatively important indicator. Therefore, in some embodiments of the present invention, the data label may include only the target computing resource, or may further include the target running duration, and the target running duration is the running duration spent when the simulation job sample is executed with the optimal computing resource parameters. The following will be described in conjunction with Figure 5 for illustration.
[0038] As Figure 5 shown, in some embodiments, in the training method of the present invention example, the process of preprocessing each simulation job sample includes: S421. Perform simulation processing on each simulation job sample to obtain the target running duration and target computing resource corresponding to each simulation job sample.
[0039] S422. Based on the target running duration and target computing resource, obtain the data label of the simulation job sample.
[0040] In the embodiments of the present invention, taking a simulation job sample in the training sample set as an example, first, multiple computing resource parameters can be preconfigured, and then the simulation process of the simulation job sample is executed respectively using each computing resource parameter, and then the running duration and resource utilization rate of the simulation job are statistically calculated.
[0041] For example, in one example, after performing simulation processing on the simulation job sample with multiple computing resource parameters, the statistical results are shown in Table 1 below: Table 1
[0042] Combined with the example in Table 1, it can be understood that for the same simulation job sample, when performing the simulation job with different computing resources (such as the number of CPU cores), the corresponding simulation situations may also be different. For example, when the number of CPU cores does not reach the required number, although the computing resource is 100% utilized, due to insufficient computing resources, the simulation efficiency is low, and the corresponding running duration is high. As the number of CPU cores gradually increases, when it reaches the required number, at this time the computing resource just meets the requirements, the computing resource is fully utilized, and at the same time the simulation efficiency reaches the upper limit and the running duration is low. And as the number of CPU cores further increases, although the simulation efficiency reaches the upper limit and the running duration is low, the CPU appears idle and redundant, resulting in a decrease in the utilization rate of the computing resource.
[0043] Therefore, in the embodiments of the present invention, in combination with the simulation statistical results shown in Table 1, the computing resource parameters with the optimal running duration and the highest resource utilization rate can be determined as the target computing resources corresponding to the simulation job sample. At the same time, the running duration corresponding at this time is determined as the target running duration. For example, in the example of Table 1, assuming that the running duration T3 is the smallest and the resource utilization rate r3 is the largest, the number of CPU cores c corresponding at this time can be determined as the target computing resource, and the running duration T3 is the target running duration.
[0044] After determining the target computing resources and the target running duration corresponding to the simulation job sample, a data label corresponding to the simulation job sample can be constructed based on the two. In the example of the present invention, the data label at least includes the target computing resources when running the simulation job sample, and may further include the target running duration.
[0045] The above only takes one simulation job sample in the training sample set as an example to illustrate the process of determining the data label. For each simulation job sample in the training sample set, the above method process is sequentially repeated to obtain the data label corresponding to each simulation job sample in the training sample set, and the present invention will not elaborate on this.
[0046] S430. Extract features from the simulation job sample to obtain the sample data features corresponding to the simulation job sample data.
[0047] In some embodiments of the present invention, the input of the resource allocation network is the data features of the simulation job sample. Therefore, before predicting the computing resource parameters of the simulation job sample, it is necessary to extract features from the simulation job sample to obtain the sample data features.
[0048] Still taking one simulation job sample in the training sample set as an example, the simulation job sample includes sample data of one or more behaviors. For example, taking the simulation job in the military deduction scenario as an example, a military deduction may include multiple actions, and each action corresponds to a behavior. Therefore, in the simulation job sample corresponding to a military deduction, there can be sample data of multiple behaviors.
[0049] In addition, it can be understood that the behavior types corresponding to different behaviors may be different, and thus have different behavior attributes. For example, still taking the above-mentioned military deduction as an example, a certain behavior is specifically moving to the target territory, and its corresponding behavior attribute is, for example, "movement"; a certain behavior is specifically attacking a certain target point, and its corresponding behavior attribute is, for example, "attack", etc.
[0050] In the embodiments of the present invention, a simulation job sample includes sample data of one or more behaviors. The sample data of each behavior may include the behavior attributes corresponding to the behavior, the start time, and the end time. The behavior attributes represent the type of the behavior. The start time indicates the start time of executing the behavior, and the end time indicates the end time of executing the behavior.
[0051] Therefore, in the process of feature extraction for the simulation job sample, a three-dimensional feature vector can be constructed based on the behavior data, start time, and end time of each behavior, and then the sample data features corresponding to the simulation job sample can be obtained according to the feature vectors of all behaviors in the simulation job sample.
[0052] The above only takes one simulation job sample in the training sample set as an example to illustrate the process of feature extraction. For each simulation job sample in the training sample set, the above method process is sequentially repeated to obtain the sample data features corresponding to each simulation job sample in the training sample set, which will not be elaborated in the present invention.
[0053] S440: Input the sample data features into the resource allocation network to be trained, and obtain the prediction result output by the resource allocation network.
[0054] Combined with Figure 3 Taking the network structure of the resource allocation network shown as an example, and still taking one simulation job sample in the training sample set as an example, input the sample data features of the simulation job sample into the resource allocation network to be trained. First, the number of nodes in the input layer is the same as the dimension of the sample data features. Therefore, each node further extracts the latent vector of the sample data features and inputs the latent vector into the fully connected layer. Each node in the fully connected layer is connected to all nodes in the input layer. Therefore, each node establishes a fully connected relationship according to the latent vectors of all nodes in the input layer and inputs the obtained calculation result into the activation function layer.
[0055] The activation function layer can adopt the ReLU activation function. ReLU is a non-linear activation function, which introduces non-linear transformation by setting negative values to zero, expressed as:
[0056] The activation function layer inputs the non-linear activation result to the output layer, and the output layer outputs the corresponding prediction result. For example, in the previous example, the prediction result may include the predicted computing resources and predicted running duration corresponding to the simulation job sample. The predicted computing resources refer to the computing resources required for the simulation job sample predicted by the resource allocation network. Similarly, the predicted running duration indicates the running duration required for the simulation job sample predicted by the resource allocation network.
[0057] S450. Adjust the network parameters of the resource allocation network based on the difference between the prediction result and the data label until the convergence condition is met, and obtain the trained resource allocation network.
[0058] Combined with the foregoing, the prediction result refers to the computing resources and running duration required for the simulation job sample predicted by the resource allocation network, which is a predicted value. And as can be seen from the foregoing, the data label refers to the optimal computing resources and running duration required for the simulation job sample, which is the true value. It can be understood that the smaller the difference between the prediction result and the data label, the higher the prediction accuracy of the resource allocation network, and thus the better the effect. For the network training objective of the resource allocation network, it is to make the prediction result as close as possible to the data label.
[0059] Based on this, during network training, a loss function can be pre-constructed. The loss function is used to represent the difference between the prediction result and the data, and by minimizing this loss function, the network parameters of the resource allocation network are adjusted through backpropagation to achieve optimized training of the resource allocation network.
[0060] In some embodiments, a mean squared error (MSE) loss function can be pre-constructed, and the mean squared error loss function is used to calculate the difference between the prediction result and the data label. The following is combined with Figure 6 for illustration.
[0061] As Figure 6 shown, in some embodiments, the network training process of the example of the present invention includes: S451. Construct a mean squared error loss function, and calculate the difference between the predicted computing resources and the target computing resources, and the difference between the predicted running duration and the target running duration based on the mean squared error loss function.
[0062] S452. Adjust the network parameters of the resource allocation network by minimizing the loss function until the convergence condition is met, and obtain the trained resource allocation network.
[0063] Taking the foregoing embodiment as an example, the prediction result output by the resource allocation network includes predicted computing resources and predicted running duration, which represent the predicted values of the resource allocation network for the simulation job sample. The data label of the simulation job sample includes target computing resources and target running duration, which represent the true values of the simulation job sample.
[0064] Therefore, in the embodiments of the present invention, the mean square error (MSE) loss function can be pre-constructed. The loss function includes two parts, namely the difference between the predicted computing resources and the target computing resources, and the difference between the predicted running duration and the target running duration. The corresponding loss value can be calculated through the mean square error loss function. For the calculation process of the MSE loss function, those skilled in the art can undoubtedly understand and fully implement it with reference to related technologies, and the present invention will not elaborate on this.
[0065] After obtaining the loss value, the network parameters of the resource allocation network can be adjusted and optimized by minimizing the loss function and according to the backpropagation algorithm, and thus an iterative training process for the resource allocation network can be completed.
[0066] The above only takes a simulation job sample in the training sample set as an example to illustrate the network training process. Similarly, by continuously repeating the above training process using the simulation job samples in the training sample set, the network parameters of the resource allocation network can be continuously optimized until the convergence condition is met, and then the network training of the resource allocation network can be completed to obtain the trained resource allocation network.
[0067] It can be understood that the convergence condition for the training of the resource allocation network can be set according to specific requirements. For example, the number of iterations reaches a preset number, or the network performance reaches a target value, etc. The present invention will not elaborate on this.
[0068] 2. Application stage It can be understood that the above training process for the resource allocation network can be completed in advance. Thus, after the training is completed, the trained resource allocation network can be deployed in the simulation system to implement the computing resource prediction method of the present invention. The following will be combined with Figure 7 for illustration.
[0069] As Figure 7 shown, in some embodiments, the computing resource prediction method exemplified by the present invention includes: S710. Obtain the simulation job sample to be processed.
[0070] Combined with Figure 1 shown, in a large-sample simulation scenario, the sample generation module can generate a large number of simulation job samples, and each simulation job sample includes sample data of one or more behaviors. In the following of the present invention, a certain simulation job sample will be taken as an example to illustrate the method process of computing resource allocation.
[0071] S720. Extract features from the simulation job sample to obtain the data features corresponding to the simulation job sample.
[0072] In the aforementioned network training phase type, the input of the resource allocation network is the data features of the simulation job samples. Therefore, it is necessary to extract the features of the simulation job samples in advance to obtain the corresponding data features. The following will be described in conjunction with Figure 8 for illustration.
[0073] As Figure 8 shown, in some embodiments, in the process of the computing resource prediction method according to the example of the present invention for extracting features from the simulation job samples, it includes: S721. Extract features from the sample data of each behavior in the simulation job sample to obtain the start time, end time, and behavior attribute corresponding to each behavior.
[0074] S722. Based on the start time, end time, and behavior attribute of each behavior, construct a feature vector corresponding to each behavior.
[0075] S723. Based on the feature vectors of each behavior in the simulation job sample, construct the data features corresponding to the simulation job sample.
[0076] In the embodiments of the present invention, the simulation job sample includes the sample data of one or more behaviors. For example, taking the simulation job in a military deduction scenario as an example, a military deduction may include multiple actions, and each action corresponds to a behavior. Therefore, in the simulation job sample corresponding to a military deduction, it can include the sample data of multiple behaviors.
[0077] In addition, it can be understood that the behavior types corresponding to different behaviors may be different, and thus have different behavior attributes. For example, still taking the aforementioned military deduction as an example, if a certain behavior is specifically moving to a target territory, its corresponding behavior attribute is, for example, "movement"; if a certain behavior is specifically attacking a certain target point, its corresponding behavior attribute is, for example, "attack", etc.
[0078] In the embodiments of the present invention, the simulation job sample includes the sample data of one or more behaviors, and the sample data of each behavior can include the behavior attribute, start time, and end time corresponding to the behavior. The behavior attribute represents the type of the behavior, the start time represents the start time of executing the behavior, and the end time represents the end time of executing the behavior.
[0079] Therefore, in the process of extracting features from the simulation job sample, a three-dimensional feature vector can be constructed based on the behavior data, start time, and end time of each behavior, and then the data features corresponding to it can be obtained according to the feature vectors of all behaviors in the simulation job sample.
[0080] S730. Input the data features into the pre-trained resource allocation network to obtain the computing resource parameters predicted and output by the resource allocation network.
[0081] In conjunction withFigure 3 For the network structure of the resource allocation network shown, after feature extraction of the simulation job sample, the data features of the simulation job sample can be input into the resource allocation network, which is the network trained by the aforementioned training method.
[0082] First, the number of nodes in the input layer of the resource allocation network is the same as the dimension of the sample data features. Thus, each node further extracts a latent vector from the data features and inputs the latent vector into the fully connected layer. Each node in the fully connected layer is connected to all nodes in the input layer. Thus, each node establishes a fully connected relationship based on the latent vectors of all nodes in the input layer and inputs the obtained calculation result into the activation function layer.
[0083] The activation function layer can adopt the ReLU activation function. ReLU is a non-linear activation function that introduces non-linear transformation by setting negative values to zero, expressed as:
[0084] The activation function layer inputs the non-linear activation result to the output layer, and the output layer outputs the corresponding prediction result.
[0085] In the example of the present invention, the prediction result predicted and output by the resource allocation network is the calculation resource parameter corresponding to the simulation job sample, and this calculation resource parameter represents the optimal calculation resources required to process the simulation job sample. For example, in the previous example, taking the number of CPU cores as the calculation resource, the calculation resource parameter output by the resource allocation network is also the optimal number of CPU cores that need to be called to process the simulation job sample.
[0086] S740. Configure the calculation resources based on the calculation resource parameters, and use the calculation resources to process the simulation job sample to obtain the corresponding simulation result.
[0087] According to the foregoing, the calculation resource parameter predicted and output by the resource allocation network represents the optimal calculation resources required to process the simulation job sample. Thus, in combination with Figure 2 As shown, the simulation system can configure the calculation resources according to this calculation resource parameter, and then the behavior simulation module can call these calculation resources to perform simulation processing on the simulation job sample to obtain the corresponding simulation result.
[0088] Specifically, in some embodiments, the resource manager of the simulation system can configure the corresponding calculation resources for the simulation processing process of the simulation job sample according to the calculation resource parameter predicted and output by the resource allocation network, and then the resource manager can call the configured calculation resources to complete the simulation processing of the simulation job sample to obtain the corresponding simulation result.
[0089] In one example, taking the number of CPU cores as an example of computing resources, the resource manager can automatically configure the computing resources according to the number of CPU cores output by the resource configuration network, and call the corresponding number of CPU cores to perform simulation processing on the simulation job samples to obtain the corresponding simulation results.
[0090] Of course, those skilled in the art can understand that the specific type of computing resources can be set according to the simulation scenario, and is not limited to the number of CPU cores. For example, it can also include IO (Input / Output) resources, GPU (graphics processing unit) resources, memory resources, etc. The present invention places no restrictions on this.
[0091] In addition, the above-described embodiment only illustrates the simulation process for one simulation job sample. For a large-sample simulation scenario, a large number of simulation job samples can sequentially repeat the above method process, and the resource configuration network is used to intelligently and dynamically configure the required computing resources for each simulation job sample. Thus, on the basis of ensuring the simulation processing efficiency, the utilization rate of computing resources can be fully improved. Those skilled in the art can understand this, and the present invention will not elaborate further.
[0092] As can be seen from the above, in the embodiment of the present invention, for a large-sample behavior simulation scenario, in the case of limited hardware resources, by using the resource configuration network to dynamically configure the required computing resources for each behavior simulation job, the intelligent allocation of computing resources is achieved, and on the basis of ensuring the simulation processing efficiency, the utilization rate of computing resources can be fully improved.
[0093] In some embodiments, during the application stage of the resource configuration network, as the large-sample behavior simulation progresses, parameters such as the resource utilization rate and simulation efficiency during the simulation process can also be regularly counted, so as to regularly update or fine-tune the network parameters of the resource configuration network. At the same time, a new training sample set can also be constructed based on historical simulation job samples to retrain the resource configuration network. The process of retraining can refer to the aforementioned network training, so as to further optimize the accuracy of the resource configuration network, and thereby improve the resource utilization rate and simulation efficiency of the simulation system.
[0094] For example, in one example, the average resource utilization rate and average efficiency during the historical simulation process can be counted every 7 days, 15 days, or 30 days, and then the network parameters of the current resource configuration network can be updated or fine-tuned based on the average resource utilization rate and average efficiency.
[0095] For another example, a training set can be constructed every 7 days, 15 days or 30 days based on the simulation job samples in the historical simulation process, and combined with the original training sample set to form a new training data set. Then, the resource allocation network is retrained based on the foregoing training method process, so that the resource allocation network can adapt to the current scenario changes and improve the network accuracy.
[0096] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A computational resource prediction method for large-sample behavior simulation, characterized in that Including: Obtain a simulation job sample to be processed, where the simulation job sample includes sample data of one or more behaviors; Extract features from the sample data of each behavior in the simulation job sample to obtain the start time, end time, and behavior attribute corresponding to each behavior, where the behavior attribute characterizes the type of the behavior; Construct a feature vector corresponding to each behavior based on the start time, the end time, and the behavior attribute of each behavior; Construct the data feature corresponding to the simulation job sample based on the feature vectors of each behavior in the simulation job sample; Input the data feature into a pre-trained resource allocation network to obtain the computing resource parameters predicted and output by the resource allocation network, where the computing resource parameters characterize the computing resources required to process the simulation job sample; Configure computing resources based on the computing resource parameters, and use the computing resources to process the simulation job sample to obtain a corresponding simulation result.
2. The method according to claim 1, wherein Inputting the data feature into a pre-trained resource allocation network to obtain the computing resource parameters predicted and output by the resource allocation network includes: Input the data feature into a pre-trained resource allocation network, and each node in the input layer of the resource allocation network extracts the latent vector of the data feature and inputs the latent vector into the fully connected layer of the resource allocation network; Each node in the fully connected layer obtains a calculation result based on the latent vectors of all nodes in the input layer, and inputs the calculation result into the activation function layer of the resource allocation network; The activation function layer performs non-linear activation on the calculation result of the fully connected layer and inputs the activation result into the output layer of the resource allocation network; The output layer outputs the corresponding computing resource parameters according to the activation result of the fully connected layer.
3. The method according to any one of claims 1 to 2, characterized in that, The training process of the resource allocation network includes: Obtain a training sample set, where the training sample set includes multiple simulation job samples; Preprocess each simulation job sample to obtain the data label corresponding to the simulation job sample; Extract features from the simulation job sample to obtain the sample data features corresponding to the simulation job sample data; Input the sample data features into the resource allocation network to be trained to obtain the prediction result output by the resource allocation network; Based on the difference between the prediction result and the data label, adjust the network parameters of the resource allocation network until the convergence condition is met to obtain the trained resource allocation network.
4. The method according to claim 3, wherein Preprocessing each simulation job sample to obtain the data label corresponding to the simulation job sample includes: Perform simulation processing on each simulation job sample to obtain the target running duration and target computing resources corresponding to each simulation job sample; Based on the target running duration and the target computing resources, obtain the data label of the simulation job sample.
5. The method according to claim 4, wherein The prediction results include the predicted computing resources and the predicted running duration corresponding to the simulation job samples; based on the difference between the prediction results and the data labels, adjusting the network parameters of the resource allocation network until the convergence condition is met, to obtain the trained resource allocation network, including: Construct a mean squared error loss function, and based on the mean squared error loss function, calculate the difference between the predicted computing resources and the target computing resources, and the difference between the predicted running duration and the target running duration; Adjust the network parameters of the resource allocation network by minimizing the loss function until the convergence condition is met, to obtain the trained resource allocation network.
6. The method according to claim 1, wherein The step of allocating computing resources based on the computing resource parameters and using the computing resources to process the simulation job samples to obtain corresponding simulation results includes: Allocate the resource manager of the simulation system based on the computing resource parameters; The resource manager invokes the computing resources corresponding to the computing resource parameters to process the simulation job samples to obtain corresponding simulation results.
7. The method according to claim 1, wherein The computing resources include one or more of CPU resources, IO resources, GPU resources, and memory resources.
Citation Information
Patent Citations
Cloud computing resource intelligent distribution method and device for complex system simulation application
CN111258767A
Cloud simulation memory resource prediction model construction method and memory resource prediction method
CN112181659A
Cloud simulation computing resource prediction method, device and equipment based on sorting learning
CN115344386A