Scheduling method, system and electronic device based on GPU virtualization in cloud computing environment

By integrating feature matrices of semantic association and topological association in the cloud computing environment and using beta-like distribution analysis, the challenges of GPU virtualization resource scheduling in the cloud computing environment are solved, achieving more efficient resource utilization and cost reduction.

CN115373813BActive Publication Date: 2025-05-09CNNC FUJIAN FUQING NUCLEAR POWER
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210354107.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-03
Publication Date
2025-05-09
Estimated Expiration
2042-04-03

AI Technical Summary

Technical Problem

In a cloud computing environment, how to effectively schedule GPU virtualization resources to improve resource utilization and reduce hardware costs?

Method used

By fusion of feature association between the first feature matrix based on semantic association and the second feature matrix based on topological association, combined with beta-like distribution analysis between nodes, responsive feature components for representing the mapping relationship between non-associated feature distributions are obtained, thereby more accurately allocating operational GPU virtual resources.

Benefits of technology

It realizes more accurate and reasonable client task allocation, improves the utilization rate of GPU virtual resources, and reduces hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115373813B_ABST
    Figure CN115373813B_ABST
Patent Text Reader

Abstract

The present application discloses a scheduling method, system and electronic device based on GPU virtualization in a cloud computing environment, which, on the one hand, can obtain an appropriate coding expression of the task properties of the single client through the feature association fusion between the first feature matrix based on semantic association and the second feature matrix based on topological association, and the query-based retrieval of the feature expression of a single client in the fused association feature space. On the other hand, the positions of the first feature vector used for resource feature expression and the third feature vector used for the association semantic feature expression of the task properties are respectively regarded as nodes, and the responsiveness feature component used to represent the mapping relationship between non-associated feature distributions can be obtained based on the beta-like distribution analysis between nodes, as a practical factor for evaluating the responsiveness behavior between node distributions, and the factor can improve the confidence mapping estimation degree between node distributions, thereby obtaining an appropriate expression of responsiveness between high-dimensional feature distributions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of GPU virtualization in cloud computing, and more specifically, to a scheduling method, system and electronic device based on GPU virtualization in a cloud computing environment. Background Art

[0002] With the rapid development of Internet applications, computers need to process a large amount of data every day. However, the relatively high maintenance costs of traditional software and hardware methods have created a great obstacle to the popularization of related applications. In order to solve this problem, Google launched the "Google 101 Project" in 2006 and proposed the concept of cloud computing. Cloud computing transforms IT capabilities into manageable logical resources, and virtualization is one of the key technologies of cloud computing.

[0003] GPU (Graphics processing unit), also known as graphics processor, was originally used to accelerate computer graphics and has a very mature programming library interface. GPU is a high-performance computing hardware unit used in many fields, such as video encoding and decoding, weather forecasting, and general computing. The number of GPU cores is increasing, and the computing power is becoming more and more powerful. Cloud platforms such as Amazon's EC2ill and Alibaba Cloud have begun to use GPU-assisted computing. If multiple users can share the GPU, the utilization rate of the GPU can be improved and the hardware cost can be reduced. Due to the demand for multi-tasking, GPU virtualization research has become a trend.

[0004] However, how to schedule GPU virtualization resources is a technical problem that needs to be solved in GPU virtualization services. Therefore, in order to more accurately and reasonably treat the allocated client allocation of GPU virtual resources, a scheduling solution based on GPU virtualization in a cloud computing environment is desired. Summary of the invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide a scheduling method, system and electronic device based on GPU virtualization in a cloud computing environment. On the one hand, through the feature association fusion between the first feature matrix based on semantic association and the second feature matrix based on topological association, and the query-style retrieval of the feature expression of a single client in the fused association feature space, the appropriate coding expression of the task nature of the single client can be obtained. On the other hand, the positions of the first feature vector used for resource feature expression and the third feature vector used for the associated semantic feature expression of the task nature are respectively regarded as nodes, and the responsiveness feature component used to represent the mapping relationship between non-associated feature distributions can be obtained based on the beta-like distribution analysis between nodes, as a practical factor for evaluating the responsiveness behavior between node distributions, and the factor can improve the confidence mapping estimation degree between node distributions, thereby obtaining an appropriate expression of responsiveness between high-dimensional feature distributions. Furthermore, it is possible to more accurately and reasonably treat the allocated client allocation operation GPU virtual resources.

[0006] According to one aspect of the present application, a scheduling method based on GPU virtualization in a cloud computing environment is provided, which includes:

[0007] Obtain the available amount of cloud GPU virtual resources at multiple predetermined time points including the current time point;

[0008] Passing the available amounts of cloud GPU virtual resources at the plurality of predetermined time points through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer to obtain a first feature vector;

[0009] Get the task descriptions of the pending computing tasks of multiple clients;

[0010] Passing the task description of the to-be-computed task of each of the clients through a context encoder including an embedding layer to obtain a plurality of feature vectors, and concatenating the plurality of feature vectors to obtain a second feature vector corresponding to the task description of the to-be-computed task of each of the clients;

[0011] Arranging the plurality of second feature vectors of the plurality of clients in two dimensions into a feature matrix and then passing the matrix through a first convolutional neural network to obtain a first feature matrix;

[0012] Acquire a topology matrix of the multiple clients, wherein the value of each non-diagonal position in the topology matrix is ​​the distance between two corresponding clients, and the value of each diagonal position in the topology matrix is ​​zero;

[0013] Passing the topological matrix through a second convolutional neural network model to obtain a topological feature matrix;

[0014] Performing matrix multiplication of the topological feature matrix and the first feature matrix to map the high-dimensional topological features of the topological feature matrix into the high-dimensional feature space of the first feature matrix to obtain a second feature matrix that integrates the topological features and the semantic features of the task properties;

[0015] Extracting the second eigenvector of the client to be assigned from the second eigenvector of the task description of the task to be calculated of each client, and multiplying the second eigenvector of the client to be assigned with the second eigenmatrix to obtain a third eigenvector;

[0016] calculating an inter-nodal response criterion factor between the first eigenvector and the third eigenvector to obtain a fourth eigenvector, the inter-nodal response criterion factor being related to a ratio between a position-wise addition between the first eigenvector and the third eigenvector and a position-wise product of the first eigenvector and the third eigenvector;

[0017] The fourth eigenvector is subjected to a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resource; and

[0018] The computing GPU virtual resource is allocated to the client to be allocated based on the product of the probability value and the available amount of the cloud GPU virtual resource at the current time point.

[0019] According to another aspect of the present application, a scheduling system based on GPU virtualization in a cloud computing environment is provided, which includes:

[0020] A resource availability acquisition unit, used to acquire the availability of cloud GPU virtual resources at multiple predetermined time points including the current time point;

[0021] a temporal encoding unit, configured to obtain a first feature vector by passing the available amounts of the cloud GPU virtual resources at the plurality of predetermined time points obtained by the resource available amount acquisition unit through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer;

[0022] A task description acquisition unit, used to acquire task descriptions of tasks to be calculated from multiple clients;

[0023] a context encoding unit, configured to obtain a plurality of feature vectors by passing the task descriptions of the tasks to be calculated of the clients obtained by the task description obtaining units through a context encoder including an embedding layer, and to concatenate the plurality of feature vectors to obtain a second feature vector corresponding to the task descriptions of the tasks to be calculated of the clients;

[0024] A first convolution unit, configured to two-dimensionally arrange the plurality of second feature vectors of the plurality of clients obtained by the context encoding unit into a feature matrix and then obtain a first feature matrix through a first convolutional neural network;

[0025] A topology matrix acquisition unit, used to acquire a topology matrix of the plurality of clients, wherein the value of each non-diagonal position in the topology matrix is ​​the distance between two corresponding clients, and the value of each diagonal position in the topology matrix is ​​zero;

[0026] A second convolution unit, used for passing the topological matrix obtained by the topological matrix acquisition unit through a second convolutional neural network model to obtain a topological feature matrix;

[0027] A mapping unit, configured to perform matrix multiplication on the topological feature matrix obtained by the second convolution unit and the first feature matrix obtained by the first convolution unit to map the high-dimensional topological features of the topological feature matrix into the high-dimensional feature space of the first feature matrix to obtain a second feature matrix that integrates the topological features and the semantic features of the task properties;

[0028] A third feature vector generating unit is used to extract the second feature vector of the client to be assigned from the second feature vector of the task description of the task to be calculated of the client obtained by each of the context encoding units, and multiply the second feature vector of the client to be assigned by the second feature matrix obtained by the mapping unit to obtain a third feature vector;

[0029] a response criterion factor calculation unit, configured to calculate an inter-node response criterion factor between the first eigenvector obtained by the temporal encoding unit and the third eigenvector obtained by the third eigenvector generation unit to obtain a fourth eigenvector, wherein the inter-node response criterion factor is related to a ratio between a positional point addition between the first eigenvector and the third eigenvector and a positional point product between the first eigenvector and the third eigenvector;

[0030] a probability value calculation unit, configured to pass the fourth eigenvector obtained by the response criterion factor calculation unit through a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resource; and

[0031] An allocating unit is used to allocate the computing GPU virtual resources to the client to be allocated based on the product of the probability value obtained by the probability value calculating unit and the available amount of the cloud GPU virtual resources at the current time point.

[0032] According to another aspect of the present application, an electronic device is provided, comprising: a processor; and a memory, in which computer program instructions are stored, and when the computer program instructions are executed by the processor, the processor executes the scheduling method based on GPU virtualization in the cloud computing environment as described above.

[0033] According to yet another aspect of the present application, a computer-readable medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes the scheduling method based on GPU virtualization in a cloud computing environment as described above.

[0034] Compared with the prior art, the scheduling method, system and electronic device based on GPU virtualization in a cloud computing environment provided by the present application, on the one hand, can obtain an appropriate coding expression of the task nature of the single client through the feature association fusion between the first feature matrix based on semantic association and the second feature matrix based on topological association, and the query-based retrieval of the feature expression of a single client in the fused association feature space. On the other hand, the positions of the first feature vector used for resource feature expression and the third feature vector used for the associated semantic feature expression of the task nature are respectively regarded as nodes, and the responsiveness feature component used to represent the mapping relationship between non-associated feature distributions can be obtained based on the beta-like distribution analysis between nodes, as a practical factor for evaluating the responsiveness behavior between node distributions, and the factor can improve the confidence mapping estimation degree between node distributions, thereby obtaining an appropriate expression of responsiveness between high-dimensional feature distributions. Furthermore, it is possible to more accurately and reasonably treat the allocated client allocation operation GPU virtual resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0036] Figure 1 is a flowchart of a scheduling method based on GPU virtualization in a cloud computing environment according to an embodiment of the present application;

[0037] Figure 2 A schematic diagram of a system architecture of a scheduling method based on GPU virtualization in a cloud computing environment according to an embodiment of the present application;

[0038] Figure 3 A block diagram of a scheduling system based on GPU virtualization in a cloud computing environment according to an embodiment of the present application;

[0039] Figure 4 is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0040] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0041] Scenario Overview

[0042] As mentioned above, with the rapid development of Internet applications, computers need to process a large amount of data every day. However, the relatively high maintenance costs of traditional software and hardware methods have created a great obstacle to the popularization of related applications. In order to solve this problem, Google launched the "Google 101 Project" in 2006 and proposed the concept of cloud computing. Cloud computing transforms IT capabilities into manageable logical resources, and virtualization is one of the key technologies of cloud computing.

[0043] GPU (Graphics processing unit), also known as graphics processor, was originally used to accelerate computer graphics and has a very mature programming library interface. GPU is a high-performance computing hardware unit used in many fields, such as video encoding and decoding, weather forecasting, and general computing. The number of GPU cores is increasing, and the computing power is becoming more and more powerful. Cloud platforms such as Amazon's EC2ill and Alibaba Cloud have begun to use GPU-assisted computing. If multiple users can share the GPU, the utilization rate of the GPU can be improved and the hardware cost can be reduced. Due to the demand for multi-tasking, GPU virtualization research has become a trend.

[0044] However, how to schedule GPU virtualization resources is a technical problem that needs to be solved in GPU virtualization services. Therefore, in order to more accurately and reasonably treat the allocated client allocation of GPU virtual resources, a scheduling solution based on GPU virtualization in a cloud computing environment is desired.

[0045] At present, deep learning and neural networks have been widely used in computer vision, natural language processing, speech signal processing and other fields. In addition, deep learning and neural networks have also shown a level close to or even beyond that of humans in image classification, object detection, semantic segmentation, text translation and other fields.

[0046] The development of deep learning and neural networks has provided new solutions and plans for scheduling based on GPU virtualization in cloud computing environments.

[0047] It should be understood that in the technical solution of the present application, the resource allocation of tasks can essentially be regarded as a problem of resource response to the nature of the task. When considering resource scheduling, it is necessary not only to consider the amount of available resources (that is, how many GPU virtual resources are available), but also to consider the distribution characteristics of the demand for GPU computing resources of the tasks to be processed by each user among multiple users.

[0048] Based on this, in the technical solution of the present application, first, the available amount of GPU virtual resources in the cloud at multiple predetermined time points including the current time point is obtained, and it is passed through a timing encoder to extract high-dimensional correlation feature information of the available amount of GPU virtual resources in the cloud at multiple predetermined time points in the time dimension and the data dimension, thereby obtaining a first feature vector.

[0049] Then, the task descriptions of the tasks to be calculated of multiple clients are obtained, and the task descriptions of the tasks to be calculated are respectively passed through the context encoder to obtain multiple feature vectors with global correlation information, and the multiple feature vectors are cascaded to obtain the second feature vectors corresponding to each task to be calculated. In this way, the multiple second feature vectors can be arranged in two dimensions and passed through the convolutional neural network to obtain the first feature matrix.

[0050] Next, a topological matrix of the multiple clients is obtained to represent the spatial distribution of each client to characterize the data propagation topological characteristics. Here, the value of each position on the non-diagonal position in the topological matrix is ​​the distance between the corresponding two clients, and the value of each position on the diagonal position is zero. The topological matrix is ​​further passed through a convolutional neural network to extract the topological structure feature information of each client, thereby obtaining a calculated topological feature matrix.

[0051] However, since the resource allocation of tasks can essentially be viewed as a problem of resource response to task properties, on the one hand, if the resource and task properties can be properly encoded and expressed in a high-dimensional space, and on the other hand, the responsiveness between the feature distributions of the encoded feature expressions of the two can be properly characterized, then the high-dimensional representation of the responsiveness can be used for classification to obtain a probabilistic representation of the allocation.

[0052] Furthermore, the topological feature matrix is ​​multiplied by the first feature matrix to map the high-dimensional topological features of the topological feature matrix to the high-dimensional feature space of the first feature matrix, thereby obtaining a second feature matrix that integrates the topological features and the semantic features of the task properties. The second feature vector of the client to be assigned is extracted from the second feature vector of the task description of the task to be calculated of each client, and is further used as a query vector to multiply it by the second feature matrix to obtain a third feature vector of the client to be assigned that integrates the topological features and the semantic features of the task properties.

[0053] Here, for the first feature vector used to express resource features , and the third feature vector for expressing the associated semantic features of the task nature , calculate the node-to-node response criterion factor, expressed as:

[0054]

[0055] in and denote vector dot product and dot addition respectively, and It means that the square of each position of the vector is calculated. is a unit vector.

[0056] In this way, the fourth eigenvector can be obtained, and then the probability value of the fourth eigenvector belonging to the allocation of cloud GPU virtual resources is obtained through a Softmax-like classification function. Further, the computing GPU virtual resources are allocated to the client to be allocated based on the product between the probability value and the available amount of cloud GPU virtual resources at the current time point.

[0057] In this way, on the one hand, through the feature association fusion between the first feature matrix based on semantic association and the second feature matrix based on topological association, and the query-based retrieval of the feature expression of a single client in the fused association feature space, an appropriate coding expression of the task nature of a single client can be obtained.

[0058] On the other hand, by regarding each position of the first eigenvector and the third eigenvector as a node, the responsiveness characteristic component used to represent the mapping relationship between non-correlated feature distributions can be obtained based on the beta-like distribution analysis between nodes, which can be used as a practical factor for evaluating the responsiveness behavior between node distributions. This factor can improve the degree of confidence mapping estimation between node distributions, thereby obtaining an appropriate expression of the responsiveness between high-dimensional feature distributions.

[0059] Based on this, the present application proposes a scheduling method based on GPU virtualization in a cloud computing environment, which includes: obtaining the available amount of cloud GPU virtual resources at multiple predetermined time points including the current time point; passing the available amount of cloud GPU virtual resources at the multiple predetermined time points through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer to obtain a first feature vector; obtaining task descriptions of tasks to be calculated for multiple clients; passing the task descriptions of the tasks to be calculated for each of the clients through a context encoder including an embedding layer to obtain multiple feature vectors, and cascading the multiple feature vectors to obtain a second feature vector corresponding to the task description of the tasks to be calculated for each of the clients; arranging the multiple second feature vectors of the multiple clients in two dimensions into a feature matrix and then passing them through a first convolutional neural network to obtain a first feature matrix; obtaining a topological matrix of the multiple clients, wherein the value of each position on the non-diagonal position in the topological matrix is ​​the distance between the corresponding two clients, and the value of each position on the diagonal position in the topological matrix is ​​zero; passing the topological matrix through a second convolutional neural network model to obtain a topological feature matrix; The invention relates to a method for implementing a plurality of tasks, comprising: performing matrix multiplication of a feature matrix with the first feature matrix to map the high-dimensional topological features of the topological feature matrix to the high-dimensional feature space of the first feature matrix to obtain a second feature matrix that integrates the topological features and the semantic features of the task properties; extracting the second feature vector of the client to be assigned from the second feature vector of the task description of the task to be calculated of each of the clients, and multiplying the second feature vector of the client to be assigned with the second feature matrix to obtain a third feature vector; calculating an inter-node response criterion factor between the first feature vector and the third feature vector to obtain a fourth feature vector, wherein the inter-node response criterion factor is related to the ratio between the position point addition between the first feature vector and the third feature vector and the position point product between the first feature vector and the third feature vector; applying a Softmax-like classification function to the fourth feature vector to obtain a probability value that the fourth feature vector belongs to the allocation of cloud GPU virtual resources; and allocating computing GPU virtual resources to the client to be assigned based on the product of the probability value and the available amount of cloud GPU virtual resources at the current time point.

[0060] After introducing the basic principles of the present application, various non-limiting embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0061] Exemplary Methods

[0062] Figure 1 The figure shows a flow chart of a scheduling method based on GPU virtualization in a cloud computing environment. Figure 1As shown, according to the scheduling method based on GPU virtualization in a cloud computing environment of an embodiment of the present application, the method includes: S110, obtaining the available amount of cloud GPU virtual resources at multiple predetermined time points including the current time point; S120, passing the available amount of cloud GPU virtual resources at the multiple predetermined time points through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer to obtain a first feature vector; S130, obtaining task descriptions of tasks to be calculated of multiple clients; S140, passing the task descriptions of the tasks to be calculated of each of the clients through a context encoder including an embedding layer to obtain multiple feature vectors, and cascading the multiple feature vectors to obtain a second feature vector corresponding to the task description of the tasks to be calculated of each of the clients; S150, arranging the multiple second feature vectors of the multiple clients in two dimensions into a feature matrix and passing it through a first convolutional neural network to obtain a first feature matrix; S160, obtaining a topological matrix of the multiple clients, wherein the value of each position at the non-diagonal position in the topological matrix is ​​the distance between the corresponding two clients, and the value of each position at the diagonal position in the topological matrix is ​​zero; S170, passing the topological matrix through a second convolutional neural network model to obtain a topological feature matrix; S1 80, matrix multiplying the topological feature matrix with the first feature matrix to map the high-dimensional topological features of the topological feature matrix into the high-dimensional feature space of the first feature matrix to obtain a second feature matrix that integrates the topological features and the semantic features of the task properties; S190, extracting the second feature vector of the client to be assigned from the second feature vector of the task description of the task to be calculated of each of the clients, and multiplying the second feature vector of the client to be assigned with the second feature matrix to obtain a third feature vector; S200, calculating the node-to-node response criterion factor between the first feature vector and the third feature vector to obtain a fourth feature vector, wherein the node-to-node response criterion factor is related to the ratio between the position point addition between the first feature vector and the third feature vector and the position point product between the first feature vector and the third feature vector; S210, passing the fourth feature vector through a Softmax-like classification function to obtain a probability value that the fourth feature vector belongs to the allocation of cloud GPU virtual resources; and, S220, allocating computing GPU virtual resources to the client to be assigned based on the product between the probability value and the available amount of cloud GPU virtual resources at the current time point.

[0063] Figure 2 The diagram shows a schematic diagram of the architecture of a scheduling method based on GPU virtualization in a cloud computing environment according to an embodiment of the present application. Figure 2 As shown, in the network architecture of the scheduling method based on GPU virtualization in the cloud computing environment, first, the available amounts of cloud GPU virtual resources at the multiple predetermined time points (for example, Figure 2P1 as shown in Figure 1) is passed through a temporal encoder consisting of a one-dimensional convolutional layer and a fully connected layer (e.g., Figure 2 E1 as shown in the figure) to obtain the first eigenvector (e.g., Figure 2 Then, the task description of the to-be-computed task of each client is obtained (for example, Figure 2 P2 as shown in ) is passed through a context encoder containing an embedding layer (e.g., Figure 2 E2 as shown in the figure) to obtain multiple feature vectors (for example, Figure 2 ), concatenate the plurality of feature vectors to obtain a second feature vector (eg, as shown in VF2) corresponding to the task description of the to-be-computed task of each client. Figure 2 Then, the plurality of second feature vectors of the plurality of clients are arranged in two dimensions into a feature matrix (for example, as Figure 2 ) and then passed through a first convolutional neural network (e.g., Figure 2 ) to obtain the first feature matrix (e.g., Figure 2 MF1 shown in FIG); then, the obtained topological matrix (for example, Figure 2 M1 as shown in the figure) through a second convolutional neural network model (e.g., Figure 2 ) to obtain a topological feature matrix (e.g., Figure 2 MF2 as shown in the figure); then, the topological feature matrix is ​​matrix-multiplied with the first feature matrix to map the high-dimensional topological features of the topological feature matrix to the high-dimensional feature space of the first feature matrix to obtain a second feature matrix (for example, as Figure 2 MF3 as shown in the figure); then, extracting the second feature vector of the client to be assigned from the second feature vector of the task description of the task to be calculated of each client (for example, Figure 2 VF4 as shown in FIG. 4 ), multiplying the second feature vector of the client to be assigned by the second feature matrix to obtain a third feature vector (for example, Figure 2 Then, the node-to-node response criterion factor between the first eigenvector and the third eigenvector is calculated to obtain a fourth eigenvector (for example, Figure 2 Then, the fourth feature vector is classified by a Softmax-like function (for example, Figure 2 The circle S shown in FIG. 1 is used to obtain the probability value of the fourth eigenvector belonging to the allocated cloud GPU virtual resource (for example, Figure 2Q as shown in ); and, finally, allocating computing GPU virtual resources to the client to be allocated based on the product of the probability value and the available amount of cloud-based GPU virtual resources at the current time point.

[0064] In step S110 and step S120, the available amount of cloud GPU virtual resources at multiple predetermined time points including the current time point is obtained, and the available amount of cloud GPU virtual resources at the multiple predetermined time points is passed through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer to obtain a first feature vector. As mentioned above, it should be understood that in the technical solution of the present application, the resource allocation of tasks can essentially be regarded as a problem of resource response to the nature of the task. When considering resource scheduling, it is necessary not only to consider the existing available resource amount (that is, how many GPU virtual resources are available), but also to consider the demand distribution characteristics of GPU computing resources for the pending tasks of each user among multiple users.

[0065] Therefore, in the technical solution of the present application, first, the available amount of cloud GPU virtual resources at multiple predetermined time points including the current time point is obtained from the cloud GPU. Then, the available amount of cloud GPU virtual resources at the multiple predetermined time points is encoded in a timing encoder to extract high-dimensional correlation feature information of the available amount of cloud GPU virtual resources at the multiple predetermined time points in the time dimension and the data dimension, so that not only can the influence of data drift along the timing direction be eliminated through the correlation information, but also the original data can be replaced by the extracted high-dimensional features that reflect the correlation information between the input data for calculation, which can eliminate the influence of the error of the original data in the data dimension, thereby obtaining the first feature vector.

[0066] Specifically, in an embodiment of the present application, the process of obtaining a first feature vector by passing the available amounts of the cloud GPU virtual resources at the plurality of predetermined time points through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer includes: first, arranging the available amounts of the cloud GPU virtual resources at the plurality of predetermined time points into a one-dimensional input vector corresponding to each of the cloud GPU virtual resources according to the time dimension. Then, the fully connected layer of the temporal encoder is used to perform full connection encoding on the input vector using the following formula to extract the high-dimensional implicit features of the feature values ​​at each position in the input vector, wherein the formula is: ,in is the input vector, is the output vector, is the weight matrix, is the bias vector, Represents matrix multiplication. Finally, the one-dimensional convolution layer of the temporal encoder is used to perform one-dimensional convolution encoding on the input vector using the following formula to extract the high-dimensional implicit correlation feature of the correlation between the eigenvalues ​​at each position in the input vector, wherein the formula is:

[0067]

[0068] in, a The convolution kernel is x Width in direction, F is the convolution kernel parameter vector, G is the local vector matrix that operates with the convolution kernel function, is the size of the convolution kernel.

[0069] In step S130 and step S140, the task descriptions of the tasks to be calculated of the multiple clients are obtained, and the task descriptions of the tasks to be calculated of each of the clients are passed through a context encoder including an embedding layer to obtain multiple feature vectors, and the multiple feature vectors are cascaded to obtain a second feature vector corresponding to the task descriptions of the tasks to be calculated of each of the clients. It should be understood that in order to more accurately and reasonably allocate computing GPU virtual resources to the clients to be allocated, when considering resource scheduling, it is necessary not only to consider the existing available resource amount, that is, how many GPU virtual resources are available, but also to consider the demand distribution characteristics of the GPU computing resources for the tasks to be processed of each of the multiple users.

[0070] Therefore, in the technical solution of the present application, it is further necessary to obtain the task descriptions of the tasks to be calculated of multiple clients from the cloud storage end, and pass the task descriptions of the multiple tasks to be calculated through a context encoder including an embedding layer to obtain multiple feature vectors with global task description associated feature information. In this way, the obtained multiple feature vectors can be cascaded to obtain a second feature vector corresponding to each task to be calculated.

[0071] Specifically, in an embodiment of the present application, the process of obtaining multiple feature vectors by passing the task description of the tasks to be calculated of each of the clients through a context encoder including an embedding layer includes: first, performing word segmentation processing on the task description of the tasks to be calculated of each of the clients to convert the task description of the tasks to be calculated of each of the clients into a word sequence composed of multiple words, so as to avoid semantic confusion in the subsequent encoding process. Then, using the embedding layer of the context encoder, each word in the word sequence is mapped to a word vector to obtain a sequence of word vectors. Finally, using the converter of the context encoder, the sequence of word vectors is semantically encoded based on global context to obtain the multiple feature vectors. It should be understood that since the converter-based encoder model can perform global feature encoding on the word vector based on the context, the multiple feature vectors obtained have global task description associated feature information.

[0072] In step S150, the multiple second feature vectors of the multiple clients are arranged in two dimensions into a feature matrix and then passed through the first convolutional neural network to obtain a first feature matrix. That is, in the technical solution of the present application, after obtaining the second feature vectors corresponding to each task to be calculated, the second feature vectors can be further arranged in two dimensions into a feature matrix to fuse the feature information of each task to be calculated. Then, the feature matrix is ​​processed through the first convolutional neural network to extract the high-dimensional correlation features of each position in the feature matrix, thereby obtaining the first feature matrix. Correspondingly, in a specific example, each layer of the first convolutional neural network performs convolution processing based on a two-dimensional convolution kernel, pooling processing along the channel dimension and activation processing on the input data in the forward pass of the layer to output the first feature matrix by the last layer of the first convolutional neural network, wherein the input of the first layer of the first convolutional neural network is the feature matrix.

[0073] In step S160 and step S170, the topological matrix of the multiple clients is obtained, the value of each position on the non-diagonal position in the topological matrix is ​​the distance between the corresponding two clients, the value of each position on the diagonal position in the topological matrix is ​​zero, and the topological matrix is ​​passed through the second convolutional neural network model to obtain a topological feature matrix. That is, in the technical solution of the present application, in order to more accurately and reasonably allocate computing GPU virtual resources to the clients to be allocated, it is also necessary to obtain the topological matrix of the multiple clients. In particular, here, the value of each position on the non-diagonal position in the topological matrix is ​​the distance between the corresponding two clients, and the value of each position on the diagonal position in the topological matrix is ​​zero. Then, the topological matrix is ​​further processed through the second convolutional neural network model to extract the topological structure feature information of the multiple clients, thereby obtaining a topological feature matrix.

[0074] In step S180, the topological feature matrix is ​​matrix-multiplied with the first feature matrix to map the high-dimensional topological features of the topological feature matrix to the high-dimensional feature space of the first feature matrix to obtain a second feature matrix that integrates the topological features and the semantic features of the task properties. It should be understood that the resource allocation of the task can essentially be regarded as a problem of the response of resources to the task properties. Therefore, in the technical solution of the present application, it is expected that the resources and task properties can be properly encoded and expressed in a high-dimensional space to improve the accuracy of allocating computing GPU virtual resources to the client to be allocated.

[0075] Therefore, in the technical solution of the present application, after obtaining the semantic features of the topological features and the nature of the task, the topological feature matrix is ​​further matrix-multiplied with the first feature matrix to map the high-dimensional topological features of the topological feature matrix to the high-dimensional feature space of the first feature matrix, thereby obtaining a second feature matrix that integrates the semantic features of the topological features and the nature of the task. It should be understood that through the feature association fusion between the first feature matrix based on semantic association and the topological feature matrix based on topological association, and the query-based retrieval of the feature expression of the single client in the fused association feature space, the appropriate coding expression of the nature of the task of the single client can be obtained.

[0076] In step S190, the second feature vector of the client to be assigned is extracted from the second feature vector of the task description of the task to be calculated of each of the clients, and the second feature vector of the client to be assigned is multiplied with the second feature matrix to obtain a third feature vector. That is, in the technical solution of the present application, it is further necessary to extract the task description features of the client to be assigned from the task description features of the task to be calculated of each of the clients, so as to obtain the second feature vector. In this way, the second feature vector having the task description features of the client to be assigned can be further matrix-multiplied with the second feature matrix having the semantic features of the fused topological features and task properties, so as to obtain the third feature vector for expressing the associated semantic features of the task properties. .

[0077] In step S200, the node-to-node response criterion factor between the first eigenvector and the third eigenvector is calculated to obtain a fourth eigenvector, and the node-to-node response criterion factor is related to the ratio between the positional point addition between the first eigenvector and the third eigenvector and the positional point product between the first eigenvector and the third eigenvector. It should be understood that the resource allocation of the task can essentially be regarded as a problem of resource response to the nature of the task. Therefore, on the one hand, if the resource and task nature can be properly encoded and expressed in a high-dimensional space, and on the other hand, the responsiveness between the feature distributions of the encoded feature expressions of the two can be properly characterized, the high-dimensional representation of the responsiveness can be used for classification to obtain a probabilistic representation of the allocation.

[0078] Therefore, in the technical solution of the present application, the first feature vector used for resource feature expression is further , and the third feature vector for the associated semantic feature expression of the task nature , calculate the node-to-node response criterion factor between the first eigenvector and the third eigenvector to obtain a fourth eigenvector. It should be understood that the positions of the first eigenvector and the third eigenvector are respectively regarded as nodes, and the responsiveness characteristic component used to represent the mapping relationship between non-correlated feature distributions can be obtained based on the beta-like distribution analysis between the nodes, as a practical factor for evaluating the responsiveness behavior between the node distributions, and the factor can improve the confidence mapping estimation degree between the node distributions, thereby obtaining an appropriate expression of the responsiveness between high-dimensional feature distributions.

[0079] Specifically, in the embodiment of the present application, the process of calculating the inter-node response criterion factor between the first eigenvector and the third eigenvector to obtain the fourth eigenvector includes: calculating the inter-node response criterion factor between the first eigenvector and the third eigenvector using the following formula to obtain the fourth eigenvector;

[0080] Wherein, the formula is:

[0081]

[0082] in is the first eigenvector, is the third eigenvector, and denote vector dot product and dot addition respectively, and It means that the square of each position of the vector is calculated. is a unit vector.

[0083] In step S210 and step S220, the fourth eigenvector is passed through a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resources, and the computing GPU virtual resources are allocated to the client to be allocated based on the product between the probability value and the available amount of the cloud GPU virtual resources at the current time point. That is, in the technical solution of the present application, after obtaining an appropriate expression of the responsiveness between high-dimensional feature distributions, the high-dimensional representation of the responsiveness is further used for classification to obtain a probabilistic representation of the allocation. Specifically, the fourth eigenvector is further passed through a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resources. In this way, the computing GPU virtual resources can be allocated to the client to be allocated based on the product between the probability value and the available amount of the cloud GPU virtual resources at the current time point.

[0084] Specifically, in an embodiment of the present application, the process of passing the fourth eigenvector through a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resources includes: passing the fourth eigenvector through a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resources by calculating the following formula; wherein the formula is: exp(-xi) / ∑ i exp(-xi).

[0085] In summary, the scheduling method based on GPU virtualization in the cloud computing environment of the embodiment of the present application is explained. On the one hand, through the feature association fusion between the first feature matrix based on semantic association and the second feature matrix based on topological association, and the query-based retrieval of the feature expression of a single client in the fused association feature space, the appropriate coding expression of the task nature of the single client can be obtained. On the other hand, the positions of the first feature vector used for resource feature expression and the third feature vector used for the associated semantic feature expression of the task nature are respectively regarded as nodes, and the responsiveness feature component used to represent the mapping relationship between non-associated feature distributions can be obtained based on the beta-like distribution analysis between nodes, as a practical factor for evaluating the responsiveness behavior between node distributions, and the factor can improve the confidence mapping estimation degree between node distributions, thereby obtaining an appropriate expression of responsiveness between high-dimensional feature distributions. Furthermore, it is possible to more accurately and reasonably treat the allocated client allocation computing GPU virtual resources.

[0086] Exemplary Systems

[0087] Figure 3 The block diagram of the scheduling system based on GPU virtualization in the cloud computing environment according to the embodiment of the present application is shown. Figure 3As shown, according to the scheduling system 400 based on GPU virtualization in a cloud computing environment of an embodiment of the present application, it includes: a resource availability acquisition unit 410, which is used to acquire the availability of cloud-based GPU virtual resources at multiple predetermined time points including the current time point; a timing encoding unit 420, which is used to obtain the availability of cloud-based GPU virtual resources at the multiple predetermined time points obtained by the resource availability acquisition unit 410 through a timing encoder including a one-dimensional convolutional layer and a fully connected layer to obtain a first feature vector; a task description acquisition unit 430, which is used to acquire task descriptions of tasks to be calculated for multiple clients; a context encoding unit 440, which is used to encode the available amount of the cloud-based GPU virtual resources at the multiple predetermined time points obtained by the resource availability acquisition unit 410 through a timing encoder including a one-dimensional convolutional layer and a fully connected layer to obtain a first feature vector; a task description acquisition unit 430, which is used to acquire task descriptions of tasks to be calculated for multiple clients; and a context encoding unit 440, which is used to encode the available amount of the cloud-based GPU virtual resources at the multiple predetermined time points obtained by the task description acquisition unit 430. The task description of the task to be calculated of the client is obtained by a context encoder including an embedding layer to obtain multiple feature vectors, and the multiple feature vectors are cascaded to obtain a second feature vector corresponding to the task description of the task to be calculated of each client; a first convolution unit 450 is used to arrange the multiple second feature vectors of the multiple clients obtained by the context encoding unit 440 into a feature matrix in two dimensions and then pass them through a first convolutional neural network to obtain a first feature matrix; a topology matrix acquisition unit 460 is used to obtain a topology matrix of the multiple clients, the value of each position on the non-diagonal position in the topology matrix is ​​the distance between the corresponding two clients, and the value of each position on the diagonal position in the topology matrix is ​​the distance between the corresponding two clients. The value of the position is zero; a second convolution unit 470 is used to pass the topology matrix obtained by the topology matrix acquisition unit 460 through a second convolution neural network model to obtain a topology feature matrix; a mapping unit 480 is used to perform matrix multiplication on the topology feature matrix obtained by the second convolution unit 470 and the first feature matrix obtained by the first convolution unit 450 to map the high-dimensional topological features of the topology feature matrix to the high-dimensional feature space of the first feature matrix to obtain a second feature matrix that integrates the topological features and the semantic features of the task properties; a third feature vector generation unit 490 is used to obtain the task vectors of the client to be calculated from each of the context encoding units 440. a response criterion factor calculation unit 500, configured to calculate an inter-node response criterion factor between the first eigenvector obtained by the timing coding unit 420 and the third eigenvector obtained by the third eigenvector generating unit 490 to obtain a fourth eigenvector, wherein the inter-node response criterion factor is related to a ratio between a position point addition between the first eigenvector and the third eigenvector and a position point product between the first eigenvector and the third eigenvector;The probability value calculation unit 510 is used to obtain the probability value of the fourth eigenvector obtained by the response criterion factor calculation unit 500 through a Softmax classification function; and the allocation unit 520 is used to allocate the computing GPU virtual resource to the client to be allocated based on the product of the probability value obtained by the probability value calculation unit 510 and the available amount of the cloud GPU virtual resource at the current time point.

[0088] In one example, in the scheduling system 400 based on GPU virtualization in the cloud computing environment, the timing encoding unit 420 is further used to: arrange the available amounts of the cloud GPU virtual resources at the multiple predetermined time points according to the time dimension into a one-dimensional input vector corresponding to each of the cloud GPU virtual resources; and use the fully connected layer of the timing encoder to fully connect the input vector using the following formula to extract the high-dimensional implicit features of the feature values ​​at each position in the input vector, wherein the formula is: ,in is the input vector, is the output vector, is the weight matrix, is the bias vector, represents matrix multiplication; using a one-dimensional convolution layer of a temporal encoder to perform one-dimensional convolution encoding on the input vector according to the following formula to extract a high-dimensional implicit correlation feature of the correlation between the eigenvalues ​​at each position in the input vector, wherein the formula is:

[0089]

[0090] in, a The convolution kernel is x Width in direction, F is the convolution kernel parameter vector, G is the local vector matrix that operates with the convolution kernel function, is the size of the convolution kernel.

[0091] In one example, in the scheduling system 400 based on GPU virtualization in the above-mentioned cloud computing environment, the context encoding unit 440 is further used to: perform word segmentation processing on the task description of the task to be calculated of each of the clients to convert the task description of the task to be calculated of each of the clients into a word sequence composed of multiple words; use the embedding layer of the context encoder to map each word in the word sequence to a word vector to obtain a sequence of word vectors; and use the converter of the context encoder to perform global context semantic encoding on the sequence of word vectors to obtain the multiple feature vectors.

[0092] In one example, in the scheduling system 400 based on GPU virtualization in the above-mentioned cloud computing environment, the first convolution unit 450 is further used to: two-dimensionally arrange the multiple second feature vectors of the multiple clients to obtain a feature matrix; and each layer of the first convolutional neural network performs convolution processing based on a two-dimensional convolution kernel, pooling processing along the channel dimension, and activation processing on the input data in the forward pass of the layer to output the first feature matrix by the last layer of the first convolutional neural network, wherein the input of the first layer of the first convolutional neural network is the feature matrix.

[0093] In one example, in the scheduling system 400 based on GPU virtualization in the cloud computing environment, the response criterion factor calculation unit 500 is further used to calculate the inter-node response criterion factor between the first eigenvector and the third eigenvector using the following formula to obtain a fourth eigenvector; wherein the formula is:

[0094]

[0095] in is the first eigenvector, is the third eigenvector, and denote vector dot product and dot addition respectively, and It means that the square of each position of the vector is calculated. is a unit vector.

[0096] In one example, in the scheduling system 400 based on GPU virtualization in the above-mentioned cloud computing environment, the probability value calculation unit 510 is further used to: pass the fourth eigenvector through a Softmax classification function to obtain the probability value of the fourth eigenvector belonging to the allocated cloud GPU virtual resource by the following formula; wherein the formula is: exp(-xi) / ∑ i exp(-xi).

[0097] Here, those skilled in the art can understand that the specific functions and operations of the various units and modules in the scheduling system 400 based on GPU virtualization in the above cloud computing environment have been described in the above reference. Figure 1 to Figure 2 The scheduling method based on GPU virtualization in a cloud computing environment has been described in detail, and therefore, its repeated description will be omitted.

[0098] As described above, the scheduling system 400 based on GPU virtualization in the cloud computing environment according to the embodiment of the present application can be implemented in various terminal devices, such as a server of a scheduling algorithm based on GPU virtualization in the cloud computing environment. In one example, the scheduling system 400 based on GPU virtualization in the cloud computing environment according to the embodiment of the present application can be integrated into the terminal device as a software module and / or a hardware module. For example, the scheduling system 400 based on GPU virtualization in the cloud computing environment can be a software module in the operating system of the terminal device, or can be an application developed for the terminal device; of course, the scheduling system 400 based on GPU virtualization in the cloud computing environment can also be one of the many hardware modules of the terminal device.

[0099] Alternatively, in another example, the scheduling system 400 based on GPU virtualization in the cloud computing environment and the terminal device may also be separate devices, and the scheduling system 400 based on GPU virtualization in the cloud computing environment may be connected to the terminal device via a wired and / or wireless network, and transmit interactive information in accordance with an agreed data format.

[0100] Exemplary Electronic Devices

[0101] Below, reference Figure 4 To describe the electronic device according to the embodiment of the present application. Figure 4 As shown, the electronic device 10 includes one or more processors 11 and a memory 12. The processor 11 may be a central processing unit (CPU) or other processing units with data processing capability and / or instruction execution capability, and may control other components in the electronic device 10 to perform desired functions.

[0102] The memory 12 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may run the program instructions to implement the functions of the scheduling method based on GPU virtualization in the cloud computing environment of each embodiment of the present application described above and / or other desired functions. Various contents such as a topological feature matrix, a third eigenvector, etc. may also be stored in the computer-readable storage medium.

[0103] In one example, the electronic device 10 may further include: an input system 13 and an output system 14 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0104] The input system 13 may include, for example, a keyboard, a mouse, and the like.

[0105] The output system 14 can output various information to the outside, including the probability value of allocating cloud GPU virtual resources, the allocated computing GPU virtual resources, etc. The output system 14 can include, for example, a display, a speaker, a printer, a communication network and its connected remote output devices, etc.

[0106] Of course, to simplify, Figure 4 Only some of the components related to the present application in the electronic device 10 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application situations, the electronic device 10 may also include any other appropriate components.

[0107] Exemplary computer program products and computer-readable storage media

[0108] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps in the functions of the scheduling method based on GPU virtualization in a cloud computing environment according to various embodiments of the present application described in the above "Exemplary Method" section of this specification.

[0109] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0110] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, wherein when the computer program instructions are executed by a processor, the processor executes the steps of the scheduling method based on GPU virtualization in a cloud computing environment described in the above “Exemplary Method” section of this specification.

[0111] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, system or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0112] The basic principles of the present application are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present application. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, not for limitation, and the above details do not limit the present application to being implemented by adopting the above specific details.

[0113] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.

[0114] It should also be noted that in the apparatus, device and method of the present application, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0115] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

[0116] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

Claims

1. A scheduling method based on GPU virtualization in a cloud computing environment, characterized in that: include: Obtain the available amount of cloud GPU virtual resources at multiple predetermined time points including the current time point; Passing the available amounts of cloud GPU virtual resources at the plurality of predetermined time points through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer to obtain a first feature vector; Get the task descriptions of the pending computing tasks of multiple clients; Passing the task description of the to-be-computed task of each of the clients through a context encoder including an embedding layer to obtain a plurality of feature vectors, and concatenating the plurality of feature vectors to obtain a second feature vector corresponding to the task description of the to-be-computed task of each of the clients; Arranging the plurality of second feature vectors of the plurality of clients in two dimensions into a feature matrix and then passing the matrix through a first convolutional neural network to obtain a first feature matrix; Acquire a topology matrix of the multiple clients, wherein the value of each non-diagonal position in the topology matrix is ​​the distance between two corresponding clients, and the value of each diagonal position in the topology matrix is ​​zero; Passing the topological matrix through a second convolutional neural network model to obtain a topological feature matrix; Performing matrix multiplication of the topological feature matrix and the first feature matrix to map the high-dimensional topological features of the topological feature matrix into the high-dimensional feature space of the first feature matrix to obtain a second feature matrix that integrates the topological features and the semantic features of the task properties; Extracting the second eigenvector of the client to be assigned from the second eigenvector of the task description of the task to be calculated of each client, and multiplying the second eigenvector of the client to be assigned with the second eigenmatrix to obtain a third eigenvector; calculating an inter-nodal response criterion factor between the first eigenvector and the third eigenvector to obtain a fourth eigenvector, the inter-nodal response criterion factor being related to a ratio between a position-wise addition between the first eigenvector and the third eigenvector and a position-wise product of the first eigenvector and the third eigenvector; The fourth eigenvector is subjected to a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resource; and Allocating computing GPU virtual resources to the client to be allocated based on the product of the probability value and the available amount of the cloud GPU virtual resources at the current time point; The step of calculating the inter-node response criterion factor between the first eigenvector and the third eigenvector to obtain a fourth eigenvector comprises: Calculating an inter-node response criterion factor between the first eigenvector and the third eigenvector using the following formula to obtain a fourth eigenvector; Wherein, the formula is: in is the first eigenvector, is the third eigenvector, and denote vector dot product and dot addition respectively, and The square of each position of the vector is calculated. is a unit vector.

2. The scheduling method based on GPU virtualization in a cloud computing environment according to claim 1, wherein: The available amounts of the cloud GPU virtual resources at the plurality of predetermined time points are passed through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer to obtain a first feature vector, including: Arranging the available amounts of the cloud GPU virtual resources at the plurality of predetermined time points into a one-dimensional input vector corresponding to each of the cloud GPU virtual resources according to the time dimension; The input vector is fully connected encoded using the fully connected layer of the temporal encoder according to the following formula to extract high-dimensional implicit features of the feature values ​​at each position in the input vector, wherein the formula is: ,in is the input vector, is the output vector, is the weight matrix, is the bias vector, Represents matrix multiplication; The one-dimensional convolution layer of the temporal encoder is used to perform one-dimensional convolution encoding on the input vector according to the following formula to extract the high-dimensional implicit correlation feature of the correlation between the feature values ​​at each position in the input vector, wherein the formula is: in, a The convolution kernel is x Width in direction, F is the convolution kernel parameter vector, G is the local vector matrix that operates with the convolution kernel function, is the size of the convolution kernel.

3. The scheduling method based on GPU virtualization in a cloud computing environment according to claim 2, wherein: The task description of each of the tasks to be calculated of the client is passed through a context encoder including an embedding layer to obtain multiple feature vectors, including: Performing word segmentation processing on the task description of the to-be-computed task of each of the clients to convert the task description of the to-be-computed task of each of the clients into a word sequence consisting of a plurality of words; Mapping each word in the word sequence to a word vector using the embedding layer of the context encoder to obtain a sequence of word vectors; and The converter of the context encoder is used to perform global context semantic encoding on the sequence of word vectors to obtain the multiple feature vectors.

4. The scheduling method based on GPU virtualization in a cloud computing environment according to claim 3, wherein: Arranging the plurality of second feature vectors of the plurality of clients in two dimensions into a feature matrix and then passing the matrix through a first convolutional neural network to obtain a first feature matrix includes: Arranging the plurality of second feature vectors of the plurality of clients in two dimensions to obtain a feature matrix; and In the forward pass of the layer, each layer of the first convolutional neural network performs convolution processing based on a two-dimensional convolution kernel, pooling processing along the channel dimension, and activation processing on the input data so that the last layer of the first convolutional neural network outputs the first feature matrix, wherein the input of the first layer of the first convolutional neural network is the feature matrix.

5. The scheduling method based on GPU virtualization in a cloud computing environment according to claim 4, wherein: The fourth eigenvector is subjected to a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resource, including: The fourth eigenvector is passed through a Softmax classification function to obtain the probability value of the fourth eigenvector belonging to the allocated cloud GPU virtual resource by the following formula; wherein the formula is: exp(-xi) / ∑ i exp(-xi).

6. A scheduling system based on GPU virtualization in a cloud computing environment, characterized in that: include: A resource availability acquisition unit, used to acquire the availability of cloud GPU virtual resources at multiple predetermined time points including the current time point; a temporal encoding unit, configured to obtain a first feature vector by passing the available amounts of the cloud GPU virtual resources at the plurality of predetermined time points obtained by the resource available amount acquisition unit through a temporal encoder including a one-dimensional convolutional layer and a fully connected layer; A task description acquisition unit, used to acquire task descriptions of tasks to be calculated from multiple clients; a context encoding unit, configured to obtain a plurality of feature vectors by passing the task descriptions of the tasks to be calculated of the clients obtained by the task description obtaining units through a context encoder including an embedding layer, and to concatenate the plurality of feature vectors to obtain a second feature vector corresponding to the task descriptions of the tasks to be calculated of the clients; A first convolution unit, configured to two-dimensionally arrange the plurality of second feature vectors of the plurality of clients obtained by the context encoding unit into a feature matrix and then obtain a first feature matrix through a first convolutional neural network; A topology matrix acquisition unit, used to acquire a topology matrix of the plurality of clients, wherein the value of each non-diagonal position in the topology matrix is ​​the distance between two corresponding clients, and the value of each diagonal position in the topology matrix is ​​zero; A second convolution unit, used for passing the topological matrix obtained by the topological matrix acquisition unit through a second convolutional neural network model to obtain a topological feature matrix; A mapping unit, configured to perform matrix multiplication on the topological feature matrix obtained by the second convolution unit and the first feature matrix obtained by the first convolution unit to map the high-dimensional topological features of the topological feature matrix into the high-dimensional feature space of the first feature matrix to obtain a second feature matrix that integrates the topological features and the semantic features of the task properties; A third feature vector generating unit is used to extract the second feature vector of the client to be assigned from the second feature vector of the task description of the task to be calculated of the client obtained by each of the context encoding units, and multiply the second feature vector of the client to be assigned by the second feature matrix obtained by the mapping unit to obtain a third feature vector; a response criterion factor calculation unit, configured to calculate an inter-node response criterion factor between the first eigenvector obtained by the temporal encoding unit and the third eigenvector obtained by the third eigenvector generation unit to obtain a fourth eigenvector, wherein the inter-node response criterion factor is related to a ratio between a positional point addition between the first eigenvector and the third eigenvector and a positional point product between the first eigenvector and the third eigenvector; A probability value calculation unit, configured to pass the fourth eigenvector obtained by the response criterion factor calculation unit through a Softmax-like classification function to obtain a probability value that the fourth eigenvector belongs to the allocated cloud GPU virtual resource; as well as an allocating unit, configured to allocate the computing GPU virtual resource to the client to be allocated based on the product of the probability value obtained by the probability value calculating unit and the available amount of the cloud GPU virtual resource at the current time point; Wherein, the response criterion factor calculation unit is used to: Calculating an inter-node response criterion factor between the first eigenvector and the third eigenvector using the following formula to obtain a fourth eigenvector; Wherein, the formula is: in is the first eigenvector, is the third eigenvector, and denote vector dot product and dot addition respectively, and The square of each position of the vector is calculated. is a unit vector.

7. The scheduling system based on GPU virtualization in a cloud computing environment according to claim 6, wherein: The timing coding unit is further used for: Arranging the available amounts of the cloud GPU virtual resources at the plurality of predetermined time points into a one-dimensional input vector corresponding to each of the cloud GPU virtual resources according to the time dimension; The input vector is fully connected encoded using the fully connected layer of the temporal encoder according to the following formula to extract high-dimensional implicit features of the feature values ​​at each position in the input vector, wherein the formula is: ,in is the input vector, is the output vector, is the weight matrix, is the bias vector, Represents matrix multiplication; The one-dimensional convolution layer of the temporal encoder is used to perform one-dimensional convolution encoding on the input vector according to the following formula to extract the high-dimensional implicit correlation feature of the correlation between the feature values ​​at each position in the input vector, wherein the formula is: in, a The convolution kernel is x Width in direction, F is the convolution kernel parameter vector, G is the local vector matrix that operates with the convolution kernel function, is the size of the convolution kernel.

8. The scheduling system based on GPU virtualization in a cloud computing environment according to claim 6, wherein: The context encoding unit is further used for: Performing word segmentation processing on the task description of the to-be-computed task of each of the clients to convert the task description of the to-be-computed task of each of the clients into a word sequence consisting of a plurality of words; Using the embedding layer of the context encoder, each word in the word sequence is mapped to a word vector to obtain a sequence of word vectors; and using the converter of the context encoder, the sequence of word vectors is encoded based on global context semantics to obtain the multiple feature vectors.

9. An electronic device, comprising: processor; as well as A memory, in which computer program instructions are stored, and when the computer program instructions are executed by the processor, the processor executes the scheduling method based on GPU virtualization in a cloud computing environment according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • NFV resources scheduling method, device and system

    CN113535399A

  • Micro-service resource management method and system based on dynamic routing and electronic equipment

    CN113778718A