Computing power scheduling method and system combining GPU virtualization and AI
By identifying and analyzing the business volume and scheduling time of the computing power demander and supplier, combining GPU virtualization technology and time slice rotation algorithm, optimizing computing power scheduling has solved the problem of computing power scheduling in the existing technology that cannot dynamically deal with business changes and hunger, and achieve more efficient and balanced resource allocation.
Patent Information
- Application Number
- CN202410648997.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-05-23
AI Technical Summary
The prior art is difficult to dynamically respond to changes on the business side in computing power scheduling, which leads to inability to effectively follow the dynamic changes on the business side, and there is a problem of non-emergency and short-term services being hungry due to the long-term failure of computing power scheduling.
By identifying the computing power demander, computing power supplier and computing power intermediary, detecting historical business volume and business volume moments, analyzing and planning business volume, setting computing power scheduling priority, identifying the business-time mapping relationship between historical business volume and historical scheduling time, using the planning scheduling time and GPU virtualization technology to allocate computing power scheduling time slices, combining the time slice rotation algorithm and the optimized computing power scheduling algorithm to complete computing power scheduling power scheduling.
It reduces the hunger problem of computing power scheduling, improves the ability to respond to dynamic changes on the business side, and ensures balanced processing and resource allocation of each business volume.
Smart Images

Figure CN118672767B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network technology, and in particular to a computing power scheduling method and system combining GPU virtualization and AI. Background Art
[0002] GPU virtualization technology refers to graphics card virtualization technology. Graphics card virtualization technology is the process of slicing graphics cards and allocating these graphics card time slices to virtual machines. Since graphics cards that support graphics card virtualization can generally be divided into time slices of different specifications as needed, they can be allocated to multiple virtual machines. Artificial intelligence (AI) is a technology that simulates human intelligent thinking. It can realize human cognition and thinking activities. Through this technology, computers can simulate human thinking and intelligence, so as to complete many complex tasks, such as image recognition, speech recognition, natural language processing, decision making, etc. Computing power scheduling refers to the process of reasonably allocating and utilizing computing power resources in a certain area or system, by allocating CPU resources to processes.
[0003] At present, the existing technology for computing power scheduling is mostly set up in advance, that is, the computing power requirements of the business end are investigated in advance, and then which businesses need to be allocated which computing power. Some problems will arise in this process. First, when the business end changes dynamically, since the computing power scheduling criteria are set in advance and static, the computing power scheduling cannot follow the dynamic changes of the business end. Second, the existing computing power scheduling algorithm prioritizes computing power scheduling based on which businesses are urgent or which businesses have large volumes. This will cause non-urgent businesses and short-term businesses to starve due to long-term lack of computing power scheduling. Therefore, the hunger problem of computing power scheduling is difficult to solve. Summary of the invention
[0004] In order to solve the above problems, the present invention provides a computing power scheduling method and system combining GPU virtualization and AI, which can reduce the hunger problem of computing power scheduling.
[0005] In a first aspect, the present invention provides a computing power scheduling method combining GPU virtualization and AI, comprising:
[0006] Identify the computing power demander, computing power supplier and computing power intermediary when performing computing power scheduling, detect the historical business volume and business volume time of the computing power demander in a preset historical period in the computing power intermediary, and analyze the planned business volume of the computing power demander in a preset planning period based on the historical business volume and the business volume time;
[0007] According to the planned business volume, the computing power scheduling priority of the computing power demander is set, the historical scheduling duration of the computing power dispatched from the computing power supplier to the computing power demander within a preset historical period is queried, and the business-duration mapping relationship between the historical business volume and the historical scheduling duration is identified;
[0008] According to the business-duration mapping relationship and the planned business volume, the planned scheduling duration for scheduling computing power from the computing power supplier to the computing power demander within the planned period is analyzed, and the computing power scheduling time slice of the computing power demander is allocated by using the planned scheduling duration and the preset GPU virtualization technology;
[0009] Based on the computing power scheduling time slice, the computing power supplier is scheduled using a preset time slice rotation algorithm and the computing power scheduling priority to obtain a computing power scheduling record, and the computing power scheduling record is used to determine whether the computing power supplier meets the requirements of the computing power demander;
[0010] When the computing power supplier does not meet the computing power demander, calculating the computing power scheduling performance of the computing power supplier in the computing power scheduling record, and determining a computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander according to the computing power scheduling performance;
[0011] The computing power scheduling algorithm is used to complete the computing power scheduling from the computing power supplier to the computing power demander, and the computing power scheduling result of the computing power supplier is obtained.
[0012] In a possible implementation manner of the first aspect, analyzing the planned business volume of the computing power demander within a preset planning period based on the historical business volume and the business volume moment includes:
[0013] Based on the traffic moment, the historical traffic is converted into a traffic sequence using the following format:
[0014]
[0015] X=(x (1) , x (2) ,...x (j) ,...x (m) )
[0016]
[0017] X=(X1,X2,...X i , ..., X n ) T
[0018] Where X represents the business volume sequence, i represents the order number of the computing power demander, n represents the total number of computing power demanders, j represents the order number of the business volume moment, m represents the total number of business volume moments, and x (j) represents the historical business volume at the jth business moment, It represents the historical business volume of the first computing power demander at the jth business volume moment. It represents the historical business volume of the second computing power demander at the jth business volume moment, It represents the historical business volume of the i-th computing power demander at the j-th business volume moment, represents the historical business volume of the nth computing power demander at the jth business volume moment, x (1) represents the historical business volume at the first business volume moment, x (2) represents the historical business volume at the second business volume moment, x (m) represents the historical business volume at the mth business moment, X i represents the historical business volume of the i-th computing power demander, It represents the historical business volume of the i-th computing power demander at the first business volume moment, It represents the historical business volume of the i-th computing power demander at the second business volume moment, represents the historical business volume of the i-th computing power demander at the m-th business volume moment, X1 represents the historical business volume of the first computing power demander, X2 represents the historical business volume of the second computing power demander, and X n Indicates the historical business volume of the nth computing power demander;
[0019] Extracting features from the traffic sequence using a preset principal component analysis method to obtain extracted features;
[0020] Inputting the extracted features into a preset long short-term memory recurrent network;
[0021] In the long short-term memory recurrent network, according to the extracted features, the planned business volume of the computing power demander within the preset planning period is calculated using the following formula:
[0022] f t =σ(W f ·[h t-1 , x t ]+b f )
[0023] i t =σ(W i ·[h t-1 , x t ]+b i )
[0024] Ct =tanh(W C ·[h t-1 , x t ]+b C )
[0025] C t =f t ·C t-1 +i t ·C t
[0026] o t =σ(W o ·[h t-1 , x t ]+b o )
[0027] h t =o t tanh(C t )
[0028] Among them, h t represents the planned traffic volume, σ represents, x t represents the extracted features at historical moment t, h t-1 represents the hidden state of the previous moment of historical moment t, W f represents the weight matrix of the forget gate, b f represents the bias term of the forget gate, σ represents the sigmoid function, and f t represents the output of the forget gate at historical time t, W i represents the weight matrix of the input gate, b i represents the bias term of the input gate, i t Represents the output result of the input gate, W C represents the weight matrix of the memory gate, b C represents the bias term of the memory gate, C t represents the cell state output by the memory gate at historical time t, C t-1 represents the cell state at the previous moment of historical moment t, W o represents the weight matrix of the output gate, b o represents the bias term of the output gate, o t Represents the output result of the output gate.
[0029] In a possible implementation manner of the first aspect, extracting features from the traffic sequence using a preset principal component analysis method to obtain extracted features includes:
[0030] Calculating the covariance matrix of the traffic sequence using a preset principal component analysis method;
[0031] According to the covariance matrix, using the preset principal component analysis method to calculate the eigenvalue and eigenvector of the traffic sequence;
[0032] Selecting a non-zero eigenvalue from the eigenvalues;
[0033] Identifying a target feature vector corresponding to the non-zero feature value in the feature vector;
[0034] The target feature vector is used as the extracted feature.
[0035] In a possible implementation manner of the first aspect, setting the computing power scheduling priority of the computing power demander according to the planned business volume includes:
[0036] Randomly arranging the planned traffic to obtain a traffic queue;
[0037] Converting the traffic queue into a large root heap;
[0038] Querying the front-to-back order of the planned business volume in the big root heap according to a preset pre-order traversal algorithm;
[0039] Based on the sequence, the computing power scheduling priority of the planned business volume is set according to a preset first-come-first-served scheduling algorithm.
[0040] In a possible implementation manner of the first aspect, the identifying a service-duration mapping relationship between the historical service volume and the historical scheduling duration includes:
[0041] Arrange the historical business volumes in ascending order to obtain a business volume sequence;
[0042] Arranging the historical scheduling durations according to the business volume sequence to obtain a duration sequence;
[0043] Normalizing the business volume sequence by using a missing value interpolation method to obtain a normalized sequence;
[0044] Inputting the normalized sequence into a preset duration analysis model, so as to output a normalized duration corresponding to the normalized sequence through the duration analysis model;
[0045] Identifying a target duration in the normalized duration corresponding to the traffic sequence;
[0046] Calculating a duration loss value between the duration sequence and the target duration, so as to optimize the duration analysis model by using the duration loss value to obtain an optimized duration analysis model;
[0047] The optimized duration analysis model is used to identify the service-duration mapping relationship between the historical service volume and the historical scheduling duration.
[0048] In a possible implementation of the first aspect, the using the planned scheduling duration and the preset GPU virtualization technology to allocate the computing power scheduling time slice of the computing power demander includes:
[0049] Allocate the initial time slice of the computing power demander by using the GPU virtualization technology;
[0050] According to the initial time slice and the planned scheduling duration, the number of computing power scheduling times of the computing power demander is calculated using the following formula:
[0051]
[0052]
[0053] Among them, K represents the number of computing power scheduling, k i represents the number of computing power scheduling of the i-th computing power demander, t i It represents the planning and scheduling duration of the i-th computing power demander, t0 represents the initial time slice, Indicates the rounding up symbol;
[0054] Determine the computing power scheduling round of the planned scheduling duration by using the computing power scheduling times;
[0055] Using the initial time slice to divide the planned scheduling duration into divided durations to obtain divided durations;
[0056] Determine the computing power scheduling round number of the split time using the computing power scheduling round;
[0057] Obtaining the computing power scheduling priority of the computing power demander;
[0058] Constructing the priority of the computing power scheduling round number using the computing power scheduling priority;
[0059] Sorting the planned scheduling durations by using the computing power scheduling round number and the priority to obtain a planned duration sequence;
[0060] According to the planned duration sequence, the computing power waiting time of the computing power demander is calculated using the following formula:
[0061]
[0062]
[0063] …
[0064]
[0065] …
[0066]
[0067]
[0068]
[0069] Among them, T represents the computing power waiting time, T i represents the computing power waiting time of the i-th computing power demander, u represents the computing power scheduling round number in the planned duration sequence, U represents the total number of computing power scheduling round numbers in the planned duration sequence, v represents the priority in the planned duration sequence, V represents the total number of priorities in the planned duration sequence, i represents the sequence number of the computing power demander, n represents the total number of computing power demanders, t0 represents the initial time slice, It represents the computing power waiting time of the i-th computing power demander under the first computing power scheduling round number, It represents the computing power waiting time of the i-th computing power demander under the second computing power scheduling round number, It represents the computing power waiting time of the i-th computing power demander under the u-th computing power scheduling round number, represents the computing power waiting time of the i-th computing power demander under the U-th computing power scheduling round number, T i Indicates the total computing power waiting time of the i-th computing power demander;
[0070] According to the computing power waiting time, the average waiting time of the computing power demander is calculated using the following formula:
[0071]
[0072] Where T′ represents the average waiting time, n represents the total number of computing power demanders, and T represents the computing power waiting time;
[0073] Based on the sum of the number of computing power scheduling times and the average waiting time, a target time slice corresponding to the number of computing power scheduling times and the average waiting time is obtained from the initial time slice, and the target time slice is used as the computing power scheduling time slice.
[0074] In a possible implementation manner of the first aspect, the computing power scheduling time slice based on the computing power scheduling time slice is used to schedule the computing power supplier by using a preset time slice rotation algorithm and the computing power scheduling priority to obtain a computing power scheduling record, including:
[0075] Based on the computing power scheduling time slice, the first computing power supply order corresponding to each computing power demander among the computing power demanders is set by using the time slice rotation algorithm and the computing power scheduling priority;
[0076] Setting a second computing power supply order between the computing power demanders corresponding to the first computing power supply order;
[0077] Randomly generate a supplier sequence of the computing power supplier;
[0078] According to the second computing power supply order, the computing power suppliers in the supplier sequence are scheduled using a preset first-come-first-served scheduling algorithm to obtain a computing power scheduling record.
[0079] In a possible implementation manner of the first aspect, the calculating the computing power scheduling performance of the computing power supplier in the computing power scheduling record includes:
[0080] Using the computing power scheduling record, identifying the target traffic volume that is blocked in the planned traffic volume;
[0081] The ratio of the target business volume to the planned business volume is used as the business blocking rate of the computing power supplier;
[0082] Using the computing power scheduling record, identifying the computing power scheduling time slice that has not been scheduled in the computing power scheduling round;
[0083] The variance of the time slice without computing power scheduling is calculated using the following formula:
[0084]
[0085] Among them, S 2 represents the variance, Y p Indicates the number of time slices that have not been scheduled in the pth computing power scheduling round. It represents the average number of time slices that have not been scheduled in P computing power scheduling rounds, where P represents the total number of computing power scheduling rounds and p represents the sequence number of the computing power scheduling round;
[0086] The service blocking rate and the variance are used as the computing power scheduling performance.
[0087] In a possible implementation manner of the first aspect, determining, according to the computing power scheduling performance, a computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander includes:
[0088] When the service blocking rate in the computing power scheduling performance is the smallest and the variance in the computing power scheduling performance is the smallest, obtaining the optimized computing power scheduling priority and the optimized computing power scheduling time slice from the computing power scheduling priority and the computing power scheduling time slice respectively;
[0089] The optimized computing power scheduling priority and the optimized computing power scheduling time slice are used to determine a computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander.
[0090] In a second aspect, the present invention provides a computing power scheduling system combining GPU virtualization and AI, the system comprising:
[0091] A business analysis module is used to identify the computing power demander, computing power supplier and computing power intermediary when computing power scheduling is performed, detect the historical business volume and business volume time of the computing power demander in the preset historical period in the computing power intermediary, and analyze the planned business volume of the computing power demander in the preset planning period based on the historical business volume and the business volume time;
[0092] A relationship identification module is used to set the computing power scheduling priority of the computing power demander according to the planned business volume, query the historical scheduling duration of the computing power dispatched from the computing power supplier to the computing power demander within a preset historical period, and identify the business-duration mapping relationship between the historical business volume and the historical scheduling duration;
[0093] A time allocation module is used to analyze the planned scheduling duration of scheduling computing power from the computing power supplier to the computing power demander within the planned period according to the business-duration mapping relationship and the planned business volume, and allocate the computing power scheduling time slice of the computing power demander by using the planned scheduling duration and the preset GPU virtualization technology;
[0094] A record judgment module is used to schedule the computing power supplier based on the computing power scheduling time slice, using a preset time slice rotation algorithm and the computing power scheduling priority, to obtain a computing power scheduling record, and to use the computing power scheduling record to determine whether the computing power supplier meets the requirements of the computing power demander;
[0095] an algorithm determination module, configured to calculate the computing power scheduling performance of the computing power supplier in the computing power scheduling record when the computing power supplier does not meet the computing power demander, and determine the computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander according to the computing power scheduling performance;
[0096] The computing power scheduling module is used to use the computing power scheduling algorithm to complete the computing power scheduling from the computing power supplier to the computing power demander, and obtain the computing power scheduling result of the computing power supplier.
[0097] Compared with the prior art, the technical principle and beneficial effects of this solution are:
[0098] The embodiment of the present invention analyzes the planned business volume of the computing power demander within a preset planning period based on the historical business volume and the business volume moment, so as to analyze the business volume that the computing power demander needs to process in a future period, so as to facilitate the subsequent analysis of how long the computing power resources need to be called in the future period. The embodiment of the present invention sets the computing power scheduling priority of the computing power demander according to the planned business volume, so as to arrange the processing priority of the planned business volume by using the order of the big root heap from large to small. Since the order of the big root heap from large to small is not completely from large to small, but smaller values are interspersed in the middle of larger values, the computing power scheduling priority can take into account the computing power scheduling needs of small business volumes and reduce the computing power scheduling hunger phenomenon. Furthermore, the embodiment of the present invention identifies the business-duration mapping relationship between the historical business volume and the historical scheduling duration, so as to use the business-duration mapping relationship to predict the planning scheduling time required for the planned business volume in the future period. Long, the embodiment of the present invention allocates the computing power scheduling time slice of the computing power demander by using the planned scheduling duration and the preset GPU virtualization technology, so as to allocate a fixed time slice for each business, thereby reducing the phenomenon of starvation of the business not being able to obtain computing power resources for a long time. The embodiment of the present invention calculates the computing power scheduling performance of the computing power supplier in the computing power scheduling record, so as to obtain the priority and time slice with the best computing power scheduling performance from the computing power scheduling priority and the computing power scheduling time slice, and uses the variance of the computing power scheduling performance without computing power scheduling time slice to observe whether the time slice allocated to the computing power demander can make the business volume of each computing power demander balanced. Further, the embodiment of the present invention determines the computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander according to the computing power scheduling performance, so as to obtain the priority and time slice with the best computing power scheduling performance from the computing power scheduling priority and the computing power scheduling time slice. Therefore, the computing power scheduling method and system combining GPU virtualization and AI proposed in the embodiment of the present invention can reduce the problem of hunger in computing power scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0100] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0101] Figure 1A flowchart of a computing power scheduling method combining GPU virtualization and AI provided by one embodiment of the present invention;
[0102] Figure 2 A module diagram of a computing power scheduling system combining GPU virtualization and AI provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0103] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0104] The embodiment of the present invention provides a computing power scheduling method combining GPU virtualization and AI, and the execution subject of the computing power scheduling method combining GPU virtualization and AI includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiment of the present invention. In other words, the computing power scheduling method combining GPU virtualization and AI can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0105] See also Figure 1 As shown in FIG. 1 , it is a flow chart of a computing power scheduling method combining GPU virtualization and AI provided by an embodiment of the present invention. Figure 1 The computing power scheduling method combining GPU virtualization and AI described in the article includes:
[0106] S1. Identify the computing power demander, computing power supplier and computing power intermediary during computing power scheduling, detect the historical business volume and business volume time of the computing power demander within a preset historical period in the computing power intermediary, and analyze the planned business volume of the computing power demander within a preset planning period based on the historical business volume and the business volume time.
[0107] In an embodiment of the present invention, the computing power demander refers to an area or system that requires computing power resources, the computing power supplier refers to a computer processing system that supplies computing power resources to the computing power demander, and the computing power intermediary refers to a computer processing system that transmits the pending business of the computing power demander to the computing power supplier. The computing power intermediary is used to share the pressure for the computing power supplier, analyze the business data of the computing power demander in advance, and reasonably arrange the computing power demander's needs for computing power resource scheduling.
[0108] Furthermore, in an embodiment of the present invention, the historical business volume refers to network business requests, including business requests sent by school units, enterprise units, factory units, etc., and the business volume moment refers to the moment when the historical business volume arrives at the computing power intermediary. It should be noted that due to the huge business volume, business volume arrives at the computing power intermediary almost every moment. Therefore, the business volume moment is a moment with a certain fixed time interval, such as a business volume moment every 10 minutes, and the historical business volume is the business volume at each business volume moment.
[0109] Furthermore, the embodiment of the present invention analyzes the planned business volume of the computing power demander within a preset planning period based on the historical business volume and the business volume moment, so as to analyze the business volume that the computing power demander needs to process in a future period, thereby facilitating subsequent analysis of how long the computing power resources need to be called in the future period.
[0110] The planned business volume refers to the business volume within a preset planning period.
[0111] In one embodiment of the present invention, the analyzing the planned business volume of the computing power demander within a preset planning period based on the historical business volume and the business volume moment includes: based on the business volume moment, converting the historical business volume into a business volume sequence using the following format:
[0112]
[0113] X=(x (1) , x (2) , x (j) ,...x (m) )
[0114]
[0115] X=(X1,X2,...X i , ..., X n ) T
[0116] Where X represents the business volume sequence, i represents the order number of the computing power demander, n represents the total number of computing power demanders, j represents the order number of the business volume moment, m represents the total number of business volume moments, and x (j) represents the historical business volume at the jth business moment, It represents the historical business volume of the first computing power demander at the jth business volume moment. It represents the historical business volume of the second computing power demander at the jth business volume moment, It represents the historical business volume of the i-th computing power demander at the j-th business volume moment, represents the historical business volume of the nth computing power demander at the jth business volume moment, x (1) represents the historical business volume at the first business volume moment, x (2) represents the historical business volume at the second business volume moment, x (m) represents the historical business volume at the mth business moment, X i represents the historical business volume of the i-th computing power demander, It represents the historical business volume of the i-th computing power demander at the first business volume moment, It represents the historical business volume of the i-th computing power demander at the second business volume moment, represents the historical business volume of the i-th computing power demander at the m-th business volume moment, X1 represents the historical business volume of the first computing power demander, X2 represents the historical business volume of the second computing power demander, and X n Indicates the historical business volume of the nth computing power demander;
[0117] The business volume sequence is feature extracted using a preset principal component analysis method to obtain extracted features; the extracted features are input into a preset long short-term memory recurrent network; in the long short-term memory recurrent network, the planned business volume of the computing power demander within a preset planning period is calculated according to the extracted features using the following formula:
[0118] ft=σ(Wf·[ht-1,xt]+bf)
[0119] i t =σ(W i ·[h t-1 , x t ]+b i )
[0120] C t =tanh(W C ·[h t-1 , x t ]+b C )
[0121] C t =f t ·C t-1 +i t ·C t
[0122] o t =σ(W o ·[h t-1 , x t ]+b o )
[0123] h t =ot tanh(C t )
[0124] Among them, h t represents the planned traffic volume, σ represents, x t represents the extracted features at historical moment t, h t-1 represents the hidden state of the previous moment of historical moment t, W f represents the weight matrix of the forget gate, b f represents the bias term of the forget gate, σ represents the sigmoid function, and f t represents the output of the forget gate at historical time t, W i represents the weight matrix of the input gate, b i represents the bias term of the input gate, i t Represents the output result of the input gate, W C represents the weight matrix of the memory gate, b C represents the bias term of the memory gate, C t represents the cell state output by the memory gate at historical time t, C t-1 represents the cell state at the previous moment of historical moment t, W o represents the weight matrix of the output gate, b o represents the bias term of the output gate, o t Represents the output result of the output gate.
[0125] In another embodiment of the present invention, the method of extracting features from the traffic volume sequence using a preset principal component analysis method to obtain extracted features includes: calculating the covariance matrix of the traffic volume sequence using a preset principal component analysis method; calculating eigenvalues and eigenvectors of the traffic volume sequence based on the covariance matrix using the preset principal component analysis method; selecting non-zero eigenvalues from the eigenvalues; identifying a target eigenvector corresponding to the non-zero eigenvalue in the eigenvector; and using the target eigenvector as the extracted feature.
[0126] S2. According to the planned business volume, the computing power scheduling priority of the computing power demander is set, the historical scheduling duration of the computing power dispatched from the computing power supplier to the computing power demander within a preset historical period is queried, and the business-duration mapping relationship between the historical business volume and the historical scheduling duration is identified.
[0127] The embodiment of the present invention sets the computing power scheduling priority of the computing power demander according to the planned business volume, so as to arrange the processing priority of the planned business volume by using the order of the big root heap from large to small. Since the order of the big root heap from large to small is not completely from large to small, but smaller values are interspersed among larger values, the computing power scheduling priority can take into account the computing power scheduling needs of small business volumes and reduce the computing power scheduling hunger phenomenon.
[0128] In one embodiment of the present invention, setting the computing power scheduling priority of the computing power demander according to the planned business volume includes: randomly arranging the planned business volume to obtain a business volume queue; converting the business volume queue into a big root heap; querying the front and back order of the planned business volume in the big root heap according to a preset pre-order traversal algorithm; based on the front and back order, setting the computing power scheduling priority of the planned business volume according to a preset first-come-first-served scheduling algorithm.
[0129] It should be noted that since the establishment of the big root heap is based on the randomly arranged business volume queues, there are multiple types of big root heaps that are constructed.
[0130] Furthermore, in an embodiment of the present invention, the historical scheduling duration refers to the total duration that the computing power demander uses the scheduled computing power to process historical business volume when the computing power is scheduled to the computing power demander.
[0131] Furthermore, the embodiment of the present invention identifies the service-duration mapping relationship between the historical service volume and the historical scheduling duration, so as to utilize the service-duration mapping relationship to predict the planned scheduling duration required for the planned service volume in the future period.
[0132] The service-duration mapping relationship refers to the relationship of inferring the historical scheduling duration through the historical service volume.
[0133] In one embodiment of the present invention, the identifying of the business-duration mapping relationship between the historical business volume and the historical scheduling duration includes: arranging the historical business volume in order from small to large to obtain a business volume sequence; arranging the historical scheduling duration according to the business volume sequence to obtain a duration sequence; normalizing the business volume sequence using a missing value interpolation method to obtain a normalized sequence; inputting the normalized sequence into a preset duration analysis model to output a normalized duration corresponding to the normalized sequence through the duration analysis model; identifying a target duration in the normalized duration that corresponds to the business volume sequence; calculating a duration loss value between the duration sequence and the target duration to optimize the duration analysis model through the duration loss value to obtain an optimized duration analysis model; and using the optimized duration analysis model to identify the business-duration mapping relationship between the historical business volume and the historical scheduling duration.
[0134] Among them, the duration analysis model refers to a long short-term memory neural network. In an embodiment of the present invention, the normalized order replaces the time series, and the historical scheduling duration is used as the output of the long short-term memory neural network under the normalized order. The target duration refers to the scheduling duration of each business volume in the business volume sequence. This is because the normalized duration is the scheduling duration under the normalized order with a fixed interval, and a part of the scheduling duration belongs to the newly interpolated business volume. Therefore, it is necessary to remove this part of the business volume and filter out the scheduling duration belonging to the business volume sequence.
[0135] Optionally, the process of normalizing the business volume sequence by using the missing value interpolation method to obtain the normalized sequence refers to the process of normalizing the growth amplitude of the business volume from small to large to a fixed growth amplitude. For example, the business volume sequence is 2, 3, 5, 9, and the normalized order is 2, 3, 4, 5, 6, 7, 8, 9. When normalizing the business volume sequence, the minimum growth amplitude in the business volume sequence is used as the fixed growth amplitude.
[0136] S3. According to the business-duration mapping relationship and the planned business volume, the planned scheduling duration for scheduling computing power from the computing power supplier to the computing power demander within the planned period is analyzed, and the computing power scheduling time slice of the computing power demander is allocated using the planned scheduling duration and the preset GPU virtualization technology.
[0137] The embodiment of the present invention allocates computing power scheduling time slices of the computing power demander by utilizing the planned scheduling duration and the preset GPU virtualization technology, so as to allocate fixed time slices to each business, thereby reducing the phenomenon of starvation of businesses due to long-term lack of computing power resources.
[0138] In one embodiment of the present invention, the use of the planned scheduling duration and the preset GPU virtualization technology to allocate the computing power scheduling time slice of the computing power demander includes: using the GPU virtualization technology to allocate the initial time slice of the computing power demander; according to the initial time slice and the planned scheduling duration, using the following formula to calculate the number of computing power scheduling of the computing power demander:
[0139]
[0140]
[0141] Among them, K represents the number of computing power scheduling, k i represents the number of computing power scheduling of the i-th computing power demander, t i It represents the planning and scheduling duration of the i-th computing power demander, t0 represents the initial time slice, Indicates the rounding up symbol;
[0142] Determine the computing power scheduling round of the planned scheduling duration by using the computing power scheduling times; divide the planned scheduling duration by using the initial time slice to obtain the divided duration; determine the computing power scheduling round number of the divided duration by using the computing power scheduling round; obtain the computing power scheduling priority of the computing power demander; construct the priority of the computing power scheduling round number by using the computing power scheduling priority; sort the planned scheduling duration by using the computing power scheduling round number and the priority to obtain a planned duration sequence; calculate the computing power waiting time of the computing power demander by using the following formula according to the planned duration sequence:
[0143]
[0144]
[0145] …
[0146]
[0147] …
[0148]
[0149]
[0150]
[0151] Among them, T represents the computing power waiting time, T i represents the computing power waiting time of the i-th computing power demander, u represents the computing power scheduling round number in the planned duration sequence, U represents the total number of computing power scheduling round numbers in the planned duration sequence, v represents the priority in the planned duration sequence, V represents the total number of priorities in the planned duration sequence, i represents the sequence number of the computing power demander, n represents the total number of computing power demanders, t0 represents the initial time slice, It represents the computing power waiting time of the i-th computing power demander under the first computing power scheduling round number, It represents the computing power waiting time of the i-th computing power demander under the second computing power scheduling round number, It represents the computing power waiting time of the i-th computing power demander under the u-th computing power scheduling round number, represents the computing power waiting time of the i-th computing power demander under the U-th computing power scheduling round number, T i Indicates the total computing power waiting time of the i-th computing power demander;
[0152] According to the computing power waiting time, the average waiting time of the computing power demander is calculated using the following formula:
[0153]
[0154] Where T′ represents the average waiting time, n represents the total number of computing power demanders, and T represents the computing power waiting time;
[0155] Based on the sum of the number of computing power scheduling times and the average waiting time, a target time slice corresponding to the number of computing power scheduling times and the average waiting time is obtained from the initial time slice, and the target time slice is used as the computing power scheduling time slice.
[0156] Among them, the initial time slice refers to a plurality of time slices of different lengths, and the computing power scheduling round refers to the number of times the time slice requirements of all computing power demanders are processed in turn at one time. For example, if there are 5 computing power demanders, each computing power demander is allocated a time slice. After processing 5 time slices, the time slice requirements of all computing power demanders are processed in turn at one time. Since the business volume of each computing power demander is huge, one time slice cannot handle the business volume of one computing power demander, so a second round of computing power scheduling is required. The length of the split time is consistent with the length of the initial time slice. The computing power scheduling round number refers to the number of the computing power scheduling round that a computing power demander needs to perform. The priority Priority refers to the priority of the computing power scheduling round number of the computing power demander in each round of computing power scheduling. For example, in the first round, the computing power scheduling round number of the computing power demander (A) is 1. The computing power scheduling round number of the computing power demander (A) is higher than the computing power scheduling round numbers of other computing power demanders (B, C, D). This is because the order of computing power scheduling priority is A, B, C, D. The planned duration sequence refers to the result of splicing the sequence of each round according to the priority of each round, and then splicing the sequence of each round according to the order of the computing power scheduling round number. The target time slice refers to several time slices selected from the initial time slice with a smaller sum of the number of computing power scheduling times and the average waiting time.
[0157] S4. Based on the computing power scheduling time slice, the computing power supplier is scheduled using a preset time slice rotation algorithm and the computing power scheduling priority to obtain a computing power scheduling record, and the computing power scheduling record is used to determine whether the computing power supplier meets the requirements of the computing power demander.
[0158] In the embodiment of the present invention, the computing power scheduling record refers to a record of computing power scheduling time slices including which computing power suppliers are scheduled to process which computing power demanders at which times.
[0159] In one embodiment of the present invention, the computing power scheduling time slice is based on the computing power scheduling time slice, and the computing power scheduling priority is used to schedule the computing power supplier to obtain a computing power scheduling record, including: based on the computing power scheduling time slice, the time slice rotation algorithm and the computing power scheduling priority are used to set the first computing power supply order corresponding to each computing power demander among the computing power demanders; the second computing power supply order between the computing power demanders corresponding to the first computing power supply order is set; the supplier sequence of the computing power suppliers is randomly generated; according to the second computing power supply order, the computing power suppliers in the supplier sequence are scheduled using a preset first-come-first-served scheduling algorithm to obtain a computing power scheduling record.
[0160] Among them, the first computing power supply order refers to the order of computing power scheduling time slices arranged according to the computing power scheduling round number, and the second computing power supply order is consistent with the above-mentioned planned duration sequence.
[0161] Optionally, the process of using the computing power scheduling record to determine whether the computing power supplier meets the requirements of the computing power demander refers to determining whether the computing power supplier fails to complete the business volume of the computing power demander within the computing power scheduling time slice.
[0162] S5. When the computing power supplier does not meet the requirements of the computing power demander, calculate the computing power scheduling performance of the computing power supplier in the computing power scheduling record, and determine a computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander based on the computing power scheduling performance.
[0163] The embodiment of the present invention calculates the computing power scheduling performance of the computing power supplier in the computing power scheduling record to obtain the priority and time slice with the best computing power scheduling performance from the computing power scheduling priority and the computing power scheduling time slice, and uses the variance of the un-computing power scheduling time slice in the computing power scheduling performance to observe whether the time slice allocated to the computing power demander can balance the business volume of each computing power demander.
[0164] In one embodiment of the present invention, the computing power scheduling performance of the computing power supplier in the computing power scheduling record includes: using the computing power scheduling record to identify the target traffic volume blocked in the planned traffic volume; taking the ratio of the target traffic volume to the planned traffic volume as the traffic blocking rate of the computing power supplier; using the computing power scheduling record to identify the computing power scheduling time slice in the computing power scheduling round where the computing power scheduling time slice is not scheduled; and using the following formula to calculate the variance of the time slice where the computing power scheduling is not performed:
[0165]
[0166] Among them, S 2 represents the variance, Yp Indicates the number of time slices that have not been scheduled in the pth computing power scheduling round. It represents the average number of time slices that have not been scheduled in P computing power scheduling rounds, where P represents the total number of computing power scheduling rounds and p represents the sequence number of the computing power scheduling round;
[0167] The service blocking rate and the variance are used as the computing power scheduling performance.
[0168] Among them, the business blocking rate refers to the proportion of user requests that cannot be met due to insufficient network resources within a certain period of time. The time slices that are not scheduled for computing power refer to the time slices that cannot be scheduled for computing power due to blocking or the time slices that are wasted due to the early processing of business volume. If the variance is large, it means that the number of time slices that are not scheduled for computing power under different computing power scheduling rounds is quite different, which may indicate that the computing power scheduling of different computing power demanders is unbalanced.
[0169] Furthermore, an embodiment of the present invention determines a computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander based on the computing power scheduling performance, so as to obtain the priority and time slice with the best computing power scheduling performance from the computing power scheduling priority and the computing power scheduling time slice.
[0170] The computing power scheduling algorithm includes an algorithm for computing power scheduling by optimizing computing power scheduling priority and optimizing computing power scheduling time slice.
[0171] In one embodiment of the present invention, the computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander is determined based on the computing power scheduling performance, including: when the service blocking rate in the computing power scheduling performance is the smallest and the variance in the computing power scheduling performance is the smallest, obtaining the optimized computing power scheduling priority and the optimized computing power scheduling time slice from the computing power scheduling priority and the computing power scheduling time slice respectively; and determining the computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander using the optimized computing power scheduling priority and the optimized computing power scheduling time slice.
[0172] S6. Use the computing power scheduling algorithm to complete the computing power scheduling from the computing power supplier to the computing power demander, and obtain the computing power scheduling result of the computing power supplier.
[0173] It can be seen that the embodiment of the present invention analyzes the planned business volume of the computing power demander in the preset planning period based on the historical business volume and the business volume moment, so as to analyze the business volume that the computing power demander needs to process in the future period, so as to facilitate the subsequent analysis of how long the computing power resources need to be called in the future period. The embodiment of the present invention sets the computing power scheduling priority of the computing power demander according to the planned business volume, so as to arrange the processing priority of the planned business volume by using the order from large to small of the big root heap. Since the order from large to small of the big root heap is not completely from large to small, but smaller values are interspersed in the middle of larger values, the computing power scheduling priority can take into account the computing power scheduling needs of small business volumes and reduce the computing power scheduling hunger phenomenon. Furthermore, the embodiment of the present invention identifies the business-duration mapping relationship between the historical business volume and the historical scheduling duration, so as to use the business-duration mapping relationship to predict the planning required for the planned business volume in the future period. Scheduling duration, the embodiment of the present invention allocates the computing power scheduling time slice of the computing power demander by utilizing the planned scheduling duration and the preset GPU virtualization technology, so as to allocate a fixed time slice for each business, thereby reducing the phenomenon of starvation of the business not being able to obtain computing power resources for a long time. The embodiment of the present invention calculates the computing power scheduling performance of the computing power supplier in the computing power scheduling record, so as to obtain the priority and time slice with the best computing power scheduling performance from the computing power scheduling priority and the computing power scheduling time slice, and uses the variance of the computing power scheduling performance without computing power scheduling time slice to observe whether the time slice allocated to the computing power demander can balance the business volume of each computing power demander. Furthermore, the embodiment of the present invention determines the computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander according to the computing power scheduling performance, so as to obtain the priority and time slice with the best computing power scheduling performance from the computing power scheduling priority and the computing power scheduling time slice. Therefore, the computing power scheduling method combining GPU virtualization and AI proposed in the embodiment of the present invention can reduce the starvation problem of computing power scheduling.
[0174] like Figure 2 As shown, it is a functional module diagram of the computing power scheduling system combining GPU virtualization and AI of the present invention.
[0175] The computing power scheduling system 200 combining GPU virtualization and AI of the present invention can be installed in an electronic device. According to the functions implemented, the computing power scheduling system combining GPU virtualization and AI can include a business analysis module 201, a relationship identification module 202, a time allocation module 203, a record judgment module 204, an algorithm determination module 205 and a computing power scheduling module 206. The module of the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0176] In the embodiment of the present invention, the functions of each module / unit are as follows:
[0177] The business analysis module 201 is used to identify the computing power demander, computing power supplier and computing power intermediary when computing power scheduling is performed, detect the historical business volume and business volume time of the computing power demander in the preset historical period in the computing power intermediary, and analyze the planned business volume of the computing power demander in the preset planning period based on the historical business volume and the business volume time;
[0178] The relationship identification module 202 is used to set the computing power scheduling priority of the computing power demander according to the planned business volume, query the historical scheduling duration of the computing power dispatched from the computing power supplier to the computing power demander within a preset historical period, and identify the business-duration mapping relationship between the historical business volume and the historical scheduling duration;
[0179] The time allocation module 203 is used to analyze the planned scheduling duration of scheduling computing power from the computing power supplier to the computing power demander within the planned period according to the business-duration mapping relationship and the planned business volume, and allocate the computing power scheduling time slice of the computing power demander by using the planned scheduling duration and the preset GPU virtualization technology;
[0180] The record judgment module 204 is used to schedule the computing power supplier based on the computing power scheduling time slice, using a preset time slice rotation algorithm and the computing power scheduling priority, to obtain a computing power scheduling record, and use the computing power scheduling record to determine whether the computing power supplier meets the requirements of the computing power demander;
[0181] The algorithm determination module 205 is used to calculate the computing power scheduling performance of the computing power supplier in the computing power scheduling record when the computing power supplier does not meet the computing power demander, and determine the computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander according to the computing power scheduling performance;
[0182] The computing power scheduling module 206 is used to complete the computing power scheduling from the computing power supplier to the computing power demander by using the computing power scheduling algorithm, and obtain the computing power scheduling result of the computing power supplier.
[0183] In detail, the modules in the computing power scheduling system 200 combining GPU virtualization and AI in the embodiment of the present invention are used in the same manner as above. Figure 1 The computing power scheduling method combining GPU virtualization and AI described in the previous section uses the same technical means and can produce the same technical effects, so I will not go into details here.
[0184] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0185] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0186] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0187] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0188] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A computing power scheduling method combining GPU virtualization and AI, characterized in that: The method comprises: Identify the computing power demander, computing power supplier and computing power intermediary during computing power scheduling, detect the historical business volume and business volume time of the computing power demander in a preset historical period in the computing power intermediary, and analyze the planned business volume of the computing power demander in a preset planning period based on the historical business volume and the business volume time, wherein the analyzing the planned business volume of the computing power demander in the preset planning period based on the historical business volume and the business volume time includes: Based on the traffic moment, the historical traffic is converted into a traffic sequence using the following format: X=(x (1) ,x (2) ,…x (j) ,…x (m) ) X=(X1,X2,...X i ,…,X n ) T Where X represents the business volume sequence, i represents the order number of the computing power demander, n represents the total number of computing power demanders, j represents the order number of the business volume moment, m represents the total number of business volume moments, and x (j) represents the historical business volume at the jth business moment, It represents the historical business volume of the first computing power demander at the jth business volume moment. It represents the historical business volume of the second computing power demander at the jth business volume moment, It represents the historical business volume of the i-th computing power demander at the j-th business volume moment, represents the historical business volume of the nth computing power demander at the jth business volume moment, x (1) represents the historical business volume at the first business volume moment, x (2) represents the historical business volume at the second business volume moment, x (m) represents the historical business volume at the mth business moment, X i represents the historical business volume of the i-th computing power demander, It represents the historical business volume of the i-th computing power demander at the first business volume moment, It represents the historical business volume of the i-th computing power demander at the second business volume moment, represents the historical business volume of the i-th computing power demander at the m-th business volume moment, X1 represents the historical business volume of the first computing power demander, X2 represents the historical business volume of the second computing power demander, and X n Indicates the historical business volume of the nth computing power demander; Extracting features from the traffic sequence using a preset principal component analysis method to obtain extracted features; Inputting the extracted features into a preset long short-term memory recurrent network; In the long short-term memory recurrent network, according to the extracted features, the planned business volume of the computing power demander within the preset planning period is calculated using the following formula: f t =σ(W f ·[h t-1 ,x t ]+b f ) i t =σ(W i ·[h t-1 ,x t ]+b i ) C t =tanh(W C ·[h t-1 ,x t ]+b C ) C t =f t ·C t-1 +i t ·C t the t =σ(W o ·[h t-1 ,x t ]+b o ) h t =o t ·tanh(C t ) Among them, h t represents the planned traffic volume, x t represents the extracted features at historical moment t, h t-1 represents the hidden state of the previous moment of historical moment t, W f represents the weight matrix of the forget gate, b f represents the bias term of the forget gate, σ represents the sigmoid function, and f t represents the output of the forget gate at historical time t, W i represents the weight matrix of the input gate, b i represents the bias term of the input gate, i t Represents the output result of the input gate, W C represents the weight matrix of the memory gate, b C represents the bias term of the memory gate, C t represents the cell state output by the memory gate at historical time t, C t-1 represents the cell state at the previous moment of historical moment t, W o represents the weight matrix of the output gate, b o represents the bias term of the output gate, o t Represents the output result of the output gate; According to the planned business volume, the computing power scheduling priority of the computing power demander is set, the historical scheduling duration of the computing power dispatched from the computing power supplier to the computing power demander within a preset historical period is queried, and the business-duration mapping relationship between the historical business volume and the historical scheduling duration is identified; According to the business-duration mapping relationship and the planned business volume, the planned scheduling duration for scheduling computing power from the computing power supplier to the computing power demander within the planned period is analyzed, and the computing power scheduling time slice of the computing power demander is allocated by using the planned scheduling duration and the preset GPU virtualization technology; Based on the computing power scheduling time slice, the computing power supplier is scheduled using a preset time slice rotation algorithm and the computing power scheduling priority to obtain a computing power scheduling record, and the computing power scheduling record is used to determine whether the computing power supplier meets the requirements of the computing power demander; When the computing power supplier does not meet the computing power demander, calculating the computing power scheduling performance of the computing power supplier in the computing power scheduling record, and determining a computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander according to the computing power scheduling performance; The computing power scheduling algorithm is used to complete the computing power scheduling from the computing power supplier to the computing power demander, and the computing power scheduling result of the computing power supplier is obtained.
2. The method according to claim 1, characterized in that The extracting features of the traffic sequence by using a preset principal component analysis method to obtain extracted features includes: Calculating the covariance matrix of the traffic sequence using a preset principal component analysis method; According to the covariance matrix, using the preset principal component analysis method to calculate the eigenvalue and eigenvector of the traffic sequence; Selecting a non-zero eigenvalue from the eigenvalues; Identifying a target feature vector corresponding to the non-zero feature value in the feature vector; The target feature vector is used as the extracted feature.
3. The method according to claim 1, characterized in that The step of setting the computing power scheduling priority of the computing power demander according to the planned business volume includes: Randomly arranging the planned traffic to obtain a traffic queue; Converting the traffic queue into a large root heap; Querying the front-to-back order of the planned business volume in the big root heap according to a preset pre-order traversal algorithm; Based on the sequence, the computing power scheduling priority of the planned business volume is set according to a preset first-come-first-served scheduling algorithm.
4. The method according to claim 1, characterized in that: The identifying a service-duration mapping relationship between the historical service volume and the historical scheduling duration includes: Arrange the historical business volumes in ascending order to obtain a business volume sequence; Arranging the historical scheduling durations according to the business volume sequence to obtain a duration sequence; Normalizing the business volume sequence by using a missing value interpolation method to obtain a normalized sequence; Inputting the normalized sequence into a preset duration analysis model, so as to output a normalized duration corresponding to the normalized sequence through the duration analysis model; Identifying a target duration in the normalized duration corresponding to the traffic sequence; Calculating a duration loss value between the duration sequence and the target duration, so as to optimize the duration analysis model by using the duration loss value to obtain an optimized duration analysis model; The optimized duration analysis model is used to identify the service-duration mapping relationship between the historical service volume and the historical scheduling duration.
5. The method according to claim 1, characterized in that The method of allocating the computing power scheduling time slice of the computing power demander by utilizing the planned scheduling duration and the preset GPU virtualization technology includes: Allocate the initial time slice of the computing power demander by using the GPU virtualization technology; According to the initial time slice and the planned scheduling duration, the number of computing power scheduling times of the computing power demander is calculated using the following formula: Among them, K represents the number of computing power scheduling, k i represents the number of computing power scheduling of the i-th computing power demander, t i It represents the planning and scheduling duration of the i-th computing power demander, t0 represents the initial time slice, Indicates the rounding up symbol; Determine the computing power scheduling round of the planned scheduling duration by using the computing power scheduling times; Using the initial time slice to divide the planned scheduling duration into divided durations to obtain divided durations; Determine the computing power scheduling round number of the split time using the computing power scheduling round; Obtaining the computing power scheduling priority of the computing power demander; Constructing the priority of the computing power scheduling round number using the computing power scheduling priority; Sorting the planned scheduling durations by using the computing power scheduling round number and the priority to obtain a planned duration sequence; According to the planned duration sequence, the computing power waiting time of the computing power demander is calculated using the following formula: Among them, T represents the computing power waiting time, T i represents the computing power waiting time of the i-th computing power demander, u represents the computing power scheduling round number in the planned duration sequence, U represents the total number of computing power scheduling round numbers in the planned duration sequence, v represents the priority in the planned duration sequence, V represents the total number of priorities in the planned duration sequence, i represents the sequence number of the computing power demander, n represents the total number of computing power demanders, t0 represents the initial time slice, It represents the computing power waiting time of the i-th computing power demander under the first computing power scheduling round number, It represents the computing power waiting time of the i-th computing power demander under the second computing power scheduling round number, It represents the computing power waiting time of the i-th computing power demander under the u-th computing power scheduling round number, represents the computing power waiting time of the i-th computing power demander under the U-th computing power scheduling round number, T i Indicates the total computing power waiting time of the i-th computing power demander; According to the computing power waiting time, the average waiting time of the computing power demander is calculated using the following formula: Where T′ represents the average waiting time, n represents the total number of computing power demanders, and T represents the computing power waiting time; Based on the sum of the number of computing power scheduling times and the average waiting time, a target time slice corresponding to the number of computing power scheduling times and the average waiting time is obtained from the initial time slice, and the target time slice is used as the computing power scheduling time slice.
6. The method according to claim 1, characterized in that The method of scheduling the computing power supplier based on the computing power scheduling time slice by using a preset time slice rotation algorithm and the computing power scheduling priority to obtain a computing power scheduling record includes: Based on the computing power scheduling time slice, the first computing power supply order corresponding to each computing power demander among the computing power demanders is set by using the time slice rotation algorithm and the computing power scheduling priority; Setting a second computing power supply order between the computing power demanders corresponding to the first computing power supply order; Randomly generate a supplier sequence of the computing power supplier; According to the second computing power supply order, the computing power suppliers in the supplier sequence are scheduled using a preset first-come-first-served scheduling algorithm to obtain a computing power scheduling record.
7. The method according to claim 1, characterized in that The calculating the computing power scheduling performance of the computing power supplier in the computing power scheduling record includes: Using the computing power scheduling record, identifying the target traffic volume that is blocked in the planned traffic volume; The ratio of the target business volume to the planned business volume is used as the business blocking rate of the computing power supplier; Using the computing power scheduling record, identifying the computing power scheduling time slice that has not been scheduled in the computing power scheduling round; The variance of the time slice without computing power scheduling is calculated using the following formula: Among them, S 2 represents the variance, Y p Indicates the number of time slices that have not been scheduled in the pth computing power scheduling round. It represents the average number of time slices that have not been scheduled in P computing power scheduling rounds, where P represents the total number of computing power scheduling rounds and p represents the sequence number of the computing power scheduling round; The service blocking rate and the variance are used as the computing power scheduling performance.
8. The method according to claim 1, characterized in that The determining, according to the computing power scheduling performance, a computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander comprises: When the service blocking rate in the computing power scheduling performance is the smallest and the variance in the computing power scheduling performance is the smallest, obtaining the optimized computing power scheduling priority and the optimized computing power scheduling time slice from the computing power scheduling priority and the computing power scheduling time slice respectively; The optimized computing power scheduling priority and the optimized computing power scheduling time slice are used to determine a computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander.
9. A computing power scheduling system combining GPU virtualization and AI, characterized in that: The system comprises: A business analysis module is used to identify a computing power demander, a computing power supplier and a computing power intermediary when performing computing power scheduling, detect the historical business volume and business volume moment of the computing power demander in a preset historical period in the computing power intermediary, and analyze the planned business volume of the computing power demander in a preset planning period based on the historical business volume and the business volume moment, wherein the analysis of the planned business volume of the computing power demander in a preset planning period based on the historical business volume and the business volume moment includes: Based on the traffic moment, the historical traffic is converted into a traffic sequence using the following format: X=(x (1) ,x (2) ,…x (j) ,…x (m) ) X=(X1,X2,...X i ,…,X n ) T Where X represents the business volume sequence, i represents the order number of the computing power demander, n represents the total number of computing power demanders, j represents the order number of the business volume moment, m represents the total number of business volume moments, and x (j) represents the historical business volume at the jth business moment, It represents the historical business volume of the first computing power demander at the jth business volume moment. It represents the historical business volume of the second computing power demander at the jth business volume moment, It represents the historical business volume of the i-th computing power demander at the j-th business volume moment, represents the historical business volume of the nth computing power demander at the jth business volume moment, x (1) represents the historical business volume at the first business volume moment, x (2) represents the historical business volume at the second business volume moment, x (m) represents the historical business volume at the mth business moment, X i represents the historical business volume of the i-th computing power demander, It represents the historical business volume of the i-th computing power demander at the first business volume moment, It represents the historical business volume of the i-th computing power demander at the second business volume moment, represents the historical business volume of the i-th computing power demander at the m-th business volume moment, X1 represents the historical business volume of the first computing power demander, X2 represents the historical business volume of the second computing power demander, and X n Indicates the historical business volume of the nth computing power demander; Extracting features from the traffic sequence using a preset principal component analysis method to obtain extracted features; Inputting the extracted features into a preset long short-term memory recurrent network; In the long short-term memory recurrent network, according to the extracted features, the planned business volume of the computing power demander within the preset planning period is calculated using the following formula: f t =σ(W f ·[h t-1 ,x t ]+b f ) i t =σ(W i ·[h t-1 ,x t ]+b i ) C t =tanh(W C ·[h t-1 ,x t ]+b C ) C t =f t ·C t-1 +i t ·C t the t =σ(W o ·[h t-1 ,x t ]+b o ) h t =o t ·tanh(C t ) Among them, h t represents the planned traffic volume, x t represents the extracted features at historical moment t, h t-1 represents the hidden state of the previous moment of historical moment t, W f represents the weight matrix of the forget gate, b f represents the bias term of the forget gate, σ represents the sigmoid function, and f t represents the output of the forget gate at historical time t, W i represents the weight matrix of the input gate, b i represents the bias term of the input gate, i t Represents the output result of the input gate, W C represents the weight matrix of the memory gate, b C represents the bias term of the memory gate, C t represents the cell state output by the memory gate at historical time t, C t-1 represents the cell state at the previous moment of historical moment t, W o represents the weight matrix of the output gate, b o represents the bias term of the output gate, o t Represents the output result of the output gate; A relationship identification module is used to set the computing power scheduling priority of the computing power demander according to the planned business volume, query the historical scheduling duration of the computing power dispatched from the computing power supplier to the computing power demander within a preset historical period, and identify the business-duration mapping relationship between the historical business volume and the historical scheduling duration; A time allocation module is used to analyze the planned scheduling duration of scheduling computing power from the computing power supplier to the computing power demander within the planned period according to the business-duration mapping relationship and the planned business volume, and allocate the computing power scheduling time slice of the computing power demander by using the planned scheduling duration and the preset GPU virtualization technology; A record judgment module is used to schedule the computing power supplier based on the computing power scheduling time slice, using a preset time slice rotation algorithm and the computing power scheduling priority, to obtain a computing power scheduling record, and to use the computing power scheduling record to determine whether the computing power supplier meets the requirements of the computing power demander; an algorithm determination module, configured to calculate the computing power scheduling performance of the computing power supplier in the computing power scheduling record when the computing power supplier does not meet the computing power demander, and determine the computing power scheduling algorithm for scheduling computing power from the computing power supplier to the computing power demander according to the computing power scheduling performance; The computing power scheduling module is used to use the computing power scheduling algorithm to complete the computing power scheduling from the computing power supplier to the computing power demander, and obtain the computing power scheduling result of the computing power supplier.
Citation Information
Patent Citations
Power load prediction method based on affair graph
CN115577754A
Decoding method, text recognition method and device
CN118038463A