Resource scheduling method, elastic computing system, computing device, medium and product
By acquiring and processing resource information and event information in the cloud computing system and predicting future resource demand, the problems of volatility and uncertainty of resource demand in cloud computing are solved, and more efficient resource scheduling and utilization are achieved.
Patent Information
- Application Number
- CN202510588394.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
In cloud computing scenarios, users' demand for resources often shows volatility and uncertainty, and it is difficult to achieve flexible predictions through static configuration rules, resulting in over-allocation of resources or insufficient supply, affecting service quality.
By obtaining resource information of system resources in the elastic computing system and event information of system events, performing time stamp alignment, building a time series of system events, performing cross-modal timing fusion processing, obtaining timing correlation characteristics, predicting resource scheduling strategies for future system events, and scheduling system resources based on this strategy.
It realizes more accurate prediction of resource requirements, improves the timeliness of resource scheduling, reduces resource waste, and improves the overall utilization efficiency of GPU clusters.
Smart Images

Figure CN120106687A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of cloud computing technology, and in particular to a resource scheduling method, an elastic computing system, a computing device, a medium, and a product. Background Art
[0002] The cloud computing scenario is an Internet-based computing model that centrally manages computing resources (such as servers, storage, networks, etc.) and provides them to users in the form of services. In the cloud computing scenario, users can obtain and release resources at any time according to their own needs, thereby flexibly adjusting the scale of computing resources and realizing elastic expansion or contraction of resources. The elastic computing system abstracts physical resources into a virtual resource pool through virtualization technology, and users can obtain computing instances, storage capacity, network bandwidth and other resources from the resource pool on demand. In practical applications, elastic computing systems are widely used in big data processing, artificial intelligence training, Web services and other fields. For example, during the AI model training process, GPU resources can be dynamically allocated according to the complexity of the training task.
[0003] However, in complex scenarios, users' demand for resources often presents volatility and uncertainty. It is difficult to achieve flexible prediction of resource demand only through static configuration rules, which can easily lead to over-allocation or under-supply of resources in actual operations, affecting service quality. Therefore, a more flexible and accurate resource scheduling method is urgently needed. Summary of the invention
[0004] In view of this, an embodiment of this specification provides a resource scheduling method. One or more embodiments of this specification also relate to a resource scheduling method applied to a model training scenario, an elastic computing system, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a resource scheduling method is provided, including: Obtain resource information of system resources in the elastic computing system and event information of system events in the elastic computing system; Align resource information and event information with timestamps to build a time series of system events, where the time series includes the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp; Based on the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp, cross-modal time series fusion processing is performed to obtain time series correlation features, and based on the time series correlation features, the resource scheduling strategy of future system events is predicted; Based on the resource scheduling strategy, system resources in the elastic computing system are scheduled for future system events.
[0006] In this way, by obtaining the resource information of system resources in the elastic computing system and the event information of system events in the elastic computing system, it is possible to obtain resource usage and system status in real time, providing a reliable data basis for subsequent predictions; by aligning the timestamps of resource information and event information and constructing a time series of system events, it is possible to integrate different types of information and align them in the time dimension, thereby realizing multimodal modeling based on time series; by performing cross-modal time series fusion processing based on timestamps, resource information under timestamps, and event information under timestamps, and obtaining time series correlation features, it is possible to cross-modally fuse time series data and text information in the system, so that the system can simultaneously understand the dynamic changes in resource usage and deep semantic information, thereby deeply understanding the root causes of changes in resource demand, and then predicting resource scheduling strategies for future system events, and improving the accuracy of prediction results; by scheduling system resources in the elastic computing system based on resource scheduling strategies, it is possible to achieve dynamic expansion and contraction, improve the timeliness of resource scheduling, and effectively reduce resource waste. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 is a system architecture diagram of an elastic computing system provided by an embodiment of this specification; Figure 2 is a flow chart of a resource scheduling method provided by an embodiment of this specification; Figure 3 is a process flow chart of a resource scheduling method provided by an embodiment of this specification; Figure 4 is a schematic diagram of a processing process of a resource management unit provided by an embodiment of this specification; Figure 5 It is a structural diagram of a resource scheduling device provided by an embodiment of this specification; Figure 6 It is a structural diagram of a resource scheduling device applied to a model training scenario provided by an embodiment of this specification; Figure 7 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0008] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0009] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0010] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0011] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0012] First, the terms involved in one or more embodiments of this specification are explained.
[0013] Temporal Point Process (TPP): is a statistical modeling framework used to analyze event sequences occurring in a continuous time dimension, which can capture the temporal dynamics and type correlation of events.
[0014] Large Language Model (LLM): A deep neural network model pre-trained on massive text data, with the ability to understand, reason, and generate text.
[0015] Time series-text fusion representation technology: refers to a cross-modal feature extraction method that converts event timestamps into discrete byte tags and jointly encodes them with text descriptions.
[0016] Dynamic scaling: refers to automatically adjusting the size of the cloud computing cluster based on real-time predicted GPU resource demand to achieve dynamic matching of resource supply and demand.
[0017] Fine-grained resource usage pattern: This refers to identifying the micro-fluctuation patterns of GPU resource usage through event time prediction and text semantic association analysis.
[0018] Application Programming Interface (API): can expose the functions or data of a software system in a specific way for use by other software systems.
[0019] Graphics Processing Unit (GPU): In fields such as deep learning and scientific computing, it can process large amounts of data in parallel and is suitable for processing tasks that require highly parallel computing.
[0020] Core processor (CPU, Central Processing Unit): As the core component of the computer, it is mainly responsible for executing the computer's instruction set, performing complex logical operations, data processing, and coordinating the work of various computer components. It is like the "brain" of the computer, controlling the operation of the entire computer system, including the startup of the operating system, the loading and running of applications, the storage and reading of data, etc.
[0021] In this specification, a resource scheduling method is provided. This specification also relates to a resource scheduling method applied to a model training scenario, an elastic computing system, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0022] See also Figure 1 , Figure 1 The system architecture diagram of an elastic computing system provided by an embodiment of the present specification is shown. The elastic computing system 100 includes a resource management unit 102 and a graphics processing unit 104 .
[0023] Resource management unit 102: used to obtain resource information of the graphics processing unit 104 and event information of system events in the elastic computing system 100; align the resource information and event information with timestamps to construct a time series of model training events, wherein the time series includes the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp; based on the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp, perform cross-modal time series fusion processing to obtain time series correlation features, and based on the time series correlation features, predict resource scheduling strategies for future system events; and schedule the graphics processing unit 104 based on the resource scheduling strategy.
[0024] Specifically, the elastic computing system 100 may include one or more graphics processing units 104. The graphics processing unit 104 may be specifically understood as a GPU. In the case where the elastic computing system 100 includes multiple graphics processing units 104, the multiple graphics processing units 104 may constitute a GPU cluster. Scheduling the graphics processing unit 104 may be understood as scheduling the computing resources of the GPU, which may include computing core allocation, video memory allocation, computing frequency adjustment, power state adjustment, etc., which may be specifically determined according to the needs of actual applications.
[0025] By applying this embodiment, by obtaining resource information of system resources and event information of system events in the elastic computing system 100, resource usage and system status can be obtained in real time, providing a reliable data basis for subsequent predictions; by aligning the timestamps of resource information and event information and constructing a time series of system events, different types of information can be integrated and aligned in the time dimension, thereby realizing multimodal modeling based on time series; by performing cross-modal time series fusion processing based on timestamps, resource information under timestamps, and event information under timestamps, time series correlation features can be obtained, and time series data and text information in the system can be cross-modally fused, so that the resource management unit 102 can simultaneously understand the dynamic change rules of resource usage and deep semantic information, thereby deeply understanding the root causes of changes in resource demand, and then predicting resource scheduling strategies for future system events, and improving the accuracy of prediction results; by scheduling the graphics processing unit 104 based on the resource scheduling strategy, more intelligent dynamic expansion and contraction decisions can be realized, and the timeliness of resource scheduling can be improved, thereby effectively reducing resource waste and improving the overall utilization efficiency of the GPU cluster.
[0026] See also Figure 2 , Figure 2 A flow chart of a resource scheduling method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0027] Step 202: Obtain resource information of system resources in the elastic computing system and event information of system events in the elastic computing system.
[0028] The embodiments of this specification can be applied to an elastic computing system with resource allocation and management functions. The elastic computing system can be an elastic computing system in a cloud computing scenario and can be applied to a model training and deployment platform of a private cloud or a public cloud.
[0029] Specifically, an elastic computing system can be understood as a system that can dynamically adjust computing resources according to computing needs, and can flexibly allocate and recycle resources to meet computing needs under different loads, improve resource utilization and system responsiveness. System resources can be understood as various resources available in an elastic computing system, such as hardware resources such as CPU, GPU, memory, disk I / O, network bandwidth, and software resources such as operating system and database. Resource information can be understood as information describing the current status and usage of system resources, such as CPU usage, GPU usage data, memory usage, disk read and write speed, etc. System events can be understood as various operations or state changes that occur in an elastic computing system, such as user requests, task start / end, system failure, etc. Event information can be understood as detailed information related to system events, including the time, type, and resources involved of the event.
[0030] According to an optional implementation, obtaining resource information of system resources in the elastic computing system may include: collecting resource information of system resources through a system monitoring tool or monitoring software.
[0031] Specifically, the system monitoring tool can be understood as the underlying command provided by the operating system, and the monitoring software can be third-party monitoring software that is called through a specific API.
[0032] Optionally, the system monitoring tool or monitoring software may periodically sample and record resource information and store it in a database or log file.
[0033] According to an optional implementation, obtaining event information of system events in the elastic computing system may include: extracting event information of system events from data sources such as system log files, application logs, and message queues.
[0034] Specifically, system log files and application logs can be understood as system operation logs of elastic computing systems. System operation logs can record various event information such as task start, end, error report, exception information, etc. In elastic computing systems, message queues can be used to pass messages between different components. Message queues can store various types of information including system event related information.
[0035] Optionally, in an elastic computing system, when a system event occurs, the relevant components will send a message describing the event to a message queue. These messages usually contain key information such as the type of event (e.g., task state change event, resource allocation event, etc.), the time when the event occurred, and the objects involved (e.g., specific task ID, resource identifier, etc.). To extract event information from a message queue, the corresponding consumer program must first be connected to the message queue. The consumer program filters the messages in the message queue according to pre-set rules (e.g., based on the subject, tag, etc. of the message) to identify messages related to system events. Then, specific event information is parsed from these qualified messages, such as extracting the timestamp of the event occurrence and a detailed description of the event from the content of the message. In this way, the system event information of the elastic computing system can be accurately obtained from the message queue.
[0036] Obtaining resource information of system resources in the elastic computing system and event information of system events in the elastic computing system provides a data basis for subsequent resource scheduling strategy prediction.
[0037] Step 204: align the timestamps of the resource information and the event information to construct a time series of the system event, wherein the time series includes the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp.
[0038] In practical applications, after collecting data from elastic computing systems and obtaining resource information and event information, the data can be cleaned and standardized, and the resource information and event information can be timestamped to build a time series of system events.
[0039] Specifically, timestamp alignment can be understood as matching and synchronizing resource information and event information in chronological order, ensuring that resource information and event information at each time point are corresponding. The time series of system events can be understood as a collection of system resource information and event information arranged in chronological order, and each time point contains corresponding timestamps, resource information, and event information. Timestamps can also be understood as time points, and the specific accuracy can be determined according to the needs of actual applications.
[0040] According to an optional implementation, aligning timestamps of resource information and event information may include: extracting timestamp fields from resource information and event information; sorting resource information and event information according to timestamps, and matching resource information and event information at the same time point or a similar time point based on the timestamp fields.
[0041] In the actual implementation process, you can choose to align resource information and event information at a time accuracy of seconds, milliseconds, etc. according to actual needs.
[0042] According to an optional implementation, constructing a time series of system events may include: for each time stamp, combining the time stamp, resource information under the time stamp, and event information under the time stamp to obtain a time series.
[0043] During the operation of the elastic computing system, the time series will be continuously updated over time. In practical applications, the size of the time series can be maintained as a preset fixed value, such as 1k, 10k, etc., which can be determined according to the needs of the actual application. For example, when the time series size is 1k, the time series can include 1000 groups of timestamps, resource information under the timestamps, and event information under the timestamps. The time series can be understood as a sliding window of fixed size, which continuously slides backward with the subsequent events added, thereby ensuring that each prediction is based on the latest collected data.
[0044] Step 206: Based on the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp, perform cross-modal time series fusion processing to obtain time series correlation features, and based on the time series correlation features, predict resource scheduling strategies for future system events.
[0045] In practical applications, on the basis of constructing time series, cross-modal time series fusion processing can be performed based on the timestamp of system events, resource information under the timestamp, and event information under the timestamp to obtain time series correlation features. Based on the time series correlation features, the resource scheduling strategy of future system events can be predicted.
[0046] Specifically, cross-modal time series fusion processing can be understood as cross-modal fusion of different types of information (such as timestamps, resource information, and event information) and extracting correlation features between different types of information. Time series correlation features can reflect the temporal relationship and change patterns of system resources and events, such as the correlation between the change trend of resource usage and the frequency of event occurrence. Resource scheduling strategies can be understood as plans for rationally allocating and managing system resources based on the current state of the system and predicted future needs. The pre-trained model can be a common model or a large model.
[0047] According to an optional implementation, a pre-trained model may be used to perform cross-modal time series fusion processing to obtain time series correlation features, and based on the time series correlation features, resource scheduling strategies for future system events may be predicted.
[0048] For ordinary models, after obtaining the time series correlation features, the predicted future system events are usually output first. After outputting the prediction results of future events, it is necessary to use additional rules or preset algorithms to determine the corresponding resource scheduling strategies. For example, based on past experience, the amount of resources required for different event types in different scenarios is summarized, and then combined with the predicted future events, appropriate system resources are allocated to these events according to the established rules. After training and learning a large amount of data, the large model can understand complex contextual information and patterns. With its powerful generalization ability, the large model can directly output the corresponding resource scheduling strategy when facing the time series correlation features obtained based on the time series of system events. The large model can perform complex reasoning and calculations within itself, comprehensively consider various factors, and directly give a resource scheduling plan that meets the actual situation of the system. This method is more efficient and intelligent, can better adapt to the complex and changing system environment, and can timely execute accurate resource scheduling strategies without formulating corresponding rules in advance.
[0049] Step 208: Based on the resource scheduling policy, schedule system resources in the elastic computing system for future system events.
[0050] In practical applications, when the resource scheduling strategy for future system events is obtained based on the constructed time series prediction, the system resources in the elastic computing system can be scheduled based on the resource scheduling strategy, thereby realizing dynamic expansion and contraction of the elastic computing system during operation.
[0051] According to an optional implementation, scheduling system resources in the elastic computing system based on a resource scheduling policy may be implemented through a specific API.
[0052] By applying this embodiment, resource information of system resources in an elastic computing system and event information of system events in the elastic computing system are obtained; timestamps of resource information and event information are aligned to construct a time series of system events, wherein the time series includes the timestamp of the system event, resource information under the timestamp, and event information under the timestamp; based on the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp, cross-modal time series fusion processing is performed to obtain time series correlation features, and based on the time series correlation features, resource scheduling strategies for future system events are predicted; based on the resource scheduling strategies, system resources in the elastic computing system are scheduled for future system events.
[0053] In this way, by obtaining the resource information of system resources in the elastic computing system and the event information of system events in the elastic computing system, it is possible to obtain resource usage and system status in real time, providing a reliable data basis for subsequent predictions; by aligning the timestamps of resource information and event information and constructing a time series of system events, it is possible to integrate different types of information and align them in the time dimension, thereby realizing multimodal modeling based on time series; by performing cross-modal time series fusion processing based on timestamps, resource information under timestamps, and event information under timestamps, and obtaining time series correlation features, it is possible to cross-modally fuse time series data and text information in the system, so that the system can simultaneously understand the dynamic changes in resource usage and deep semantic information, thereby deeply understanding the root causes of changes in resource demand, and then predicting resource scheduling strategies for future system events, and improving the accuracy of prediction results; by scheduling system resources in the elastic computing system based on resource scheduling strategies, it is possible to achieve dynamic expansion and contraction, improve the timeliness of resource scheduling, and effectively reduce resource waste.
[0054] In an optional embodiment, the event information of the system event in the elastic computing system includes at least one of task execution information, error report and exception information.
[0055] Specifically, system events can be understood as various operations, state changes or specific situations that occur during the operation of the elastic computing system. These events can reflect the operating status, performance and possible problems of the system. Event information can be understood as detailed description information related to system events. Administrators and systems can often analyze, diagnose and take corresponding treatment measures for system events through event information. Resource information of system resources in the elastic computing system can include GPU usage data, which can specifically include indicators such as video memory usage and computing load. Task execution information can be understood as various information covering the entire process of a task from start to finish, such as the start time, completion time, execution progress, resources used (such as CPU, memory, GPU, etc.), and the execution result (success or failure) of the task. Through this information, the execution efficiency and resource consumption of the task can be understood. Error report can be understood as a report containing detailed error information generated when an error occurs during the operation of the system. Error reports usually record the time when the error occurred, the type of error (such as program crash, memory overflow, file read / write error, etc.), the location of the error (such as the specific line number of code, process ID, etc.), and the possible cause of the error, which helps to quickly locate and solve the problem. Exception information can be understood as information about situations that do not meet normal expectations during system operation. Exception information may not be as serious as an error report, but it will also affect the normal operation of the system, such as the running time of a process exceeds the normal range, resource utilization suddenly fluctuates greatly, etc. Exception information can help discover potential system problems in advance.
[0056] Optionally, the context description information of the system event can be composed through the resource information and the event information.
[0057] By applying this embodiment, by obtaining task execution information, it is beneficial to subsequently mine the deep correlation between task execution status and timing information, thereby helping to predict the task execution time and resource requirements of future system events and reasonably arrange scheduling strategies; by obtaining error reports and exception information, it is beneficial to mine the deep correlation between potential system problems and timing information, thereby improving the accuracy of prediction of abnormal situations and peak loads and improving the accuracy of resource scheduling.
[0058] In an optional embodiment, based on the timestamp of the system event, the resource information under the timestamp and the event information under the timestamp, cross-modal time series fusion processing is performed to obtain time series correlation features, which may include: encoding the timestamp, the resource information under the timestamp and the event information under the timestamp respectively to obtain timestamp features, type features and text content features; encoding the timestamp to obtain timestamp features, encoding the resource information under the timestamp to obtain type features, encoding the event information under the timestamp to obtain text content features; performing attention feature calculation on the cross-modal time series fusion features to extract time series correlation features.
[0059] Specifically, the timestamp can be understood as a mark indicating the time when an event occurs, which is usually a value accurate to a specific time unit (such as seconds or milliseconds). In this embodiment, the timestamp is used to record the specific moment when the system event occurs, and is a key element in constructing a time series. The time series correlation feature can be understood as a feature that reflects the temporal correlation and change law of system events. The time series correlation feature can capture the dynamic relationship between system resources, events and time series, and provide a basis for subsequent predictions and decisions. The resource information under the timestamp may include information related to resource usage, such as resource allocation, release, and errors. Since resource usage often corresponds to a specific event type, the resource information under the timestamp can also be understood as an event type. The event information under the timestamp can include data indicators of resource usage and detailed event information (such as task start, end, error information, abnormal content, etc.). Therefore, the event information under the timestamp can also be understood as the context information of the event.
[0060] According to an optional embodiment, encoding the timestamp to obtain the timestamp feature, encoding the resource information under the timestamp to obtain the type feature, encoding the event information under the timestamp to obtain the text content feature, may include: Encode the timestamp at the byte level to obtain the timestamp feature; encode the resource information under the timestamp in the text dimension to obtain the type feature; encode the event information under the timestamp in the text dimension to obtain the text content feature.
[0061] It should be noted that the purpose of encoding the timestamp at the byte level is to enable the encoded timestamp features to be aligned and integrated with the type features and text content features in the text dimension. Since timestamps are often discrete data, byte-level encoding can be used to obtain timestamp features in the text dimension.
[0062] In the actual implementation process, since the timestamp, resource information under the timestamp, and event information under the timestamp are different types of data, the corresponding specific encoding processes are different. The timestamp features, type features, and text content features obtained by encoding are all features under the text dimension, which can be text tokens.
[0063] According to an optional embodiment, the timestamp feature, type feature and text content feature are subjected to feature fusion processing to obtain a cross-modal temporal fusion feature, which can be implemented in a variety of ways, such as simple concatenation, weighted summation, fusion using a fully connected layer, etc. Taking concatenation as an example, the timestamp feature, type feature and text content feature can be sequentially concatenated into a longer vector to obtain a cross-modal temporal fusion feature.
[0064] According to an optional embodiment, performing attention feature calculation on the cross-modal time series fusion feature and extracting the time series correlation feature may include: using an attention mechanism (such as a multi-head attention mechanism) to calculate the cross-modal time series fusion feature. First, the cross-modal time series fusion feature is input into the attention layer, and the attention weight of each feature element is calculated. Then, the feature elements are weighted and summed according to the attention weight to obtain a feature vector processed by the attention mechanism, that is, the time series correlation feature.
[0065] By applying this embodiment, different types of data can be converted into feature representations under a unified text dimension by encoding timestamps, resource information, and event information respectively, thereby realizing multimodal modeling and facilitating subsequent processing and analysis; by performing feature fusion processing on timestamp features, type features, and text content features, cross-modal time series fusion features are obtained, attention features are calculated on the cross-modal time series fusion features, and time series correlation features are extracted, which can mine the potential correlation between time, resources, and events, thereby overcoming the problem of separation of time series dynamics and text semantic information in the prediction process, and providing more comprehensive information for subsequent predictions and decisions.
[0066] In an optional embodiment, encoding the timestamp to obtain the timestamp feature, encoding the resource information under the timestamp to obtain the type feature, encoding the event information under the timestamp to obtain the text content feature, may include: Discretely encode the timestamp to obtain the timestamp feature; semantically encode the resource information under the timestamp to obtain the type feature; semantically encode the event information under the timestamp to obtain the text content feature.
[0067] According to an optional embodiment, discretely encoding the timestamp to obtain the timestamp feature may include: converting the timestamp into a floating point number; converting the floating point number into a binary representation; and splitting the binary representation to obtain a byte-level encoding representation.
[0068] Exemplarily, the timestamp may be converted into a 32-bit floating point number, and the 32-bit floating point number may be further converted into a 4-byte token.
[0069] According to another optional embodiment, discretely encoding the timestamp to obtain the timestamp feature may also include: dividing the timestamp into different time intervals; mapping each timestamp to a corresponding time interval, and performing one-hot encoding on each time interval. For example, the time interval may be divided by year, month, day, hour, etc.
[0070] According to an optional embodiment, semantic encoding is performed on the resource information under the timestamp to obtain the type feature, which may include: mapping the resource information under the timestamp into a natural language description, and obtaining a text encoding corresponding to the natural language description.
[0071] According to another optional embodiment, semantic encoding is performed on the resource information under the timestamp to obtain the type feature, which may also include: using a hash function to map the resource information under the timestamp to a fixed-length code.
[0072] According to an optional embodiment, semantic encoding of the event information at the timestamp to obtain text content features may include: using a word segmenter to process context text corresponding to the event information at the timestamp to obtain text content features.
[0073] According to another optional embodiment, semantic encoding is performed on the event information at the timestamp to obtain text content features, which may also include: using a pre-trained language model (such as BERT) to encode the event information at the timestamp.
[0074] In the actual implementation process, the timestamp feature obtained by discretely encoding the timestamp is a byte token, and the type feature and text content feature obtained by semantically encoding the resource information under the timestamp and the event information under the timestamp are text tokens.
[0075] By applying this embodiment, the timestamp feature is obtained by discretely encoding the timestamp, and a byte-level encoding strategy can be adopted to encode the timestamp into a byte token; by semantically encoding the resource information under the timestamp to obtain the type feature, and semantically encoding the event information under the timestamp to obtain the text content feature, the resource information and the event information can be encoded into text tokens, which is conducive to realizing cross-modal feature fusion of different types of information in the text dimension. Through multimodal modeling, the deep correlation between different information can be extracted, so that the log text and time series information can be better combined, the root cause of the change in resource demand can be better understood, and the accuracy of prediction and scheduling can be improved.
[0076] In an optional embodiment, predicting resource scheduling strategies for future system events based on timing correlation features may include: predicting the probability distribution of occurrence of different system events based on timing correlation features; predicting future system events based on timing correlation features and event occurrence probability distribution, and generating resource scheduling strategies for future system events.
[0077] Specifically, the time series correlation feature can be understood as a quantitative representation of the correlation and change rules between system resources and events in the time dimension. The time series correlation feature integrates multiple data such as timestamps, resource information and event information, which can reflect the dynamic changes of the system over time and provide a basis for subsequent predictions and decisions. Future system events can be understood as system events that have not yet occurred, but have a certain probability of occurring in the future based on the system's historical data and current status, such as task launch, system failure, resource bottleneck, etc. Future system events can be the next event to occur corresponding to the currently constructed time series, or a set of events containing multiple future events. The occurrence probability distribution can be understood as a probability model that describes the probability of occurrence of future system events. It can present the probability of occurrence of each possible event in mathematical form, help quantify uncertainty, and provide a probabilistic basis for resource scheduling. Resource scheduling strategy can be understood as a resource scheduling plan for future system events, which can ensure the reasonable allocation and use of system resources.
[0078] According to an optional embodiment, predicting the occurrence probability distribution of different system events based on time series correlation characteristics may include: predicting the occurrence probability distribution of different system events based on time series correlation characteristics through the strength prediction layer of the resource prediction model.
[0079] According to another optional embodiment, predicting the occurrence probability distribution of different system events based on time series correlation features may also include: predicting the occurrence probability distribution of different system events based on time series correlation features through a long short-term memory network (LSTM).
[0080] Specifically, the strength prediction layer can be used to calculate and output the probability of occurrence of various possible situations of future system events based on the received time series correlation features. For example, for possible events such as task load changes, resource demand fluctuations, or system failures, the strength prediction layer will give the probability values of these events occurring at different time points, converting the originally complex time series correlation features into a quantitative representation of the possibility of future events.
[0081] In the actual implementation process, the intensity prediction layer will further process and map the input time series correlation features. It may map the data in the feature space to the probability space through a series of mathematical operations and nonlinear transformations, so that the output results meet the requirements of probability distribution, that is, the sum of the probabilities of all possible events is 1, and the probability of each event is between 0 and 1.
[0082] According to an optional embodiment, based on timing correlation characteristics and occurrence probability distribution, predicting future system events and generating resource scheduling strategies for future system events may include: outputting the next event to occur, the resource demand corresponding to the next event to occur, and the resource scheduling strategy based on timing correlation characteristics and occurrence probability distribution through a resource prediction model.
[0083] According to another optional embodiment, based on the time series correlation characteristics and the occurrence probability distribution, predicting future system events and generating resource scheduling strategies for future system events can also be achieved through rule engines, reinforcement learning, integrated learning, etc.
[0084] Alternatively, the rule engine can be understood as a series of rules defined based on the system's task execution logic and requirements. Input the time series correlation features and occurrence probability distribution into the rule engine, and the engine will automatically match the rules that meet the conditions; based on the matched rules, predict future system events and generate corresponding resource scheduling strategies.
[0085] Optionally, reinforcement learning algorithms (such as Q-learning, deep Q network, etc.) can be used to train the agent so that it can learn the optimal decision-making strategy. In actual application, the agent predicts future system events based on the current system state and generates corresponding resource scheduling strategies.
[0086] Optionally, you can also select multiple different types of base learners (such as decision trees, support vector machines, etc.); use training data to train each base learner; and integrate the prediction results of the base learners (such as voting, weighted averaging, etc.) to obtain the final future system event prediction results and resource scheduling strategies.
[0087] Optionally, based on the time series correlation feature, the event type and timestamp of the next event to occur can be output; based on the occurrence probability distribution, the resource demand and resource scheduling strategy corresponding to the next event to occur can be output.
[0088] Specifically, the resource prediction model can be understood as a large model with generalization capabilities, and the next event to occur can be understood as the next predicted event that may occur corresponding to the current time series at the current time point. The resource scheduling strategy may include an execution operation for the next event to occur.
[0089] It should be noted that the next event to occur may change with the passage of time. For example, if the last event in the time series does not change (that is, no new events are collected), the time interval between the timestamp of the last event and the current time is 10 seconds and the time interval between the timestamp of the last event and the current time is 40 seconds, the predicted next event to occur may be different.
[0090] According to another optional embodiment, when the resource prediction model is a common model rather than a large model, based on the timing correlation characteristics and the occurrence probability distribution, future system events are predicted, and a resource scheduling strategy for future system events is generated. It may also include: through the resource prediction model, based on the timing correlation characteristics and the occurrence probability distribution, outputting the event type and timestamp of the next event to occur; according to the event type and timestamp of the next event to occur, based on the mapping relationship between the event type and the resource requirement, determining the resource requirement corresponding to the next event to occur; and determining the corresponding resource scheduling strategy based on the event type, timestamp and resource requirement of the next event to occur.
[0091] By applying this embodiment, the probability distribution of occurrence of future system events is predicted based on the time series correlation characteristics, and the next possible system event can be determined based on the occurrence probability distribution, so that resources can be planned in advance for the system event and corresponding resource scheduling strategies can be generated, dynamic resource allocation and scheduling can be achieved, and the timeliness and flexibility of resource allocation can be improved. In addition, through the time series correlation characteristics, more accurate results can be predicted based on the potential correlation between time series information and resources and events, thereby improving the accuracy of resource scheduling.
[0092] In an optional embodiment, the resource scheduling strategy includes the timestamp, event type and resource requirement of future system events; generating a resource scheduling strategy for future system events may include: predicting the timestamp and event type of future system events based on timing correlation features; and predicting the resource requirement of future system events based on the occurrence probability distribution.
[0093] Specifically, the timestamp of a future system event can be understood as the specific time at which a future system event is predicted to occur, which can be accurate to seconds, milliseconds and other time units. It is used to locate events in the time dimension and help the system prepare for resource allocation in advance. The event type of a future system event can be understood as the classification of system events that may occur in the future, such as task start, task end, system failure, resource expansion request, etc. Different event types have different resource requirements and processing methods. The resource demand of future system events can be understood as the number of various system resources (such as the number of GPU cores, the number of CPU cores, memory size, disk storage space, network bandwidth, etc.) required during the execution of the predicted future system event.
[0094] According to an optional implementation, predicting the timestamp of future system events and the event type of future system events based on timing correlation features may include: using a resource prediction model to predict the timestamp of future system events and the event type of future system events based on timing correlation features.
[0095] In the actual implementation process, the resource prediction model can output the timestamp token and event type token of future system events based on the time series correlation features. In addition, it can also output the description information token for future system events, etc.
[0096] According to an optional implementation, predicting resource requirements of future system events based on occurrence probability distribution may include: using a resource prediction model to predict resource requirements of future system events based on occurrence probability distribution.
[0097] According to an optional implementation, generating a resource scheduling strategy for future system events based on timestamp, event type and resource demand may include: utilizing a resource prediction model to output a resource scheduling strategy for future system events based on timing correlation features and occurrence probability distribution.
[0098] When the resource prediction model is a large model, the resource prediction model can output the corresponding resource scheduling strategy based on its own generalization ability while outputting the prediction results such as timestamp, event type and resource demand. When the resource prediction model is a common model, the system can determine the matching resource scheduling strategy based on the prediction results such as timestamp, event type and resource demand output by the resource prediction model.
[0099] Optionally, the timestamp, event type, and resource demand of the predicted future system events can be combined to form a complete resource scheduling policy data structure. For example, a table or JSON format data containing these three fields can be constructed; the generated resource scheduling policy is stored in a database or file system for subsequent system reading and execution. At the same time, the policy can also be displayed to the system administrator in a visual way for easy viewing and adjustment.
[0100] By applying this embodiment, the occurrence time and event type of future events can be accurately predicted through the timing correlation characteristics, so that the required amount of resources can be accurately allocated to specific types of events at the right time, avoiding over-allocation or under-allocation of resources, improving resource utilization, and reducing resource waste; by predicting the resource demand of future system events based on the probability distribution of occurrence, the system can arrange maintenance tasks and prepare backup resources in advance, reduce the impact of system failures on operation, and enhance the stability and reliability of the system.
[0101] In cloud computing scenarios, based on resource prediction results and resource scheduling strategies for future events, corresponding resource expansion and contraction operations can be performed at the appropriate time, thereby responding in advance to possible resource over-allocation or undersupply situations.
[0102] In an optional embodiment, the resource scheduling strategy includes resource requirements of future system events; scheduling system resources in the elastic computing system based on the resource scheduling strategy may include: When the resource demand exceeds the current resource amount of the system resources, perform resource expansion operation on the system resources; When the resource demand does not exceed the current resource amount of the system resources, a resource reduction operation is performed on the system resources.
[0103] Specifically, the current amount of system resources can be understood as the total amount of various resources that the elastic computing system actually has and can use at the current moment. Such as the number of currently available GPU cores, the remaining memory size, free disk space, etc. Resource expansion operations can be understood as a series of actions taken to increase system resources to meet demand when the system predicts that the resource demand for future events exceeds the current amount of resources. Specifically, it can include adding physical hardware devices (such as increasing the number of GPU cores, increasing server memory modules, adding new hard disks, etc.), or leasing additional virtual resources from cloud service providers (such as adding virtual machine instances). Resource reduction operations can be understood as operations to reduce system resources to avoid idle waste of resources when the resource demand does not exceed the current amount of resources. For example, reclaiming excess GPU cores, shutting down some idle virtual machines, releasing excess memory space, or reducing rented cloud storage capacity.
[0104] According to an optional implementation, when the resource demand exceeds the current resource amount of the system resources, performing a resource expansion operation on the system resources may include: when the resource demand exceeds the current resource amount of the system resources, using a specific API interface of the cloud service provider to expand the resources.
[0105] It should be noted that after adding new resources, the resources need to be configured so that they can be recognized and used normally by the system. For example, the newly added hard disk needs to be partitioned, formatted, and mounted to the specified directory of the operating system; the newly started virtual machine needs to be configured with the network and the necessary software dependencies installed to ensure that the new resources are seamlessly integrated with the existing system resources.
[0106] According to another optional implementation, the target time point or target time period for executing the resource expansion operation may be determined according to the timestamp of the future event, and the resource expansion operation may be executed when the target time point or target time period is reached.
[0107] According to an optional implementation, taking the GPU resource as an example, when the resource demand does not exceed the current resource demand of the system resource, performing a resource scaling operation on the system resource may include: determining the resource to be recycled based on the resource scheduling policy; terminating the non-critical tasks or processes related to the resource to be recycled, and releasing the GPU's video memory and computing resources. Furthermore, the video memory and computing resources may be marked as available so as to be reallocated to other tasks in need.
[0108] According to another optional implementation, a target time point or a target time period for executing the resource reduction operation may be determined according to the timestamp of the future event, and the resource reduction operation may be executed when the target time point or the target time period is reached.
[0109] By applying this embodiment, through dynamic resource expansion and contraction operations, it is possible to ensure that system resources always match actual demand. When the resource demand is high, the capacity is expanded in time to meet the system operation requirements; when the demand is low, the capacity is contracted to avoid idle resources and waste, which greatly improves the resource utilization efficiency and reduces the operating cost.
[0110] In an optional embodiment, after scheduling the system resources in the elastic computing system based on the resource scheduling policy, it may also include: collecting resource information of the system resources in future system events to obtain execution monitoring results for the resource scheduling policy.
[0111] Specifically, the resource information of future system events can be understood as the resource usage of system resources when future system events occur, which can specifically include data indicators such as GPU usage. The execution monitoring results can be understood as the conclusions on the actual execution effect of the resource scheduling strategy obtained by collecting resource information of system resources in future system events, sorting and analyzing them, which can reflect whether the resource scheduling strategy effectively meets the resource requirements of system events, and whether there is any waste or shortage of resources.
[0112] According to an optional implementation, collecting resource information of system resources in future system events to obtain execution monitoring results for resource scheduling strategies may include: determining monitoring indicators; collecting resource information corresponding to the monitoring indicators; comparing the resource information with the predicted resource demand in the resource scheduling strategy to obtain execution monitoring results for the resource scheduling strategy.
[0113] Optionally, determining the monitoring indicator may include determining a key resource indicator. In practical applications, the key resource indicator may be determined according to specific requirements in different scenarios. For example, in a cloud computing scenario, the key resource indicator applied to a model training scenario may be a GPU utilization indicator.
[0114] Furthermore, the monitoring indicators can be further refined according to different types of future system events, such as monitoring the response time of resource allocation for task start events, monitoring the stability and sustainability of resources during task operation, and monitoring the timeliness of resource release for task end events.
[0115] Optionally, resource information corresponding to the monitoring indicators can be collected by using system tools, or by deploying third-party monitoring software and using cloud platform APIs.
[0116] Optionally, after collecting the resource information corresponding to the monitoring indicators, the collected raw resource data can be cleaned to remove abnormal values, duplicate data and erroneous data. Then the data can be sorted according to the time sequence, event type and other dimensions to facilitate subsequent analysis.
[0117] Furthermore, the collected resource information can be compared with the expected resource usage in the resource scheduling strategy. The deviation between the actual resource usage and the expected resource usage can be calculated to analyze whether the resource utilization is within a reasonable range and whether the resource scheduling strategy has achieved the expected goal. Based on the comparative analysis results, an execution monitoring result report can also be generated. The report can include various indicator data of resource usage, comparison with expectations, advantages and disadvantages of the resource scheduling strategy, and suggestions for improvement of the deficiencies.
[0118] According to another optional implementation, resource information of system resources in future system events is collected to obtain execution monitoring results for resource scheduling strategies, which may also include: using machine learning algorithms to perform anomaly detection on system resource information and evaluating the execution effect of resource scheduling strategies.
[0119] Optionally, historical resource information of system resources in future system events can be collected and divided into training sets and test sets. Standardize the data to ensure data consistency. Select appropriate anomaly detection algorithms, such as isolation forest, local anomaly factor (LOF), etc. Use the training set data to train the model. Input the collected real-time resource information into the trained model to detect whether there are anomalies. If anomalies are detected, it indicates that there may be problems with the resource scheduling strategy. Further, the time, type, and degree of anomaly occurrence can be analyzed to generate an execution monitoring result report.
[0120] By applying this embodiment, by collecting resource information and generating execution monitoring results, the effect of the resource scheduling strategy in actual application can be intuitively understood. It is clear whether the strategy successfully guarantees the resource requirements of future system events, providing a strong basis for the optimization and adjustment of the strategy. Based on the execution monitoring results, the resource scheduling strategy can be optimized to make resource allocation more accurate and reasonable, improve resource utilization, and reduce costs. By real-time monitoring of the operating status of system resources in future system events, it is also beneficial to timely discover potential system failures or performance bottlenecks. Once abnormal fluctuations or exceeding the normal range of resource usage are found, measures can be taken quickly to adjust, avoid system failures due to resource problems, and ensure the stable operation of the system.
[0121] In an optional embodiment, based on the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp, cross-modal time series fusion processing is performed to obtain time series correlation features, and based on the time series correlation features, the resource scheduling strategy of future system events is predicted. It can include: inputting the time series of the system event into a resource prediction model, and performing cross-modal time series fusion processing based on the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp by the resource prediction model to obtain time series correlation features, and based on the time series correlation features, predicting the resource scheduling strategy of future system events.
[0122] Specifically, the resource prediction model is a pre-trained model, which can be a common model or a large language model.
[0123] According to an optional implementation, the resource prediction model can be trained through the following steps: Obtaining a sample time series and a label resource scheduling policy corresponding to the sample time series, wherein the sample time series includes a sample timestamp of a system event of the elastic computer system, resource information under the sample timestamp, and event information under the sample timestamp; The sample timestamp, resource information under the sample timestamp, and event information under the sample timestamp are input into the initial resource prediction model, and cross-modal time series fusion processing is performed to obtain sample time series correlation features. Based on the sample time series correlation features, the predicted resource scheduling strategy for future system events is predicted; based on the label resource scheduling strategy and the predicted resource scheduling strategy, the initial resource prediction model is trained to obtain a trained resource prediction model.
[0124] Furthermore, predicting a prediction resource scheduling strategy for future system events based on sample time series correlation features may include: predicting a prediction resource scheduling strategy for future system events and an event probability distribution of future system events based on sample time series correlation features.
[0125] Accordingly, based on the label resource scheduling strategy and the prediction resource scheduling strategy, training the initial resource prediction model to obtain the trained resource prediction model can include: calculating the first loss value based on the label resource scheduling strategy and the prediction resource scheduling strategy; calculating the second loss value based on the event probability distribution and the sample time series; training the initial resource prediction model according to the first loss value and the second loss value to obtain the trained resource prediction model.
[0126] According to an optional implementation, the reasoning process of the resource prediction model may include: receiving a time series of system events, performing cross-modal timing fusion processing based on the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp, to obtain timing correlation features; based on the timing correlation features, predicting the resource scheduling strategy for future system events.
[0127] In practical applications, the resource prediction model can output the prediction results for the sample time series, and calculate the first loss value based on the prediction results and labels; the resource prediction model can also output the event probability distribution corresponding to the sample time series, and calculate the second loss value based on the event probability distribution and the target event probability distribution corresponding to the sample time series, so as to complete the model training by adjusting parameters based on the first loss value and the second loss value.
[0128] By applying this embodiment, through pre-training the resource prediction model, the time point process can be deeply integrated with the model, breaking through the limitations of the separation of time series data and text information in the traditional GPU resource scheduling system. By uniformly encoding multimodal information such as GPU usage changes, system logs and error reports into event sequences, and with the help of a specially designed byte-level time encoding mechanism, the system can simultaneously understand the dynamic changes in resource usage and deep semantic information. This fusion processing method not only overcomes the defects of existing solutions that rely only on a single numerical prediction or static rules, but also can deeply understand the root causes of changes in resource demand by analyzing text information such as task descriptions and error logs. The model can accurately capture abnormal situations and load peaks during training, predict changes in resource demand in advance, and thus achieve more intelligent dynamic expansion and contraction decisions, improve the accuracy of resource prediction and the timeliness of scheduling, effectively reduce resource waste, and improve the overall utilization efficiency of the GPU cluster.
[0129] The following combination Figure 3 , taking the application of the resource scheduling method provided in this specification in the model training scenario as an example, the resource scheduling method is further explained. Figure 3 A process flow chart of a resource scheduling method provided by an embodiment of the present specification is shown, which is applied to a resource management unit of an elastic computing system and specifically includes the following steps.
[0130] Step 302: Obtain resource information of a graphics processing unit in the elastic computing system and event information of a model training event in the elastic computing system.
[0131] Step 304: align the timestamps of the resource information and the event information to construct a time series of the model training events, wherein the time series includes the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp.
[0132] Step 306: Based on the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp, perform cross-modal time series fusion processing to obtain time series correlation features, and based on the time series correlation features, predict the resource scheduling strategy for future model training events.
[0133] Step 308: Based on the resource scheduling strategy, schedule the graphics processing unit in the elastic computing system for future model training events.
[0134] Specifically, the resource management unit can be understood as a key module or component responsible for unified management and scheduling of various resources in the elastic computing system, and can be specifically used to manage the graphics processing unit. The elastic computing system may include one or more graphics processing units. The graphics processing unit can be specifically understood as a GPU. In the case where the elastic computing system includes multiple graphics processing units, multiple graphics processing units can constitute a GPU cluster in the elastic computing system. Scheduling the graphics processing unit can be understood as scheduling the computing resources of the GPU, which can include computing core allocation, video memory allocation, computing frequency adjustment, and power state adjustment, etc., which can be determined according to the needs of actual applications. The model training event can be understood as any event that may occur in the process of processing the model training task using the elastic computing system, including a new training task start event, a training task end event, a training task pause event, a training task recovery event, a training task abnormal interruption event (such as caused by hardware failure, software crash), a training task resource shortage alarm event, a training data loading failure event, a performance bottleneck triggering event during training (such as memory overflow, network delay is too high), and so on.
[0135] It should be noted that the specific implementation of steps 302 to 308 is the same as the specific implementation of steps 202 to 208 above. The specific implementation of steps 302 to 308 can refer to steps 202 to 208, and this specification will not repeat them here.
[0136] By applying this embodiment, multimodal information such as GPU usage changes, system logs, and error reports are uniformly encoded into an event sequence, and with the help of a specially designed byte-level time encoding mechanism, the system can simultaneously understand the dynamic changes in resource usage and deep semantic information. This fusion processing method not only overcomes the defects of existing solutions that rely only on a single numerical prediction or static rules, but also can deeply understand the root causes of changes in resource demand by analyzing text information such as task descriptions and error logs. The model can accurately capture abnormal situations and load peaks during training, predict changes in resource demand in advance, and thus achieve more intelligent dynamic expansion and contraction decisions, improve the accuracy of resource prediction and the timeliness of scheduling, effectively reduce resource waste, and improve the overall utilization efficiency of the GPU cluster.
[0137] See also Figure 4 , Figure 4 The resource management unit 400 includes a processing layer 402 , an output layer 404 and an execution layer 406 .
[0138] Processing layer 402: used to collect GPU usage data and system operation log data in the elastic computing system; clean and standardize the GPU usage data and system operation log data, and align them in time series; build an event sequence based on timestamp, event type and event description information; encode the timestamp, event type and event description respectively to obtain timestamp features, type features and text content features, and input the timestamp features, type features and text content features into the resource prediction model; the resource prediction model performs cross-modal feature fusion processing on the timestamp features, type features and text content features to obtain time series correlation features, and obtains event probability distribution based on the time series correlation features.
[0139] Specifically, GPU usage data can include indicators such as video memory usage and computing load; system operation log data can include event information such as task start, end, and error, as well as error reports and exception information, etc. In the constructed event sequence, each event contains three core elements: timestamp, which is used to record the exact time when the event occurred; event type: which can reflect information such as resource allocation, release, and error; and event description information, which contains detailed context information of the event.
[0140] Optionally, the resource prediction model can be understood as a large language model that has been trained. The resource prediction model may include a pre-trained encoder, decoder, hidden state extraction layer, and intensity prediction layer. Optionally, the decoder may use a QwenLM decoder, or other decoders may be used according to the needs of actual applications, and this specification does not make any limitation on this.
[0141] In practical applications, the resource prediction model can use an encoder to adopt a byte-level encoding strategy for timestamps to convert 32-bit floating-point numbers into 4-byte tokens; map different types of system events into natural language descriptions to obtain text tokens corresponding to the event types; and use the built-in word segmenter of the large language model to process log texts to obtain text tokens corresponding to event description information.
[0142] Furthermore, the resource prediction model can perform alignment and fusion processing on the timestamp features, type features and text content features in the text dimension to obtain cross-modal temporal fusion features, and then extract the temporal dependency between time, type and contextual text content in the cross-modal temporal fusion features through the hidden state extraction layer to obtain temporal correlation features. The resource prediction model can also estimate the probability distribution of future events through the intensity prediction layer to obtain event probability distribution.
[0143] Output layer 404: used to predict the next event and resource demand according to the time dependency, obtain the next event prediction result and resource demand prediction result, and output dynamic scheduling decision based on the next event prediction result and resource demand prediction result.
[0144] Specifically, the next event prediction result may include the timestamp and event type of the next event. The resource demand prediction result may include the resource demand corresponding to the next event. The dynamic scheduling decision may include the resource scheduling operation to be performed for the next event, which may specifically include resource expansion or resource reduction operations.
[0145] Execution layer 406: When a high load is predicted, trigger resource expansion operations in advance; when a load reduction is predicted, perform resource recovery operations in a timely manner; monitor the policy execution effect for the next event.
[0146] For example, in a large-scale model training scenario, the system may encounter the following sequence of events: Receive a new training task start event, including information such as model scale and dataset size; The event of rapid increase in GPU usage is detected; Capture the warning log caused by insufficient video memory; The resource management unit automatically predicts the next resource demand trend; Start new GPU resource allocation in advance before resource shortage occurs; After the training task is completed, it is predicted that the resource demand will decrease and resources will be automatically recycled.
[0147] By applying this embodiment, by deeply integrating the time point process with the large language model, it is possible to break through the limitations of the separation of time series data and text information in the traditional GPU resource scheduling system. By uniformly encoding multimodal information such as GPU usage changes, system logs and error reports into event sequences, and using a specially designed byte-level time encoding mechanism, the system can simultaneously understand the dynamic changes in resource usage and deep semantic information. This fusion processing method not only overcomes the defects of existing solutions that rely only on a single numerical prediction or static rules, but also can deeply understand the root causes of changes in resource demand by analyzing text information such as task descriptions and error logs. The model can accurately capture abnormal situations and load peaks during training, predict changes in resource demand in advance, and thus achieve more intelligent dynamic expansion and contraction decisions, improve the accuracy of resource prediction and the timeliness of scheduling, effectively reduce resource waste, and improve the overall utilization efficiency of the GPU cluster.
[0148] Corresponding to the above method embodiment, this specification also provides a resource scheduling device embodiment, Figure 5FIG. 1 shows a schematic diagram of a resource scheduling device provided by an embodiment of the present specification. Figure 5 As shown, the device comprises: The first acquisition module 502 is configured to acquire resource information of system resources in the elastic computing system and event information of system events in the elastic computing system.
[0149] The first construction module 504 is configured to align the timestamps of the resource information and the event information to construct a time sequence of the system event, wherein the time sequence includes the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp.
[0150] The first prediction module 506 is configured to perform cross-modal time series fusion processing based on the timestamp of the system event, the resource information under the timestamp and the event information under the timestamp, obtain time series correlation features, and predict the resource scheduling strategy of future system events based on the time series correlation features.
[0151] The first scheduling module 508 is configured to schedule system resources in the elastic computing system for future system events based on a resource scheduling policy.
[0152] Optionally, the first prediction module 506 is further configured to: encode the timestamp to obtain timestamp features, encode the resource information under the timestamp to obtain type features, encode the event information under the timestamp to obtain text content features; perform feature fusion processing on the timestamp features, type features and text content features to obtain cross-modal time series fusion features; perform attention feature calculation on the cross-modal time series fusion features to extract time series correlation features.
[0153] Optionally, the first prediction module 506 is further configured to: discretely encode the timestamp to obtain timestamp features; semantically encode the resource information under the timestamp to obtain type features; and semantically encode the event information under the timestamp to obtain text content features.
[0154] Optionally, the first prediction module 506 is further configured to: predict the occurrence probability distribution of different system events based on the time series correlation characteristics; predict future system events based on the time series correlation characteristics and the occurrence probability distribution, and generate resource scheduling strategies for future system events.
[0155] Optionally, the resource scheduling strategy includes the timestamp, event type and resource requirement of future system events; the first prediction module 506 is further configured to: predict the timestamp and event type of future system events based on time series correlation characteristics; and predict the resource requirement of future system events based on the occurrence probability distribution.
[0156] Optionally, the resource scheduling strategy includes the resource demand of future system events; the first scheduling module 508 is further configured to: when the resource demand exceeds the current resource demand of the system resources, perform a resource expansion operation on the system resources; when the resource demand does not exceed the current resource demand of the system resources, perform a resource reduction operation on the system resources.
[0157] Optionally, the resource scheduling device further includes a monitoring module configured to collect resource information of system resources in future system events and obtain execution monitoring results for the resource scheduling strategy.
[0158] Optionally, the event information of the system event in the elastic computing system includes at least one of task execution information, error report and exception information.
[0159] Optionally, the first prediction module 506 is further configured to: input the time series of system events into the resource prediction model, perform cross-modal time series fusion processing based on the timestamp of the system event, the resource information under the timestamp, and the event information under the timestamp, obtain time series correlation features, and predict the resource scheduling strategy of future system events based on the time series correlation features.
[0160] By applying this embodiment, system resources in the elastic computing system are scheduled based on the resource scheduling strategy, so that dynamic expansion and contraction can be achieved and the timeliness of resource scheduling can be improved, thereby effectively reducing resource waste.
[0161] The above is a schematic scheme of a resource scheduling device of this embodiment. It should be noted that the technical scheme of the resource scheduling device and the technical scheme of the resource scheduling method described above are of the same concept, and the details not described in detail in the technical scheme of the resource scheduling device can be found in the description of the technical scheme of the resource scheduling method described above.
[0162] Corresponding to the above method embodiment, this specification also provides a resource scheduling device embodiment applied to a model training scenario, Figure 6 FIG. 1 shows a schematic diagram of a resource scheduling device for a model training scenario provided by an embodiment of the present specification. Figure 6 As shown, the device comprises: The second acquisition module 602 is configured to acquire resource information of a graphics processing unit in the elastic computing system and event information of a system event in the elastic computing system.
[0163] The second construction module 604 is configured to align the timestamps of the resource information and the event information, and construct a time series of the model training events, wherein the time series includes the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp.
[0164] The second prediction module 606 is configured to perform cross-modal time series fusion processing based on the timestamp of the model training event, the resource information under the timestamp and the event information under the timestamp, obtain the time series correlation features, and predict the resource scheduling strategy of future system events based on the time series correlation features.
[0165] The second scheduling module 608 is configured to schedule the graphics processing unit in the elastic computing system based on the resource scheduling policy.
[0166] By applying this embodiment, multimodal information such as GPU usage changes, system logs, and error reports are uniformly encoded into an event sequence, and with the help of a specially designed byte-level time encoding mechanism, the system can simultaneously understand the dynamic changes in resource usage and deep semantic information. This fusion processing method not only overcomes the defects of existing solutions that rely only on a single numerical prediction or static rules, but also can deeply understand the root causes of changes in resource demand by analyzing text information such as task descriptions and error logs. The model can accurately capture abnormal situations and load peaks during training, predict changes in resource demand in advance, and thus achieve more intelligent dynamic expansion and contraction decisions, improve the accuracy of resource prediction and the timeliness of scheduling, effectively reduce resource waste, and improve the overall utilization efficiency of the GPU cluster.
[0167] The above is a schematic scheme of a resource scheduling device applied to a model training scenario in this embodiment. It should be noted that the technical scheme of the resource scheduling device applied to the model training scenario and the technical scheme of the resource scheduling method applied to the model training scenario belong to the same concept, and the details of the technical scheme of the resource scheduling device applied to the model training scenario that are not described in detail can be found in the description of the technical scheme of the resource scheduling method applied to the model training scenario.
[0168] Figure 7 The block diagram of a computing device 700 according to an embodiment of the present specification is shown. The components of the computing device 700 include but are not limited to a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and the database 750 is used to store data.
[0169] The computing device 700 also includes an access device 740 that enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).
[0170] In one embodiment of the present specification, the above components of the computing device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 7 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0171] The computing device 700 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 700 may also be a mobile or stationary server.
[0172] The processor 720 is used to execute the following computer executable instructions, which implement the steps of the above method when executed by the processor.
[0173] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above method.
[0174] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which can implement the steps of the above method when executed by a processor.
[0175] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above method.
[0176] An embodiment of the present specification also provides a computer program product, including a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0177] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the above method belong to the same concept, and the details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of the above method.
[0178] The above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0179] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0180] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0181] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0182] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A resource scheduling method, comprising: Acquire resource information of system resources in the elastic computing system and event information of system events in the elastic computing system; Performing time stamp alignment on the resource information and the event information to construct a time sequence of the system event, wherein the time sequence includes the time stamp of the system event, the resource information under the time stamp, and the event information under the time stamp; Based on the timestamp of the system event, the resource information under the timestamp and the event information under the timestamp, a cross-modal time series fusion process is performed to obtain time series correlation features, and based on the time series correlation features, a resource scheduling strategy for future system events is predicted; Based on the resource scheduling policy, system resources in the elastic computing system are scheduled according to the future system event.
2. According to the method of claim 1, the performing cross-modal time series fusion processing based on the timestamp of the system event, the resource information under the timestamp and the event information under the timestamp to obtain the time series correlation feature comprises: Encode the timestamp to obtain timestamp features, encode resource information under the timestamp to obtain type features, and encode event information under the timestamp to obtain text content features; Performing feature fusion processing on the timestamp feature, the type feature, and the text content feature to obtain a cross-modal temporal fusion feature; Attention features are calculated on the cross-modal temporal fusion features to extract the temporal correlation features.
3. According to the method of claim 2, encoding the timestamp to obtain timestamp features, encoding resource information under the timestamp to obtain type features, encoding event information under the timestamp to obtain text content features, comprises: Discretely encode the timestamp to obtain a timestamp feature; Performing semantic encoding on the resource information under the timestamp to obtain type features; The event information under the timestamp is semantically encoded to obtain text content features.
4. The method according to any one of claims 1 to 3, wherein the resource scheduling strategy for predicting future system events based on the time series correlation characteristics comprises: Based on the time series correlation characteristics, predict the occurrence probability distribution of different system events; Based on the time series correlation characteristics and the occurrence probability distribution, the future system event is predicted, and a resource scheduling strategy for the future system event is generated.
5. The method according to claim 4, wherein the resource scheduling strategy includes a timestamp, event type, and resource requirement of the future system event; The resource scheduling strategy for generating the future system event includes: Predicting the timestamp of the future system event and the event type of the future system event based on the time series correlation feature; Based on the occurrence probability distribution, the resource demand of the future system event is predicted.
6. The method according to claim 1, wherein the resource scheduling strategy includes resource requirements of the future system events; The scheduling of system resources in the elastic computing system based on the resource scheduling policy includes: When the resource demand exceeds the current resource amount of the system resource, performing a resource expansion operation on the system resource; When the resource demand does not exceed the current resource amount of the system resource, a resource scaling operation is performed on the system resource.
7. The method according to claim 1, after scheduling the system resources in the elastic computing system based on the resource scheduling policy, further comprising: Resource information of the system resources in the future system events is collected to obtain execution monitoring results of the resource scheduling strategy. 8 . The method according to claim 1 , wherein the event information of the system event in the elastic computing system comprises at least one of task execution information, error report and exception information.
9. The method according to claim 1, performing cross-modal time series fusion processing based on the timestamp of the system event, the resource information under the timestamp and the event information under the timestamp to obtain time series correlation features, and predicting resource scheduling strategies for future system events based on the time series correlation features, including: The time series of the system events is input into a resource prediction model, and the resource prediction model performs cross-modal time series fusion processing based on the timestamp of the system events, the resource information under the timestamp, and the event information under the timestamp to obtain time series correlation features, and based on the time series correlation features, predicts the resource scheduling strategy of future system events.
10. A resource scheduling method applied to a model training scenario, applied to a resource management unit of an elastic computing system, comprising: Obtain resource information of a graphics processing unit in an elastic computing system and event information of a model training event in the elastic computing system; Performing timestamp alignment on the resource information and the event information to construct a time series of the model training event, wherein the time series includes the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp; Based on the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp, a cross-modal time series fusion process is performed to obtain time series correlation features, and based on the time series correlation features, a resource scheduling strategy for future model training events is predicted; Based on the resource scheduling strategy, a graphics processing unit in the elastic computing system is scheduled for the future model training event.
11. An elastic computing system comprising a resource management unit and a graphics processing unit; The resource management unit is used to obtain resource information of the graphics processing unit and event information of model training events in the elastic computing system; The resource information and the event information are timestamped to construct a time series of the model training event, wherein the time series includes the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp; based on the timestamp of the model training event, the resource information under the timestamp, and the event information under the timestamp, cross-modal time series fusion processing is performed to obtain time series correlation features, and based on the time series correlation features, a resource scheduling strategy for future model training events is predicted; based on the resource scheduling strategy, the graphics processing unit is scheduled for the future model training event.
12. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Dynamic interval elastic capacity expansion and contraction method and system based on long-time sequence prediction
CN117056021A
Node resource prediction method and system based on time sequence intervention analysis model
CN117130882A
Elastic GPU management method based on space-time diagram network resource prediction and reinforcement learning
CN118606039A
System resource dynamic allocation method and optimization system in cloud network convergence environment
CN119094612A
Cross-cloud resource scheduling method and system based on deep learning
CN119299519A