Task processing duration prediction method, apparatus, device and storage medium
By introducing a computing node group model in the 6G network, combining computing power information and task information, the task processing time prediction problem of heterogeneous computing nodes is solved, and more efficient task allocation and management is achieved.
Patent Information
- Application Number
- PCT/CN2024/141759
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-24
- Publication Date
- 2025-07-03
AI Technical Summary
In 6G network, due to the heterogeneity of cloud, edge and end execution environments, the CPU architecture, main frequency and accelerator instruction set of the computing node are heterogeneous, making it difficult to accurately predict the task processing time, affecting the task allocation efficiency.
By introducing a model corresponding to the computing node group, combining the computing power information and task information of the computing node group, the task processing time of the computing node group is predicted and the prediction accuracy is improved.
It realizes accurate prediction of the task processing time of computing node group, supports more efficient task allocation and management, and improves the management, orchestration and scheduling capabilities of computing nodes.
Smart Images

Figure CN2024141759_03072025_PF_FP_ABST
Abstract
Description
Method, device, equipment and storage medium for predicting task processing time
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311873115.7 and application name “Method, device, equipment and storage medium for predicting task processing time”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communication technology, and in particular to a method, apparatus, device, and storage medium for predicting task processing time. Background Art
[0003] With the development of artificial intelligence (AI), intelligent services are being widely applied to various intelligent business scenarios. This has led to a dramatic increase in the diversity of intelligent services and computing power requirements. To achieve the goal of intelligent and inclusive sixth-generation mobile networks (6G) and provide users with ubiquitous intelligent services, 6G networks must integrate computing nodes across the cloud, edge, and end devices as computing resources for management, orchestration, and scheduling.
[0004] Due to the diversity of cloud, edge, and end execution environments in 6G networks, the central processing unit (CPU) architecture, main frequency, and instruction sets of various accelerators of the computing nodes that execute tasks are significantly heterogeneous. These heterogeneous computing nodes have different computing capabilities, and different types of tasks have different task processing times on different types of computing nodes. Therefore, in order to assign different types of tasks to the appropriate computing nodes for execution to achieve higher performance and lower power consumption, it is necessary to predict the task processing time of the computing nodes executing tasks.
[0005] How to accurately predict the task processing time of computing nodes is a problem that needs to be solved urgently. Summary of the Invention
[0006] The embodiments of the present application provide a method, apparatus, device, and storage medium for predicting task processing duration, in order to accurately predict the task processing duration of a computing node in a wireless network.
[0007] In the first aspect, an embodiment of the present application provides a method for predicting task processing time. The execution subject of the method may be a communication device, which may be the first device hereinafter, which may be a communication device in a radio access network (RAN) or a core network, or the communication device may be a component in a communication device, such as a chip, a chip system, or other functional module capable of calling and executing a program. For ease of understanding, the following exemplary description is given using the communication device as the execution subject.
[0008] Exemplarily, the method includes: a communication device receives a prediction request for task processing time, the prediction request carries first task information, and through a first model, based on the computing power information and the first task information of the computing node group, predicts the task processing time of the computing node group, the computing node group includes M computing nodes, M is a positive integer, and the first model is determined based on the computing node group to achieve accurate prediction of the task processing time when the computing node group performs different tasks.
[0009] In one possible implementation, the first task information may include at least one of resource parameters, task parameters, and input data information. The resource parameters, task parameters, and input data information are all associated with the task processing duration, enabling the first model to accurately predict the task processing duration. Resource parameters may indicate the type and quantity of resources used to perform task processing; task parameters indicate the parameters used for task processing; and input data information indicates the size of the input data required for task processing.
[0010] Optionally, when executing task processing through the third model, the task parameters may include model structure information of the third model.
[0011] Optionally, when task processing is performed through the first code, the task parameters may include information of the first code.
[0012] In one possible implementation, the computing nodes in the computing node group may be correlated with each other. Since the first model corresponds to the computing node group, the higher the correlation between the computing nodes in the computing node group, the higher the prediction accuracy of the task processing time. Therefore, when at least one computing power parameter of the M computing nodes is the same, the correlation between the computing nodes in the computing node group can be ensured, thereby improving the prediction accuracy of the task processing time.
[0013] In one possible implementation, the first model may be determined based on a second model in at least one first computing node, where the second model is used to predict a task processing duration of the corresponding first computing node, and the M computing nodes include at least one first computing node. Exemplarily, the communication device may determine a first model corresponding to a computing node group based on the second model in the at least one first computing node to obtain the first model, thereby facilitating subsequent use of the first model to predict task processing durations for computing nodes in the computing node group.
[0014] In one possible implementation, the communication device may first obtain the second model before determining the first model based on the second model. Exemplarily, the communication device may send a first request requesting information about the second model and receive a first response including information about the second model.
[0015] Optionally, the information of the second model includes model structure information and / or model parameter information of the second model.
[0016] Optionally, the information of the second model includes a model description of the second model, where the model description of the second model is used to indicate the type of input data of the second model and the type of output data of the second model.
[0017] In one possible implementation, the communication device may receive first information from each of N computing nodes during the process of acquiring the second model, where the first information includes information indicating whether the computing node includes the second model, and the N computing nodes include the above-mentioned M computing nodes, and the communication device may determine at least one first computing node among the M computing nodes based on the N first information, so as to obtain the second model in the at least one first computing node.
[0018] In one possible implementation, the communication device may receive computing power information from each of the N computing nodes, the computing power information including at least one computing power parameter of the corresponding computing node, and determine a computing node group and computing power information of the computing node group based on the computing power information of each of the N computing nodes. It should be understood that the computing power information of each of the N computing nodes may also be used to group the N computing nodes to obtain at least one computing node group.
[0019] In one possible implementation, the communication device may receive status information from each of the N computing nodes during the process of acquiring the second model, where the status information includes network transmission bandwidth and / or the remaining power of the corresponding computing node. The status information is used to determine at least one first computing node, thereby acquiring the second model in the at least one first computing node. This application does not limit the manner in which the communication device determines at least one first computing node. For example, the communication device may determine and indicate a first computing node in combination with the above-mentioned first information, computing power information, and status information, or randomly select a first computing node.
[0020] In one possible implementation, the communication device expects to obtain the second model based on a computing node among the M computing nodes that does not deploy the second model. For example, a computing node that does not deploy the second model is more likely to be used to perform task processing, such as if the computing node has a larger network transmission bandwidth and / or more remaining power. Therefore, considering that the first model corresponding to the computing node group is determined based on the second model of the computing node, the task processing duration can be more accurately predicted; for example, when none of the M computing nodes deploy the second model. In this case, the communication device can request at least one computing node among the M computing nodes to obtain the second model through model training.
[0021] Exemplarily, the communication device sends a second request carrying first instruction information for training a second model, and receives a second response to obtain the second model. Wherein, when the first instruction information indicates that the second model is to be trained, the second response includes information about the second model; or, when the first instruction information indicates that a first training dataset is to be obtained, the second response includes the first training dataset used to train the second model, where each first training data in the first training dataset includes a feature and a label, where the label is a measured value of the task processing duration corresponding to the feature.
[0022] In one possible implementation, the second request further carries second indication information, which is used to instruct the generation of labels in the first training dataset. Optionally, the second indication information may carry a type identifier of a benchmark task, which may be used to identify the benchmark task, which is used to generate a measurement value of the task processing duration based on the computing power information of the first computing node and the test task information.
[0023] In one possible implementation, the second request also includes model structure information of the second model, so that the first computing node can clearly understand the required model structure of the second model, and then train the second model based on the first training data set and the model structure of the second model.
[0024] In one possible implementation, the second request also includes a model description of the second model, so that the first computing node can clearly understand the type of input data and output data required for the second model, so that the second model obtained through training meets the requirements.
[0025] In a possible implementation, the communication device may send a first prediction response, where the first prediction response carries information about the task processing duration corresponding to the computing node group, so as to transmit the task processing duration of the computing node group.
[0026] Optionally, the first prediction response may include an identifier of the computing node group and information about the task processing duration, so as to clarify the correspondence between the task processing duration and the computing node group.
[0027] In one possible implementation, the communication device may send third indication information. When the third indication information indicates updating the second model, the communication device receives updated second model information from the first computing node and updates the first model based on the updated second model information. Updating the second model through the first computing node, and thereby updating the first model, can save resource overhead for the communication device.
[0028] In one possible implementation, the communication device may send a third indication message. When the third indication message indicates that the second model is not to be updated, the communication device receives the measured value of the task processing time corresponding to the first task information sent by the first computing node, and trains the first model based on the measured value of the task processing time corresponding to the first task information to obtain an updated first model. The communication device does not need to wait for the first computing node to update the second model before updating the first model, thereby reducing processing delay.
[0029] In a possible implementation, the first request carries the third indication information.
[0030] In one possible implementation, when the first task information corresponds to the first task, the communication device sends fourth indication information, the first task includes unexecuted tasks, and the fourth indication information carries the type identifier of the first task, so as to perform incremental learning on the unexecuted tasks, improve the prediction accuracy of the second model, and thereby improve the prediction accuracy of the first model.
[0031] In a second aspect, an embodiment of the present application provides a method for predicting task resource information. The execution subject of the method may be a communication device, which may be the first device described below. The communication device may be a communication device in a RAN or core network, or the communication device may be a component in the communication device, such as a chip, a chip system, or other functional module capable of calling and executing a program. For ease of understanding, the following exemplary description is given using the communication device as the execution subject.
[0032] Exemplarily, the method includes: a communication device receiving a prediction request for task resource information, the prediction request carrying second task information, and predicting, using a fourth model, task resource information for a computing node group based on computing power information and the second task information of the computing node group, the computing node group including M computing nodes, where M is a positive integer, and the fourth model is determined based on the computing node group. Accurate prediction of task resource information when the computing node group executes different tasks is achieved.
[0033] In one possible implementation, the second task information includes at least one of expected duration, task parameters, and input data information. The expected duration, task parameters, and input data information are associated with task resource information to facilitate accurate prediction of the task resource information by the fourth model. Task parameters indicate the parameters used for task processing; input data information indicates the size of the input data required for task processing.
[0034] Optionally, when executing task processing through the third model, the task parameters may include model structure information of the third model.
[0035] Optionally, when task processing is performed through the first code, the task parameters may include information of the first code.
[0036] In one possible implementation, the computing nodes in the computing node group may be correlated with each other. Since the fourth model corresponds to the computing node group, the higher the correlation between the computing nodes in the computing node group, the higher the prediction accuracy of the task resource information. Therefore, when at least one computing power parameter of the M computing nodes is the same, the correlation between the computing nodes in the computing node group can be ensured, thereby improving the prediction accuracy of the task resource information.
[0037] In one possible implementation, the fourth model can be trained based on the third training data set, the labels in the third training data set are determined by the first model of the computing node group, and the first model is used to predict the task processing time of the corresponding computing node group.
[0038] In a possible implementation, the communication device may send a second prediction response, where the second prediction response carries the task resource information corresponding to the computing node group, so as to transmit the task resource information.
[0039] Optionally, the second prediction response includes the identifier of the computing node group and the task resource information to clarify the correspondence between the task resource information and the computing node group.
[0040] In one possible implementation, the first model may be determined based on a second model in at least one first computing node, where the second model is used to predict the task processing duration of the corresponding first computing node, and the M computing nodes include at least one first computing node.
[0041] In one possible implementation, the communication device may first obtain the second model before determining the first model based on the second model. Exemplarily, the communication device may send a first request requesting information about the second model and receive a first response including information about the second model.
[0042] Optionally, the information of the second model includes model structure information and / or model parameter information of the second model.
[0043] Optionally, the information of the second model includes a model description of the second model, where the model description of the second model is used to indicate the type of input data of the second model and the type of output data of the second model.
[0044] In one possible implementation, the communication device may receive first information from each of N computing nodes during the process of acquiring the second model, where the first information includes information indicating whether the computing node includes the second model, and the N computing nodes include the above-mentioned M computing nodes, and the communication device may determine at least one first computing node among the M computing nodes based on the N first information.
[0045] In one possible implementation, the communication device can receive computing power information from each of N computing nodes, where the computing power information includes at least one computing power parameter of the corresponding computing node, and determine the computing node group and the computing power information of the computing node group based on the computing power information of each of the N computing nodes.
[0046] In one possible implementation, the communication device may receive status information from each of the N computing nodes during the process of acquiring the second model. The status information includes the network transmission bandwidth and / or the remaining power of the corresponding computing node. The status information is used to determine at least one first computing node, and then acquire the second model in the at least one first computing node.
[0047] In one possible implementation, the communication device expects to obtain the second model based on a computing node among the M computing nodes that does not deploy the second model. For example, a computing node that does not deploy the second model is more likely to be used to perform task processing, such as if the computing node has a larger network transmission bandwidth and / or more remaining power. Therefore, considering that the first model corresponding to the computing node group is determined based on the second model of the computing node, the task processing duration can be more accurately predicted; for example, when none of the M computing nodes deploy the second model. In this case, the communication device can request at least one computing node among the M computing nodes to obtain the second model through model training.
[0048] Exemplarily, the communication device sends a second request carrying first instruction information for training a second model, and receives a second response to obtain the second model. Wherein, when the first instruction information indicates that the second model is to be trained, the second response includes information about the second model; or, when the first instruction information indicates that a first training dataset is to be obtained, the second response includes the first training dataset used to train the second model, where each first training data in the first training dataset includes a feature and a label, where the label is a measured value of the task processing duration corresponding to the feature.
[0049] In one possible implementation, the second request further carries second indication information, which is used to instruct the generation of labels in the first training dataset. Optionally, the second indication information may carry a type identifier of a benchmark task, which may be used to identify the benchmark task, which is used to generate a measurement value of the task processing duration based on the computing power information of the first computing node and the test task information.
[0050] In a possible implementation, the second request further includes model structure information of the second model.
[0051] In a possible implementation manner, the second request further includes a model description of the second model.
[0052] In a third aspect, an embodiment of the present application provides a communication device, comprising: a transceiver module for receiving a prediction request for a task processing time, the prediction request carrying first task information; a processing module for predicting the task processing time of a computing node group based on the computing power information and the first task information of the computing node group through a first model, the computing node group including M computing nodes, M being a positive integer, and the first model being determined based on the computing node group.
[0053] In a possible implementation, the first task information includes resource parameters, where the resource parameters indicate the type and quantity of resources for executing task processing.
[0054] In a possible implementation, the first task information includes task parameters, where the task parameters are used to indicate parameters used for task processing, and the task parameters are associated with the task processing duration.
[0055] In a possible implementation, the task parameters include model structure information of a third model, and the third model is used for task processing.
[0056] In a possible implementation, the task parameter includes information of a first code, and the first code is used for task processing.
[0057] In a possible implementation, the first task information includes information about input data, where the information about the input data indicates the size of input data required for task processing.
[0058] In one possible implementation, at least one computing power parameter of the M computing nodes is the same.
[0059] In one possible implementation, the first model is determined based on a second model in at least one first computing node, the second model is used to predict the task processing duration of the corresponding first computing node, and the M computing nodes include at least one first computing node.
[0060] In a possible implementation, the transceiver module is further used to: send a first request, where the first request is used to request sending information of the second model; and receive a first response, where the first response includes the information of the second model.
[0061] In one possible implementation, the transceiver module is also used to: receive first information from each of N computing nodes, the first information including information indicating whether the computing node includes the second model, and the N computing nodes include M computing nodes; the processing module is also used to determine at least one first computing node among the M computing nodes based on the N first information.
[0062] In one possible implementation, the transceiver module is also used to receive computing power information from each of the N computing nodes, where the computing power information includes at least one computing power parameter of the corresponding computing node; the processing module is also used to determine the computing node group and the computing power information of the computing node group based on the computing power information of each of the N computing nodes.
[0063] In one possible implementation, the transceiver module is further used to receive status information from each of the N computing nodes, where the status information includes network transmission bandwidth and / or remaining power of the corresponding computing node, and the status information is used to determine at least one first computing node.
[0064] In a possible implementation manner, the information of the second model includes model structure information and / or model parameter information of the second model.
[0065] In a possible implementation, the information of the second model includes a model description of the second model, where the model description of the second model is used to indicate a type of input data of the second model and a type of output data of the second model.
[0066] In one possible implementation, the transceiver module is also used to: send a second request, the second request carries first indication information for training a second model; and receive a second response; wherein, when the first indication information indicates that the second model is obtained by training, the second response includes information about the second model, or, when the first indication information indicates that the first training data set is obtained, the second response includes the first training data set used to train the second model, and each first training data in the first training data set includes a feature and a label, and the label is a measured value of the task processing time corresponding to the feature.
[0067] In a possible implementation, the second request further carries second indication information, where the second indication information is used to instruct generation of labels in the first training data set.
[0068] In a possible implementation manner, the second request further includes model structure information of the second model and a model description of the second model.
[0069] In a possible implementation, the transceiver module is further configured to send a first prediction response, where the first prediction response carries information about the task processing duration corresponding to the computing node group.
[0070] In a possible implementation, the first prediction response includes an identifier of the computing node group and information about the task processing duration.
[0071] In one possible implementation, the transceiver module is also used to send a third indication message; when the third indication message indicates that the second model is updated, the information of the updated second model sent by the first computing node is received; and the first model is updated according to the information of the updated second model; or, when the third indication message indicates that the second model is not updated, the measured value of the task processing time corresponding to the first task information sent by the first computing node is received; and the first model is trained according to the measured value of the task processing time corresponding to the first task information to obtain the updated first model; wherein the first computing node belongs to M computing nodes.
[0072] In a possible implementation manner, the first request carries third indication information.
[0073] In a possible implementation, when the first task information corresponds to a first task, the transceiver module is further configured to send fourth indication information, the first task includes an unexecuted task, and the fourth indication information carries a type identifier of the first task.
[0074] In a fourth aspect, an embodiment of the present application provides a communication device, comprising: a transceiver module for receiving a prediction request for task resource information, the prediction request carrying second task information; a processing module for predicting the task resource information of a computing node group based on the computing power information and the second task information of the computing node group through a fourth model, the computing node group including M computing nodes, M being a positive integer, and the fourth model being determined based on the computing node group.
[0075] In a possible implementation, the second task information includes an expected duration, and the expected duration is associated with the task resource information.
[0076] In a possible implementation, the second task information includes task parameters, where the task parameters are used to indicate parameters used for task processing, and the task parameters are associated with task resource information.
[0077] In a possible implementation, the task parameters include model structure information of a third model, and the third model is used for task processing.
[0078] In a possible implementation, the task parameter includes information of a first code, and the first code is used for task processing.
[0079] In a possible implementation, the second task information includes information about input data, where the information about the input data indicates the size of input data required for task processing.
[0080] In one possible implementation, at least one computing power parameter of the M computing nodes is the same.
[0081] In one possible implementation, the fourth model is trained based on the third training data set, the labels in the third training data set are determined by the first model of the computing node group, and the first model is used to predict the task processing time of the corresponding computing node group.
[0082] In a possible implementation, the transceiver module is further configured to send a second prediction response, where the second prediction response carries task resource information corresponding to the computing node group.
[0083] In a possible implementation, the second prediction response includes an identifier of the computing node group and task resource information.
[0084] In one possible implementation, the first model is determined based on a second model in at least one first computing node, the second model is used to predict the task processing duration of the corresponding first computing node, and the M computing nodes include at least one first computing node.
[0085] In a possible implementation, the transceiver module is further used to: send a first request, where the first request is used to request sending information of the second model; and receive a first response, where the first response includes the information of the second model.
[0086] In one possible implementation, the transceiver module is also used to: receive first information from each of N computing nodes, the first information including information indicating whether the computing node includes the second model, and the N computing nodes include M computing nodes; the processing module is also used to determine at least one first computing node among the M computing nodes based on the N first information.
[0087] In one possible implementation, the transceiver module is also used to receive computing power information from each of the N computing nodes, where the computing power information includes at least one computing power parameter of the corresponding computing node; the processing module is also used to determine the computing node group and the computing power information of the computing node group based on the computing power information of each of the N computing nodes.
[0088] In one possible implementation, the transceiver module is further used to receive status information from each of the N computing nodes, where the status information includes network transmission bandwidth and / or remaining power of the corresponding computing node, and the status information is used to determine at least one first computing node.
[0089] In a possible implementation manner, the information of the second model includes model structure information and / or model parameter information of the second model.
[0090] In a possible implementation, the information of the second model includes a model description of the second model, where the model description of the second model is used to indicate a type of input data of the second model and a type of output data of the second model.
[0091] In one possible implementation, the transceiver module is also used to: send a second request, the second request carries first indication information for training a second model; and receive a second response; wherein, when the first indication information indicates that the second model is obtained by training, the second response includes information about the second model, or, when the first indication information indicates that the first training data set is obtained, the second response includes the first training data set used to train the second model, and each first training data in the first training data set includes a feature and a label, and the label is a measured value of the task processing time corresponding to the feature.
[0092] In a possible implementation, the second request further carries second indication information, where the second indication information is used to instruct generation of labels in the first training data set.
[0093] In a possible implementation manner, the second request further includes model structure information of the second model and a model description of the second model.
[0094] In a fifth aspect, an embodiment of the present application provides a communication device, comprising: a processor, configured to execute the method in the first aspect, the second aspect, or each possible implementation manner by running a computer program or through a logic circuit.
[0095] In a possible implementation manner, the communication device further includes: a memory, wherein the memory is used to store the computer program.
[0096] In a possible implementation, the communication device further includes: a communication interface, where the communication interface is used to input and / or output signals.
[0097] In a sixth aspect, an embodiment of the present application provides a chip, comprising: a processor for calling and executing computer instructions from a memory, so that a device equipped with the chip executes a method as in the first aspect, the second aspect, or each possible implementation.
[0098] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium for storing computer program instructions, wherein the computer program enables a computer to execute the method in the first aspect, the second aspect, or each possible implementation manner.
[0099] In an eighth aspect, an embodiment of the present application provides a computer program that enables a computer to execute the method in the first aspect, the second aspect, or each possible implementation manner described above.
[0100] In a ninth aspect, an embodiment of the present application provides a computer program product, comprising computer program instructions, which enable a computer to execute the method in the first aspect, the second aspect, or each possible implementation manner.
[0101] The beneficial effects of the contents of the above-mentioned second to ninth aspects and each possible implementation method can be found in the beneficial effects brought about by the above-mentioned first aspect and each possible implementation method of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] FIG1 shows a possible, non-limiting system schematic.
[0103] FIG2 is a schematic diagram of the architecture of a communication system provided in an embodiment of the present application.
[0104] FIG3 is a schematic diagram of an interactive flow of a method for predicting task processing time provided in an embodiment of the present application.
[0105] FIG4 is a schematic diagram of an interactive flow of a method for predicting task processing time provided in an embodiment of the present application.
[0106] FIG5 is a schematic diagram of an interactive flow of a method for predicting task processing time provided in an embodiment of the present application.
[0107] FIG6 is a schematic diagram of an interactive flow of a method for predicting task resource information provided in an embodiment of the present application.
[0108] FIG7 is a schematic diagram of an interactive flow of a method for predicting task resource information provided in an embodiment of the present application.
[0109] FIG8 is a schematic block diagram of a communication device provided in an embodiment of the present application.
[0110] FIG9 is another schematic block diagram of a communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0111] The technical solution in this application will be described below with reference to the accompanying drawings.
[0112] Figure 1 shows a possible, non-limiting system diagram. As shown in Figure 1, the communication system 10 includes a RAN 100 and a core network (CN) 200. The RAN 100 includes at least one RAN node (such as 110a and 110b in Figure 1, collectively referred to as 110) and at least one terminal (such as 120a-120j in Figure 1, collectively referred to as 120). The RAN 100 may also include other RAN nodes, such as wireless relay equipment and / or wireless backhaul equipment (not shown in Figure 1). The terminal 120 is connected to the RAN node 110 wirelessly. The RAN node 110 is connected to the core network 200 wirelessly or by wire. The core network equipment in the core network 200 and the RAN node 110 in the RAN 100 can be different physical devices, or they can be the same physical device that integrates the core network logical functions and the radio access network logical functions.
[0113] The RAN 100 may be a cellular system related to the Third Generation Partnership Project (3GPP), such as a 4G, 5G, or 6G mobile communication system. The RAN 100 may also be an open access network (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system. The RAN 100 may also be a communication system that integrates two or more of the above systems.
[0114] RAN node 110, sometimes also referred to as access network equipment, RAN entity, or access node, constitutes part of the communication system and facilitates wireless access for terminals. Multiple RAN nodes 110 in the communication system 10 can be of the same type or different types. In some scenarios, the roles of RAN node 110 and terminal 120 are relative. For example, network element 120i in Figure 1 can be a helicopter or drone, which can be configured as a mobile base station. For terminal 120j accessing RAN 100 via network element 120i, network element 120i is a base station; however, for base station 110a, network element 120i is a terminal. RAN node 110 and terminal 120 are sometimes referred to as communication devices. For example, network elements 110a and 110b in Figure 1 can be understood as communication devices with base station functionality, and network elements 120a-120j can be understood as communication devices with terminal functionality.
[0115] In one possible scenario, a RAN node may be a base station, an evolved NodeB (eNodeB), an access point (AP), a transmission reception point (TRP), a next generation NodeB (gNB), a next generation base station in a sixth generation (6G) mobile communication system, a base station in a future mobile communication system, or an access node in a WiFi system. A RAN node may be a macro base station (such as 110a in FIG1 ), a micro base station or an indoor station (such as 110b in FIG1 ), a relay node or a donor node, or a wireless controller in a CRAN scenario. Optionally, a RAN node may also be a server, a wearable device, a vehicle or an onboard device. For example, an access network device in vehicle to everything (V2X) technology may be a road side unit (RSU). All or part of the functions of the RAN node in this application may also be implemented by software functions running on hardware, or by virtualized functions instantiated on a platform (such as a cloud platform). The RAN node in this application may also be a logical node, a logical module or software that can implement all or part of the RAN node functions.
[0116] In another possible scenario, multiple RAN nodes collaborate to assist the terminal in achieving wireless access, and different RAN nodes respectively implement part of the functions of the base station. For example, the RAN node can be a centralized unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU). The CU and DU can be set separately, or they can be included in the same network element, such as a baseband unit (BBU). The RU can be included in a radio frequency device or radio frequency unit, such as a remote radio unit (RRU), an active antenna unit (AAU), or a remote radio head (RRH). In the ORAN system, CU may also be called open CU (open-CU, O-CU), DU may also be called open DU (open-DU, O-DU), CU-CP may also be called open CU-CP (open-CU-CP, O-CU-CP), CU-UP may also be called open CU-UP (open-CU-UP, O-CU-UP), and RU may also be called open RU (open-RU, O-RU).
[0117] A terminal may also be called terminal equipment, user equipment (UE), access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device.
[0118] The terminal device may be a device that provides voice / data connectivity to users, such as a handheld device or vehicle-mounted device with wireless connection function. At present, some examples of terminals include: mobile phones, tablet computers, computers with wireless transceiver functions (such as laptops, PDAs, etc.), drones, customer-premises equipment (CPE), smart point of sale (POS) machines, mobile internet devices (MID), virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), and so on. assistant, PDA), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, vehicle-mounted devices, wearable devices, terminal devices in a 5G network or terminal devices in a system evolved after 5G, etc.
[0119] Figure 2 is a schematic diagram of the architecture of a communication system provided by an embodiment of the present application. To manage, orchestrate, and schedule intelligent services, the 6G network deploys orchestration, scheduling, status aggregation, and computing power evaluation functions. These functions can be implemented through different functional modules.
[0120] Some or all of the functional modules in the orchestration function, scheduling function, status aggregation function and computing power evaluation function can be centrally deployed or distributed in the core network, or some or all of the functional modules can be centrally deployed or distributed in the RAN, or some or all of the functional modules can be distributed in the core network and RAN, or some or all of the functional modules can be deployed in other locations. This application does not limit this.
[0121] Optionally, any one or more of the above functional modules may be deployed in an access and mobility management function (AMF), a session management function (SMF), or a network repository function (NRF). Alternatively, when the RAN includes a CU and a DU, any of the above functions may be implemented by the CU and / or the DU. In an ORAN system, any of the above functions may be implemented by the O-CU and / or the O-DU.
[0122] Part or all of the orchestration, scheduling, status aggregation, and computing power evaluation functions can be deployed on the control or management plane of the 6G network, or other planes added in the future.
[0123] Any one or more of the above functional modules can manage, orchestrate, and schedule computing resource pools on the cloud, edge, and end sides. A computing resource pool may include one or more computing nodes. For example, a computing resource pool on the core network or cloud side includes computing nodes 1 to m, which include but are not limited to computing resource pools within the core network, or computing resource pools in data centers such as public clouds, private clouds, or hybrid clouds outside the core network; a RAN or edge computing resource pool includes computing nodes 1 to n, which include but are not limited to computing resource pools or computing boards within RAN nodes, or computing resource pools in multi-access edge computing (MEC) or other edge data centers close to RAN nodes; and a terminal computing resource pool includes computing nodes 1 to p. m, n, and p are all positive integers. A computing node may be a device with computing capabilities, such as a physical machine, virtual machine, or other type of device. A computing node in a terminal computing resource pool may be a terminal device.
[0124] Intelligent services, such as those in 6G networks, can be described through workflows. Workflows can consist of one or more tasks, where a task refers to the process of achieving a specific goal through the collaboration of multi-dimensional resources at the network level. Depending on the purpose, tasks can be divided into different types of tasks, such as AI reasoning, AI training, computing, and perception. The orchestration function can orchestrate workflows, such as assessing the resource requirements of tasks. The scheduling function can determine scheduling information for tasks in the workflow, including allocating computing nodes to the execution units required for the tasks. The state aggregation function can record information related to computing nodes, tasks, etc., which can be used by other devices, network functions, or modules to save, query, update, or delete data. The computing power assessment function can predict the time it takes for a computing node to process a task in the workflow, or predict the resource requirements of the computing node to process the task.
[0125] The modules in the compute nodes can be used to implement the aforementioned intelligent services. The modules in the compute nodes that process tasks when implementing intelligent services are called execution units. Tasks can be of different types, and execution units can be reused to execute tasks of the same type.
[0126] When the execution unit performs task processing, it can be implemented using AI models, program codes, etc., and the instance of the execution unit is executed in the host environment of the computing node in a process or thread manner; the AI model, program code, etc. used when performing task processing can also be constructed as a container image of the execution unit, and the instance of the execution unit runs in a container manner; the AI model, program code, etc. used when performing task processing can also be constructed as a virtual machine image of the execution unit, and the instance of the execution unit runs in a virtual machine manner; the AI model, program code, etc. used when performing task processing can also be constructed into other formats of the execution unit, and the instance of the execution unit runs in other ways. This application does not limit this.
[0127] To accelerate task processing, the execution units within compute nodes exhibit the aforementioned heterogeneity. Furthermore, compute node processors, such as CPUs, have varying architectures and clock speeds. Furthermore, various accelerators used in intelligent computing, such as graphics processing units (GPUs), neural processing units (NPUs), and tensor processing units (TPUs), exacerbate this heterogeneity. Consequently, managing, orchestrating, and scheduling heterogeneous compute nodes is gaining increasing attention. Accurately predicting the task processing duration of compute nodes facilitates this management, orchestration, and scheduling.
[0128] Therefore, in response to the above problems, an embodiment of the present application introduces a first model corresponding to the computing node group, and predicts the task processing time of the computing node group based on the first model and the computing power information and task information of the computing node group, so as to accurately predict the task processing time of the computing node group when performing different tasks.
[0129] This application can be applied to the prediction of task processing duration for computing nodes in any computing resource pool, and is particularly applicable to the prediction of task processing duration for heterogeneous computing nodes.
[0130] It should be understood that the above is only explained with the 6G network as an example, and this application is not limited to this. For example: the technical solution provided in this application can also be applied to: long term evolution (LTE) system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD), sidelink (SL) communication system, universal mobile telecommunication system (UMTS), world-wide interoperability for microwave access (WiMAX) communication system, fifth generation (5G) mobile communication system or new radio access technology (NR), etc. Among them, the 5G mobile communication system can include non-standalone (NSA) and / or standalone (SA). The technical solution provided in this application can also be applied to future communication systems.
[0131] The following describes the method for predicting task processing time provided by the embodiment of the present application with reference to the accompanying drawings.
[0132] It should be understood that the following description is for ease of understanding and explanation only, and the method provided in the embodiments of the present application is described using the interaction between a first device and a second device as an example. The first device may be deployed with a functional module that implements a computing power evaluation function, and the second device may be deployed with functional modules that implement other functions, such as a functional module for a scheduling function and / or a functional module for an orchestration function. For example, both the first device and the second device may be the RAN node 110 in Figure 1, or devices in the core network 200.
[0133] It should also be understood that this should not constitute any limitation on the execution subject of the method provided in this application. As long as it is possible to execute the method provided in the embodiment of this application by running a program having the code of the method provided in the embodiment of this application, it can serve as the execution subject of the method provided in the embodiment of this application. For example, the above-mentioned first device may also be a component in a device (such as a device in the RAN node 110 or the core network 200), such as a chip, a chip system or other functional module that can call a program and execute a program; the above-mentioned second device may also be a component in a device (such as a device in the RAN node 110 or the core network 200), such as a chip, a chip system or other functional module that can call a program and execute a program.
[0134] FIG3 is a schematic diagram of an interactive flow of a method for predicting task processing duration provided by an embodiment of the present application. Referring to FIG3 , the method 300 includes at least some of the following steps.
[0135] S310: The second device sends a task processing duration prediction request to the first device, the prediction request carrying first task information. Correspondingly, the first device receives the task processing duration prediction request from the second device.
[0136] S320, the first device predicts the task processing time of the computing node group according to the computing power information and the first task information of the computing node group through the first model, and the first model is determined based on the computing node group.
[0137] S330: The first device sends a first prediction response to the second device, the first prediction response carrying information about the task processing duration corresponding to the computing node group. Correspondingly, the second device receives the first prediction response from the first device.
[0138] The first task information carries information about the task to be executed, so that the first device can determine the task processing time of the computing node group for the specific task and improve the accuracy of the prediction. It should be understood that this application does not limit the task to be executed carried by the first task information to be the first task below. Only when the task to be executed carried by the first task information is not executed or is not successfully executed, the task to be executed carried by the first task information is the first task below.
[0139] It should be noted that the "first" in the first task information is only used to distinguish it from the second task information used to predict task resource information below. Similarly, the use of prefixes such as "first" and "second" in this application is to facilitate the distinction and description of different things belonging to the same name category, and does not restrict the order, size or quantity of things. For example, the "first model", "second model", "third model" and "fourth model" below are just different models, and there is no priority relationship between the four.
[0140] Exemplarily, the first task information may include at least one of the following information about the task to be executed:
[0141] 1. Resource parameters, used to indicate the type and quantity of resources used to perform task processing. For example, resource parameters can be a two-tuple list consisting of the type and quantity of resources. The resources can be the resources required for the execution of the execution unit in the aforementioned example. The types of resources can include CPU, GPU, NPU, system memory, etc. The type of resources can also include the architecture and / or model of the processor, such as the architecture or model of the CPU, the architecture or model of the GPU, the architecture or model of the NPU, etc. For example, different positive integers can be used to represent different resource types.
[0142] 2. Task parameters: These are used to indicate the parameters used in task processing. Different types of execution units correspond to different types of task parameters.
[0143] For example, when the execution unit includes a third model for task processing, the task parameters may include model structure information of the third model. For example, the task parameters may indicate the structure of the third model, including a feature matrix and / or an adjacency matrix, etc., wherein the feature matrix includes the characteristics of different operators of the neural network in the AI model, such as operator type, operator size (i.e., the dimension of the input and output tensors of the operator) and operator computational complexity, and the adjacency matrix includes the connection relationship between operators. For another example, the task parameters may include the operator type, operator size and operator computational complexity of the neural network in the AI model. For another example, when the execution unit includes a first code for task processing, the task parameters may include information of the first code, and the information of the first code may include an abstract syntax tree of the code, which may be converted from the first code using a compilation tool. The first code may be a section of code for performing task processing, and the first code may be designed based on any programming language or programming environment.
[0144] 3. Input data information, which indicates the size of input data required for task processing. For example, when the execution unit executes task processing based on the third model, the input data information may indicate the size of data input to the third model.
[0145] It should be understood that the information included in the first task information, such as resource parameters, task parameters, and input data information, are all associated with the task processing duration.
[0146] It should also be understood that when the first task information does not carry one or more of the information in the above examples, the uncarried information may be preset or pre-agreed. For example, when the first task information does not carry input data information, the data input to the third model or the first code may be data of a unit size. If the third model or the first code requires a fixed input data size, the unit size is the fixed input data size. If the third model or the first code supports variable-length input data sizes, such as supporting an array or list of input data, or supporting multiple batches of input data, the unit size is the size of one input data item in the array or list of input data, or the size of one batch of input data among multiple batches of input data.
[0147] Exemplarily, the prediction request in S310 may also carry a task type identifier, which is used to uniquely identify a type of task. Since the task type has a corresponding relationship with the execution unit, the type identifier can also be used to uniquely identify the type of execution unit that executes the task.
[0148] Exemplarily, a prediction request for task processing time can be sent by the party demanding the task processing time prediction. For example, the scheduling function can send the prediction request to obtain the predicted task processing time, and then schedule the task according to the predicted task processing time; for another example, the orchestration function can send the prediction request to obtain the predicted task processing time, and then schedule the task according to the predicted task processing time; for another example, when the task processing time acts on orchestration and scheduling, the orchestration function and the scheduling function can collaboratively send the prediction request, such as the orchestration function can send job description information to the scheduling function, and then the scheduling function sends the prediction request to the computing power evaluation function. This application does not limit this.
[0149] The computing node group may include M computing nodes, where M is a positive integer. The computing nodes in the computing node group may be correlated or uncorrelated. Generally speaking, since the first model corresponds to the computing node group, the higher the correlation between the computing nodes in the computing node group, the higher the prediction accuracy of the task processing time. The present application does not limit the measurement standard for the correlation between computing nodes. For example, when the value of at least one computing power parameter among M computing nodes is the same or the difference is less than a threshold, it is determined that the computing nodes in the computing node group are correlated. Optionally, the first device may group the computing nodes based on the correlation between the computing nodes to obtain at least one computing node group to ensure that the computing nodes in the computing node group are correlated. For the specific grouping method, please refer to the example below.
[0150] The computing node group may be preset or pre-grouped. For example, the first device pre-groups multiple computing nodes to obtain at least one computing node group. The computing node group mentioned above may be any preset computing node group or any one of the at least one computing node group obtained by grouping. The specific implementation of the first device obtaining at least one computing node group by pre-grouping can be seen in the examples below.
[0151] The first model can be a pre-trained AI model for predicting task processing time. The input to the first model can include computing power information of the computing node group and the first task information, and the output of the first model can include information about the task processing time. It should be noted that the computing power information of the computing node group can include computing power information for each computing node in the computing node group. To save overhead, if the computing power information of the computing nodes in the computing node group is different, the computing power information of the computing node group can be the computing power information of any one computing node in the computing node group. Alternatively, the computing power information of the computing node group can be determined based on the computing power information of M computing nodes, such as the computing power information of the computing node group being the average of the computing power information of the M computing nodes. The computing power information of the computing node includes at least one computing power parameter of the computing node. The at least one computing power parameter can include at least one of: processor architecture name, processor name, number of processors, extended instruction set, system memory size, accelerator model, accelerator name, number of accelerators, accelerator memory size, hard disk storage size, or operating system type.
[0152] Exemplarily, the computing power information of a computing node may be expressed by traversing all combinations of at least one computing power parameter and encoding each combination. For example, the computing power information of a computing node may be expressed using a two-dimensional matrix consisting of multiple one-dimensional arrays, where each one-dimensional array in the two-dimensional matrix represents, from left to right, a primary index, a secondary index, ..., a K-level index of the computing power information of the computing node. The first-level index is the first number of the one-dimensional array, and different positive integers are used to represent resource types, such as processor, system memory, accelerator or accelerator memory. The second-level index is the second number of the one-dimensional array, and different positive integers are used to represent the architecture name or model of the resource in the first-level index. If there is no such resource type or the resource type has no architecture name or model, it can be represented by 0. The third-level index is the third number of the one-dimensional array, and different positive integers are used to represent the name of the resource in the first-level index. If there is no such resource type or the resource type has no name, it can be represented by 0. The fourth-level index is the fourth number of the one-dimensional array, and different positive integers are used to represent the number or size of resources in the first-level index. If there is no such resource type, it can be represented by 0. Other levels of indexes can also be used to express other computing power parameters, which will not be repeated here.
[0153] It should be understood that the embodiments of the present application do not limit the expression of the computing power information of the computing nodes. It should also be understood that the computing power information of the computing node group and the computing power information of the computing nodes can be expressed in the same or similar manner.
[0154] The task processing duration information output by the first model, that is, the information indicating the task processing duration of the computing node group, can indicate the task processing duration when any computing node among the M computing nodes processes the task corresponding to the first task information.
[0155] It should be noted that the embodiments of the present application do not limit the number of computing node groups. The prediction request for task processing time may include multiple first task information. Different first task information may correspond to different computing node groups, that is, different types of tasks may be executed by computing nodes in different computing node groups. The first device may predict the task processing time of the computing node group based on the first model corresponding to each computing node group, the computing power information of the computing node group and the first task information corresponding to the computing node group.
[0156] The first device sends the predicted task processing duration information to the second device, and the second device can schedule and / or arrange the task based on the task processing duration. This application does not limit the first device to sending the task processing duration information to the second device. For example, the first device can also send the task processing duration information to other devices, such as sending the task processing duration information to a display device to present the task processing duration.
[0157] Exemplarily, the information on the task processing duration can be carried by the first prediction response, and the first device sends the first prediction response to the second device to transmit the information on the task processing duration of the computing node group. Optionally, the first prediction response includes the identifier of the computing node group and the information on the task processing duration. For example, the first prediction response may include a two-tuple consisting of the identifier of the computing node group and the information on the task processing duration. When the first prediction response includes information on multiple task processing durations of multiple computing node groups, the first prediction response may include a two-tuple list consisting of the identifiers of the multiple computing node groups and the information on the task processing duration corresponding to each computing node group.
[0158] In an embodiment of the present application, the first device predicts the task processing time of the computing node group based on the first model, the computing power information of the computing node group and the first task information, thereby achieving accurate prediction of the task processing time of the computing node group when executing the task.
[0159] Figure 4 is a schematic diagram of the interactive flow of a method for predicting task processing time provided by an embodiment of the present application. In Figure 4, the first device is used as a computing power evaluation function, and the second device is used as a scheduling function as an example, and the embodiment of the present application is described in combination with other devices in the network (such as computing nodes). Among them, the workflow sender can be a user or other functional module, such as an orchestration function. It should be understood that the functions and devices shown in Figure 4 are only examples and do not constitute any limitation to the present application. The present application may include more or fewer functions and devices than Figure 4.
[0160] The following describes the scheduling process in conjunction with FIG4, using the predicted task processing duration as an example for implementing scheduling. Referring to FIG4, S408 to S410 are implemented in the same or similar manner as S310 to S330 in the embodiment shown in FIG3, and the technical solutions related to S310 to S330 are also applicable in this embodiment.
[0161] Exemplarily, before the scheduling function sends a prediction request for the task processing time to the computing power evaluation function, the scheduling function may further include S407. In S407, the workflow sender may send workflow description information to the scheduling function, and correspondingly, the scheduling function receives the workflow description information sent by the workflow sender.
[0162] The workflow description information may include a task template, which includes the first task information and the task type identifier in the aforementioned example. It should be understood that when a workflow includes multiple tasks, the workflow description information may include multiple such task templates, one or more such task templates may form a task template list, and the workflow description information may also include dependencies between the multiple tasks.
[0163] The scheduling function can parse the workflow description information to obtain task templates corresponding to different tasks. Optionally, the scheduling function determines the identifier of at least one computing node group for each task template to form a pre-selected group identifier list. Exemplarily, the scheduling function can filter computing nodes based on resource parameters in the task template. For example, the scheduling function determines, from multiple computing nodes, at least one computing node whose remaining resources meet the resource parameters in the task template. The identifier of the computing node group to which the at least one computing node obtained by filtering belongs can form the pre-selected group identifier list.
[0164] Optionally, in order to reduce processing complexity, some pre-selected computing node groups may be randomly selected from the pre-selected computing node groups indicated by the filtered pre-selected group identifier list to form a new pre-selected group identifier list.
[0165] Optionally, the scheduling function may perform a deduplication operation on the pre-selected group identifier list to save processing overhead.
[0166] Furthermore, the scheduling function can generate a prediction request for the task processing time based on the results of parsing the workflow description information. The prediction request for the task processing time can carry a list of pre-selected group identifiers. The computing power evaluation function predicts the task processing time of the corresponding computing node group based on the pre-selected group identifier list in the prediction request for the task processing time. When the prediction request for the task processing time does not carry the pre-selected group identifier list, the computing power evaluation function can use all preset or pre-grouped computing node groups as pre-selected computing node groups, or the computing power evaluation function can filter the pre-selected computing node groups from all computing node groups. The filtering logic is similar to that of the scheduling function in the aforementioned example and will not be repeated for the sake of brevity.
[0167] It should be understood that the embodiments of the present application do not limit the screening logic or screening algorithm for the pre-selected computing node group (or pre-selected group identifier).
[0168] Optionally, the scheduling function may save the correspondence between the task type identifier and the task template through the state aggregation function.
[0169] For example, when predicting the task processing time of a computing node group, the computing power evaluation function can first determine the first model corresponding to the computing node group and the computing power information corresponding to the computing node group, and then use the computing power information and the first task information of the computing node group as inputs of the first model to predict the task processing time.
[0170] Optionally, the state aggregation function can save a first correspondence, in which each computing node group in at least one computing node group has a corresponding first model, and the computing power evaluation function can read the first correspondence from the state aggregation function, and then determine the first model corresponding to the computing node group to be predicted (such as the above-mentioned pre-selected computing node group) based on the first correspondence. Optionally, the state aggregation function can save a second correspondence, in which each computing node group in at least one computing node group has corresponding computing power information, and the computing power evaluation function can read the second correspondence from the state aggregation function, and then determine the computing power information corresponding to the computing node group to be predicted (such as the above-mentioned pre-selected computing node group) based on the second correspondence.
[0171] Optionally, in the first corresponding relationship, each computing node group in the at least one computing node group further corresponds to a model description of the first model, where the model description is used to indicate a type of input data and a type of output data of the first model.
[0172] Optionally, the type of input data includes but is not limited to at least one of the following:
[0173] Compute the computing power information of the node group;
[0174] Resource parameters;
[0175] Task parameters;
[0176] Information about input data.
[0177] Optionally, the type of output data includes but is not limited to task processing time.
[0178] Exemplarily, in S409, the computing power evaluation function can traverse each pre-selected group identifier in the pre-selected group identifier list, and obtain the first model and computing power information of the computing node group corresponding to the pre-selected group identifier based on the first correspondence and the second correspondence, and then use the computing power information and the first task information of the computing node group as the input of the first model to determine the task processing time of each computing node group, and then obtain a two-tuple list consisting of the computing node group identifier (or group identifier) and the task processing time information.
[0179] Exemplarily, in S411, the scheduling function may combine a list of two-tuples consisting of group identifiers and task processing duration information, and the computing nodes in the computing node group corresponding to each group identifier in the two-tuple list (e.g., the M computing nodes described above), and use a preset scheduling algorithm to determine scheduling information. The scheduling information is used to instruct the execution unit that performs task processing to be assigned to the corresponding computing node. Exemplarily, the scheduling information may include an execution unit identifier, a task type identifier, and a computing node identifier.
[0180] In some embodiments, the first model corresponding to the computing node group can be pre-acquired by the computing power evaluation function. For example, the first model can be determined based on the second model in at least one first computing node, and the second model is used to predict the task processing time of the corresponding first computing node. The M computing nodes include the at least one first computing node.
[0181] The following is an exemplary description of the process of obtaining the first model with reference to S401 to S406 in Figure 4. It should be noted that when the computing power evaluation function obtains the first model, at least some of the steps S401 to S406 may be executed.
[0182] Each of the N computing nodes can report information to the computing power evaluation function. The computing power evaluation function groups at least some of the N computing nodes based on the information reported by the N computing nodes to obtain at least one computing node group. Each computing node group may include one or more computing nodes. For example, any one of the at least one computing node group may include M computing nodes. The computing power evaluation function can also determine at least one first computing node in each computing node group based on the information reported by the N computing nodes. In this embodiment, the first computing node is a computing node that provides a second model for determining the first model. The N computing nodes can be all or part of the computing nodes in the resource pool. The N computing nodes should include the above-mentioned M computing nodes. By grouping the N computing nodes, the M computing nodes in the N computing nodes are divided into one computing node group.
[0183] The reported information may include, but is not limited to, at least one of the following: first information, computing power information, and status information. It should be understood that the first information, computing power information, and status information may be sent in the same message, or they may be sent separately; this application does not limit this. The following provides an exemplary description of S401-1 to S401-3.
[0184] In S401-1, each of the N computing nodes may send first information to the computing power evaluation function, where the first information includes information indicating whether the computing node includes the second model. Optionally, when the computing node includes the second model, the first information also includes information indicating whether the computing node supports sending the second model. If the computing node does not support sending the second model, it may be deemed that the computing node does not include the second model. It should be understood that in S402, when the computing power evaluation function determines at least one first computing node among the M first computing nodes, all computing nodes including the second model indicated by the first information may be determined as first computing nodes, or they may be screened from the computing nodes including the second model indicated by the first information.
[0185] In S401-2, each of the N computing nodes may send computing power information to the computing power evaluation function. The computing power information is similar to the computing power information of the computing node in the above example and will not be described again for the sake of brevity.
[0186] As previously mentioned, the higher the correlation between the computing nodes in a computing node group, the higher the prediction accuracy of the task processing duration. Therefore, in some embodiments, at least one computing power parameter of the M computing nodes is the same to ensure the correlation between the computing nodes in the computing node group.
[0187] In view of this, in S402, the computing power evaluation function can group the N computing nodes based on the information reported by each computing node to obtain at least one computing node group, and one computing node group in the at least one computing node group includes the above-mentioned M computing nodes. Exemplarily, the computing power evaluation function can group computing nodes with all the same computing power parameters into one group, or group computing nodes with some of the same computing power parameters into one group. For example, when the computing power information includes the architecture name of the processor, the name of the processor, the number of processors, the extended instruction set, the system memory size, the model of the accelerator, the name of the accelerator, the number of accelerators, the memory size of the accelerator, the hard disk storage size and the operating system type, the computing nodes with the same computing power parameters in the above computing power information are grouped into one group, or the computing nodes with the same architecture name of the processor, the name of the processor, the number of processors, the system memory size, the model of the accelerator, the number of accelerators and the operating system type are grouped into one group, or the computing nodes with the same architecture name of the processor, the number of processors, the system memory size, the model of the accelerator and the number of accelerators are grouped into one group.
[0188] It should be understood that this application does not limit the grouping method of computing nodes, and the above grouping is only for illustrative purposes. For example, the computing power evaluation function can group computing nodes based on the remaining power of the computing nodes or the historical data of the computing nodes (such as the processing time of historical tasks).
[0189] Optionally, the state aggregation function can save a third correspondence relationship, in which each computing node group in the at least one computing node group obtained by grouping includes at least one computing node, such as the M computing nodes corresponding to one computing node group in the third correspondence relationship. For example, the third correspondence relationship can be a correspondence between the identifier of a computing node group and the node identifier of at least one computing node included in the computing node group.
[0190] Exemplarily, as shown in S405 in FIG4 , the computing power evaluation function may determine the computing power information of each computing node group obtained by grouping based on the computing power information of each computing node in the N computing nodes. As an example, the computing power evaluation function may use the computing power information of any computing node in the computing node group as the computing power information of the computing node group; as another example, the computing power evaluation function may determine the computing power information of the computing node group based on the computing power information of at least two computing nodes in the computing node group, for example, by performing a quantitative average of the computing power information of the at least two computing nodes to obtain the computing power information of the computing node group. It should be understood that the computing power evaluation function may also determine the computing power information of the computing node group based on other methods, and this application does not limit this.
[0191] It should be understood that the present application does not limit the execution order of S405 and S402 to S404.
[0192] Optionally, the computing power evaluation function can determine a second corresponding relationship based on the computing power information corresponding to each computing node group obtained by grouping, and save the second corresponding relationship to the state aggregation function.
[0193] In S401-3, each of the N computing nodes may send status information to the computing power evaluation function, where the status information includes network transmission bandwidth and / or remaining power of the computing node. The status information may be used to determine at least one first computing node. Exemplarily, in S402, the computing power evaluation function can screen the first computing node from the M computing nodes based on the first information and the computing power information. For example, the computing power evaluation function can obtain a list of alternative computing nodes based on the first information of each computing node in the M computing nodes, and the list of alternative computing nodes includes computing nodes deployed with the second model. Further, the computing power evaluation function can determine any computing node in the list of alternative computing nodes as the first computing node, or the computing power evaluation function can screen one or more first computing nodes in the list of alternative computing nodes based on the status information of each computing node in the list of alternative computing nodes. For example, the computing power evaluation function uses the first one or more computing nodes in the alternative computing nodes arranged in descending order of network transmission bandwidth as the first computing node, or the computing power evaluation function uses the first one or more computing nodes in the alternative computing nodes arranged in descending order of remaining power as the first computing node, or the computing power evaluation function uses the first one or more computing nodes in the alternative computing nodes arranged in descending order of the weighted sum of network transmission bandwidth and node remaining power as the first computing node.
[0194] Optionally, the information reported by each of the N computing nodes to the computing power evaluation function may also include an identifier of the computing node, where the identifier of the computing node is used to uniquely identify a computing node.
[0195] It should be noted that the identifier of the computing node is only an example and is not limited to this application. It can be replaced by any other information that can indicate the computing node. The type identifier of any task (such as the first task or the benchmark task) below is also an example and can be replaced by any other information that can indicate the type of the task.
[0196] Optionally, the computing power evaluation function may save information reported by each of the N computing nodes to the status aggregation function.
[0197] In S403, the computing power evaluation function sends a first request to each first computing node in the computing node group determined in S402 based on at least one first computing node. The first request is used to request the first computing node to send information of the second model. In S404, each first computing node in the at least one first computing node can send a first response to the computing power evaluation function in response to the first request. The first response includes information of the second model.
[0198] The information of the second model may include model structure information and / or model parameter information. Generally speaking, the second model can be expressed based on the information of the second model. In other words, the computing power evaluation function can construct the second model based on the received information of the second model. It should be understood that model information not included in the information of the second model can be agreed upon by the protocol, pre-configured, or preset in the computing power evaluation function. For example, the model structure information can be agreed upon by the protocol, pre-configured, or preset in the computing power evaluation function.
[0199] In S406, the computing power evaluation function may determine the first model of the computing node group to which the at least one first computing node belongs based on the information of the second model sent by the at least one first computing node. In some embodiments, when the computing power evaluation function receives the information of the second model sent by a first computing node, the second model may be used as the first model; when the computing power evaluation function receives the information of the second model sent by at least two first computing nodes respectively, any one of the at least two second models may be used as the first model, or the first model may be determined based on the information of the at least two second models, for example, the model parameters of the at least two second models may be merged to obtain the first model. This application does not limit the implementation method of the model parameter merging. For example, the model parameters of the at least two second models may be quantized and averaged to achieve the merging of the model parameters. In other embodiments, the computing power evaluation function can determine the first model based on the information of the second model sent by at least one first computing node and the information of the known model. Similar to the above, the computing power evaluation function can use at least one second model and any one of the known models as the first model, or merge the model parameters based on the information of at least one second model and the known model to obtain the first model, wherein the known model can be a model pre-acquired by the computing power evaluation function, such as a model pre-trained by the computing power evaluation function or received from other devices or functional modules.
[0200] Exemplarily, the first response may also carry a model description of the second model. The model description of the second model is used to indicate the type of input data of the second model and the type of output data of the second model. The type of input data of the second model is similar to the type of input data of the first model, and the type of output data of the second model is similar to the type of output data of the first model. For the sake of brevity, these details are not repeated here. The model description of the first model can be determined based on the model description of the second model.
[0201] For example, when the first response does not carry the model description of the second model, the model description may be agreed upon in the protocol, pre-configured, or preset in the device, and this application does not limit this.
[0202] Optionally, the first request may further include third indication information, which is used to indicate whether the first computing node updates the second model. When the third indication information indicates that the first computing node updates the second model, the first computing node may update the second model (or perform incremental learning); when the third indication information indicates that the first computing node does not update the second model, the first computing node does not update the second model, or the first computing node assists the computing power evaluation function in updating the first model.
[0203] The embodiment of the present application provides two possible methods for incremental learning of the first model. The incremental learning process of the first model is exemplarily described below in conjunction with Figure 4.
[0204] Update method 1:
[0205] S413-1, the first computing node updates the second model. Exemplarily, the first computing node may update the second model based on a second training data set, where each second training data in the second training data set includes a feature and a label, where the label is a measured value of a task processing time corresponding to the feature.
[0206] Exemplarily, the first computing node can construct the second training data set. For example, after receiving the scheduling information, the first computing node can load the corresponding execution unit and, in combination with the resource parameters in the first task information, create an execution unit instance, and then execute the corresponding task. When executing the task, the start time and completion time of the task are recorded to obtain a measurement value of the task processing duration of the task. The computing power information of the computing node and the first task information are the features of the second training data, and the measurement value of the task processing duration is the label of the second training data.
[0207] Optionally, the label of the second training data in the second training data set may also be input by the user, which is not limited in this application.
[0208] Optionally, the first computing node may determine to update the second model when the third indication information instructs the first computing node to update the second model.
[0209] Optionally, when determining to update the second model, the first computing node performs incremental learning to update the second model when determining that the number of second training data in the second training data set is greater than a threshold. The threshold can be any value such as 1, 10, or 100.
[0210] Optionally, to improve the prediction accuracy of the second model, the first computing node may perform incremental learning on unexecuted tasks. Exemplarily, the computing power evaluation function identifies the first task information and, upon determining that the first task information corresponds to an unexecuted first task, sends fourth indication information to the first computing node, instructing the first computing node to generate a label for the second training data for the first task. Optionally, the fourth indication information may include a type identifier for the first task.
[0211] Optionally, the computing power evaluation function may save the correspondence between the type identifier of the first task and the first task information to the status aggregation function.
[0212] S414-1, the first computing node sends the updated second model information to the computing power evaluation function.
[0213] S415-1: The computing power evaluation function updates the first model based on the information of the updated second model. For example, the computing power evaluation function may use the updated second model as the updated first model, or the computing power evaluation function may determine the updated first model based on multiple updated second models, such as by merging model parameters of multiple updated second models to obtain the updated first model.
[0214] Update method 2:
[0215] At step S413-2, the first computing node determines a measured value of the task processing duration corresponding to the first task information. In other words, the first computing node determines a label for the second training data. Optionally, the first computing node may construct the second training data based on the measured value of the task processing duration and add the second training data to the second training dataset.
[0216] At step S414-2, the first computing node may send the measured value of the task processing duration corresponding to the first task information to the computing power evaluation function. Optionally, if the first computing node constructs second training data based on the measured value of the task processing duration, step S414-2 may be replaced by the first computing node sending the second training data or a second training dataset to the computing power evaluation function.
[0217] S415-2, the computing power evaluation function can train the first model according to the measured value of the task processing time corresponding to the first task information to obtain an updated first model.
[0218] Exemplarily, the computing power evaluation function may construct second training data based on the measured value of the task processing time corresponding to the first task information, and add the second training data to the second training data set, thereby updating the first model based on the second training data set. Optionally, the computing power evaluation function may update the first model based on the second training data set, which may include: directly using the second training data set to iteratively train the first model, or training the second model sent by the corresponding computing node based on the second training data set to obtain an updated second model, and then updating the first model based on the updated second model. This application does not limit this.
[0219] It should be noted that the construction method of the second training data and the second training data set in the second updating method can refer to the description of the above-mentioned updating method one, which will not be repeated for the sake of brevity.
[0220] Optionally, when the first computing node sends the second training data or the second training data set to the computing power evaluation function, the computing power evaluation function does not need to execute the process of constructing the second training data set.
[0221] Optionally, the label of the second training data in the second training data set may also be input by the user, which is not limited in this application.
[0222] Optionally, when the third indication information indicates that the first computing node does not update the second model, the first computing node may determine a measured value of the task processing duration of the first task information and send the measured value of the task processing duration.
[0223] Optionally, the computing power evaluation function performs incremental learning to update the first model when determining that the number of second training data in the second training data set is greater than a threshold. The threshold can be any value such as 1, 10, or 100.
[0224] Optionally, to improve the prediction accuracy of the second model, incremental learning can be performed for unexecuted tasks. Exemplarily, the computing power assessment function identifies the first task information and, upon determining that the first task information corresponds to an unexecuted first task, sends fourth indication information to the first computing node, instructing the first computing node to generate a label for the second training data for the first task. Optionally, the fourth indication information may include a type identifier for the first task.
[0225] Optionally, the computing power evaluation function may save the correspondence between the type identifier of the first task and the first task information to the status aggregation function.
[0226] In some scenarios, the computing power assessment function expects to obtain the second model based on a computing node among the M computing nodes that does not have the second model deployed. For example, a computing node that does not have the second model deployed is more likely to be used for task processing, such as if the computing node has a large network transmission bandwidth and / or a large amount of remaining power. Therefore, considering that the first model corresponding to the computing node group is determined based on the second model of the computing node, the task processing duration can be more accurately predicted. For another example, none of the M computing nodes have the second model deployed.
[0227] In view of this, the computing power evaluation function can request at least one of the M computing nodes to obtain a second model through model training. This embodiment is described below with reference to Figure 5. S501 to S502 in Figure 5 are similar to S401 and S402 in Figure 4, except that the second model is not deployed in the first computing node.
[0228] At step S503, the computing power evaluation function sends a second request to the first computing node. This second request carries first instruction information for training the second model. This first instruction information indicates whether the first computing node has trained the second model. The following describes two possible training methods based on the instructions in the first instruction information.
[0229] Training method 1:
[0230] S504-1: The first computing node trains and obtains a second model based on the second request. Exemplarily, when the first instruction information instructs the first computing node to train and obtain the second model, the first computing node trains and obtains the second model based on the first training dataset. Each first training data point in the first training dataset includes a feature and a label, where the label is a measured value of the task processing duration corresponding to the feature.
[0231] Exemplarily, the first computing node can construct the first training data set. For example, the second request can carry test task information and / or second indication information, and the second indication information is used to indicate the generation of labels in the first training data set. Optionally, the second indication information can carry a type identifier of the benchmark test task. The test task information includes at least one of resource parameters, task parameters, and input data information. The type identifier of the benchmark test task is used to determine the benchmark test task, and the benchmark test task is used to generate a measurement value of the task processing time based on the computing power information and test task information of the first computing node.
[0232] During the process of constructing the first training data set on the first computing node, the first computing node may use the computing power information and test task information of the first computing node as inputs to the benchmark test task model, and output the measured value of the task processing time through the benchmark test task model. The computing power information and test task information of the first computing node serve as features of the first training data, and the measured value of the task processing time serves as a label for the first training data.
[0233] Optionally, if the second request does not include resource parameters, the first computing node can determine all possible resource parameters, such as all possible combinations of resource types and resource quantities, by traversing the list of available resources in the first computing node. If there are many possible resource parameters, possible combinations of resource parameters can be screened to improve processing efficiency.
[0234] Optionally, if the second request does not carry information about the input data, the first computing node may generate information about all possible input data based on a preset threshold of the input data. If there is a large amount of possible input data, the possible values of the input data information may be filtered to improve processing efficiency.
[0235] Optionally, the second request may also carry a model structure of a second model, and the first computing node trains the second model based on the first training data set and the model structure of the second model.
[0236] Optionally, the second request may further carry a model description of the second model. The model description is similar to the model description in the above example and will not be repeated for brevity.
[0237] For example, the first computing node can download the execution unit of the benchmark task based on the type identifier of the benchmark task, then traverse the resource parameters and input data information. For each possible two-tuple consisting of resource parameter and input data information, it combines the resource parameters to create an execution unit instance of the benchmark task, combines the input data information, executes the benchmark task, and records the start and completion time of the benchmark task to obtain the measured value of the task processing duration of the benchmark task. At the end of the traversal, a four-tuple list consisting of the benchmark task type identifier, resource parameters, input data information, and the measured value of the task processing duration is obtained.
[0238] Optionally, when the first computing node determines to train the second model, the first computing node performs model training to obtain the second model when the number of first training data in the first training data set is determined to be greater than a threshold. The threshold can be any value such as 1, 10, or 100.
[0239] Optionally, the labels in the first training data set may also be input by a user or sent by other devices.
[0240] S505-1, the first computing node may send a second response to the computing power evaluation function, where the second response includes information about the second model.
[0241] Training method 2:
[0242] S504-2: The first computing node determines a first training dataset for training the second model based on the second request. Optionally, when the first instruction information instructs the first computing node not to train the second model or to obtain the first training dataset, the first computing node determines the first training dataset for training the second model. It should be understood that this step is identical to the process of the first computing node constructing the first training dataset in training method one and is not further described for the sake of brevity.
[0243] S505-2, the first computing node sends a second response to the computing power evaluation function, where the second response includes the above-mentioned first training data set.
[0244] S506-2: The computing power evaluation function trains the first training data set to obtain a second model. It should be understood that the process of training the computing power evaluation function to obtain the second model based on the first training data set is similar to the process of training the first computing node to obtain the second model described above, and will not be repeated for the sake of brevity.
[0245] S507 and S508 in FIG5 are similar to S405 and S406 in the embodiment shown in FIG4 , and are not described again for the sake of brevity.
[0246] The scheduling process and incremental learning process in the embodiment shown in FIG5 can refer to the relevant description in FIG4 , which will not be repeated for the sake of brevity.
[0247] As mentioned above, how to achieve management, orchestration, and scheduling of heterogeneous computing nodes has received increasing attention. Under the condition of clear task processing time requirements, accurate prediction of the task resource information of the computing nodes will help to achieve the management, orchestration, and scheduling of the computing nodes, where the task resource information may include the type and quantity of resources used by the computing nodes to perform tasks. To address this issue, the embodiment of the present application introduces a fourth model corresponding to the computing nodes, and predicts the task resource requirements of the computing node group based on the fourth model and the computing power information of the computing node group and the second task information, thereby achieving accurate prediction of the task resource information of the computing node group when performing different tasks.
[0248] It should be understood that the present application can be applied to the prediction of task resource information for computing nodes in any computing resource pool, and is particularly applicable to the prediction of task resource information for heterogeneous computing nodes.
[0249] The following describes the method for predicting task resource information provided by the embodiment of the present application with reference to the accompanying drawings.
[0250] It should be understood that the following description is for ease of understanding and explanation only, and the method provided in the embodiments of the present application is described using the interaction between a first device and a second device as an example. The first device may be deployed with a functional module that implements a computing power evaluation function, and the second device may be deployed with functional modules that implement other functions, such as a functional module for a scheduling function and / or a functional module for an orchestration function. For example, both the first device and the second device may be the RAN node 110 in Figure 1, or devices in the core network 200.
[0251] It should also be understood that this should not constitute any limitation on the execution subject of the method provided in this application. As long as it is possible to execute the method provided in the embodiment of this application by running a program having the code of the method provided in the embodiment of this application, it can serve as the execution subject of the method provided in the embodiment of this application. For example, the above-mentioned first device may also be a component in a device (such as a device in the RAN node 110 or the core network 200), such as a chip, a chip system or other functional module that can call a program and execute a program; the above-mentioned second device may also be a component in a device (such as a device in the RAN node 110 or the core network 200), such as a chip, a chip system or other functional module that can call a program and execute a program.
[0252] FIG6 is a schematic diagram of an interactive flow of a method for predicting task resource information provided by an embodiment of the present application. Referring to FIG6 , the method 600 includes at least some of the following steps.
[0253] S610: The second device sends a prediction request for task resource information to the first device, where the prediction request carries second task information. Correspondingly, the first device receives the prediction request for task resource information from the second device.
[0254] S620, the first device predicts the task resource information of the computing node group according to the computing power information and the second task information of the computing node group through the fourth model, and the fourth model is determined based on the computing node group.
[0255] S630: The first device sends a second prediction response to the second device. The second prediction response carries task resource information corresponding to the computing node group. Correspondingly, the second device receives the second prediction response from the first device.
[0256] The second task information carries relevant information of the task to be executed, so that the first device can determine the task resource information of the computing node group for the specific task, thereby improving the accuracy of the prediction.
[0257] Illustratively, the second task information may include at least one of the following information:
[0258] 1. Expected duration, which is associated with task resource information.
[0259] 2. Task parameters
[0260] 3. Information on input data.
[0261] The information of task parameters and input data has been described in the above examples and will not be repeated for the sake of brevity.
[0262] It should be understood that the information included in the second task information, such as expected duration, task parameters, and input data, is associated with the task resource information.
[0263] It should also be understood that when the second task information does not carry one or more of the information in the above examples, the uncarried information may be preset or pre-agreed. For example, when the second task information does not carry input data information, the data input to the third model may be unit size data.
[0264] Exemplarily, the prediction request in S610 may further carry a task type identifier, which has been described in the above example and will not be repeated for brevity.
[0265] Exemplarily, a prediction request for task resource information can be sent by a demander of the task resource information prediction. For example, the scheduling function can send the prediction request to obtain the predicted task resource information, and then schedule the task based on the predicted task resource information; for another example, the orchestration function can send the prediction request to obtain the predicted task resource information, and then orchestrate the task based on the predicted task resource information; for another example, when the task resource information acts on the orchestration and scheduling, the orchestration function and the scheduling function can collaboratively send the prediction request, such as the orchestration function can send the prediction request to obtain the predicted task resource information, and after arranging the task based on the predicted task resource information, send the job description information to the scheduling function, and the scheduling function will schedule the task. This application does not limit this.
[0266] The computing node group may include M computing nodes. As previously mentioned, the computing nodes in the computing node group may be related or unrelated. Whether the computing nodes in the computing node group are related can be described in the above example, and will not be repeated for the sake of brevity.
[0267] The fourth model can be a pre-trained AI model for predicting task resource information. The input of the fourth model may include the computing power information of the computing node group and the second task information, and the output of the fourth model may include the task resource information. The computing power information of the computing node group can be found in the description of the previous example and will not be repeated for the sake of brevity.
[0268] The task resource information output by the fourth model, that is, the task resource information of the computing node group, may refer to the task resource information when any one of the M computing nodes processes the task corresponding to the second task information.
[0269] It should be noted that the embodiments of the present application do not limit the number of computing node groups. The second prediction request may include multiple second task information. Different second task information may correspond to different computing node groups, that is, different types of tasks may be executed by computing nodes in different computing node groups. The first device may predict the task resource information of the computing node group based on the fourth model corresponding to each computing node group, the computing power information of the computing node group and the second task information corresponding to the computing node group.
[0270] The first device sends the predicted task resource information to the second device, and the second device can schedule and / or arrange the task based on the task resource information. This application does not limit the first device to sending the task resource information to the second device. For example, the first device can also send the task resource information to other devices, such as sending the task resource information to a display device for presenting the task resource information.
[0271] Exemplarily, the task resource information can be carried by a second prediction response, and the first device sends the second prediction response to the second device to transmit the task resource information of the computing node group. Optionally, the second prediction response includes the identifier of the computing node group and the task resource information. For example, the second prediction response may include a two-tuple consisting of the identifier of the computing node group and the task resource information. When the second prediction response includes multiple task resource information of multiple computing node groups, the second prediction response may include a two-tuple list consisting of the identifiers of the multiple computing node groups and the task resource information corresponding to each computing node group.
[0272] In an embodiment of the present application, the first device predicts the task resource information of the computing node group based on the fourth model, the computing power information of the computing node group and the second task information, thereby achieving accurate prediction of the task resource information of the computing node group when performing different tasks.
[0273] Figure 7 is a schematic diagram of the interactive flow of a method for predicting task resource information provided by an embodiment of the present application. In Figure 7, the first device is used as a computing power evaluation function, and the second device is used as an orchestration function as an example, and the embodiment of the present application is described in combination with other devices in the network (such as computing nodes). Among them, the workflow sender can be a user or other functional module. It should be understood that the functions and devices shown in Figure 7 are only examples and do not constitute any limitation to the present application. The present application may include more or fewer functions and devices than Figure 7.
[0274] The following describes the scheduling process in conjunction with Figure 7, using the example of using predicted task resource information to implement orchestration. Referring to Figure 7, S709 to S711 are implemented in the same or similar manner as S610 to S630 in the embodiment shown in Figure 6. The technical solutions associated with S610 to S630 are also applicable to this embodiment.
[0275] Exemplarily, before the orchestration function sends the prediction request for task resource information to the computing power evaluation function, it may further include S708. In S708, the workflow sender may send workflow description information to the orchestration function, and correspondingly, the orchestration function receives the workflow description information sent by the workflow sender.
[0276] The workflow description information may include a task deadline template, which includes the second task information and the task type identifier in the aforementioned example. It should be understood that when a workflow includes multiple tasks, the workflow description information may include multiple task deadline templates, one or more of which may form a task deadline template list. The workflow description information may also include dependencies between the multiple tasks.
[0277] The orchestration function can parse the workflow description information to obtain task deadline templates corresponding to different tasks.
[0278] Optionally, the orchestration function determines, for each task deadline template, an identifier of at least one computing node group to form a preselected group identifier list. Exemplarily, the orchestration function may screen the computing node groups based on the expected duration in the task deadline template. For example, the orchestration function may determine, from multiple computing node groups, at least one computing node whose remaining resources meet preset resource parameters. The identifier of the computing node group to which the at least one screened computing node belongs may form the preselected group identifier list. The preset resource parameters may be agreed upon by a protocol, preconfigured, or preset in the orchestration function.
[0279] Optionally, in order to reduce processing complexity, some pre-selected computing node groups may be randomly selected from the pre-selected computing node groups indicated by the filtered pre-selected group identifier list to form a new pre-selected group identifier list.
[0280] Optionally, the orchestration function may perform a deduplication operation on the pre-selected group identifier list to save processing overhead.
[0281] Furthermore, the orchestration function can generate a prediction request for task resource information based on the results of parsing the workflow description information. The prediction request for task resource information can carry a list of pre-selected group identifiers. The computing power evaluation function predicts the task resource information of the corresponding computing node group based on the pre-selected group identifier list in the prediction request for task resource information. When the prediction request for task resource information does not carry the pre-selected group identifier list, the computing power evaluation function can use all computing node groups as pre-selected computing node groups, or the computing power evaluation function can filter pre-selected computing node groups from all computing node groups. The filtering logic is similar to the filtering logic of the orchestration function in the aforementioned example and will not be repeated for the sake of brevity.
[0282] It should be understood that the embodiments of the present application do not limit the screening logic or screening algorithm for the pre-selected computing node group (or pre-selected group identifier).
[0283] Optionally, the orchestration function may save the correspondence between the task type identifier and the task time limit template through the state aggregation function.
[0284] For example, when predicting the task resource information of a computing node group, the computing power evaluation function can first determine the fourth model corresponding to the computing node group and the computing power information corresponding to the computing node group, and then use the computing power information and the second task information of the computing node group as inputs of the fourth model to predict the task resource information.
[0285] Optionally, the state aggregation function may store a fourth correspondence, in which each computing node group in at least one computing node group has a corresponding fourth model, and the computing power evaluation function may read the fourth correspondence from the state aggregation function, and then determine the fourth model corresponding to the computing node group to be predicted (such as the above-mentioned pre-selected computing node group) based on the fourth correspondence. Optionally, the state aggregation function may store a fifth correspondence, in which each computing node group in at least one computing node group has corresponding computing power information, and the computing power evaluation function may read the fifth correspondence from the state aggregation function, and then determine the computing power information corresponding to the computing node group to be predicted (such as the above-mentioned pre-selected computing node group) based on the fifth correspondence.
[0286] Optionally, in the fourth corresponding relationship, each computing node group in the at least one computing node group further corresponds to a model description of the fourth model, where the model description is used to indicate a type of input data and a type of output data of the fourth model.
[0287] Optionally, the type of input data includes but is not limited to at least one of the following:
[0288] Compute the computing power information of the node group;
[0289] expected duration;
[0290] Task parameters;
[0291] Information about input data.
[0292] Optionally, the type of output data includes but is not limited to task resource information.
[0293] Exemplarily, in S710, the computing power evaluation function can traverse each pre-selected group identifier in the pre-selected group identifier list, and obtain the fourth model and computing power information of the computing node group corresponding to the pre-selected group identifier according to the fourth correspondence and the fifth correspondence, respectively, and then use the computing power information and the second task information of the computing node group as the input of the fourth model to determine the task resource requirements of each computing node group, and then obtain a two-tuple list consisting of the computing node group identifier (or group identifier) and the task resource requirement.
[0294] Illustratively, in S712, the orchestration function may combine the two-tuple list consisting of the group identifier and the task resource information, and the compute nodes in the compute node group corresponding to each group identifier in the two-tuple list (e.g., the M compute nodes described above), and determine the orchestration information using a preset orchestration algorithm. Illustratively, the two-tuple list consisting of the group identifier and the task resource information may be used to construct a task template during the orchestration process; this application does not limit the use of the task resource information.
[0295] In some embodiments, the fourth model corresponding to the computing node group may be pre-trained by the computing power assessment function, as shown in S707 of FIG7 . For example, the fourth model may be trained based on the first model corresponding to the computing node group, where the first model is used to predict the task processing time of the computing node group. The method for obtaining the first model corresponding to the computing node group can be found in the aforementioned example and will not be further described for the sake of brevity.
[0296] Exemplarily, the computing power assessment function can be trained based on a third training dataset to obtain a fourth model, wherein labels in the third training dataset can be determined by the first model of the computing node group corresponding to the fourth model. Each third training data in the third training dataset includes a feature and a label, where the label is a measured value of the task resource information corresponding to the feature.
[0297] Exemplarily, the computing power evaluation function can construct a third training data set for training a fourth model corresponding to the computing node group based on the first model corresponding to the computing node group. For example, the computing power evaluation function obtains the computing power information and the first model of the computing node group corresponding to the computing node group by traversing the first corresponding relationship and the second corresponding relationship saved by the state aggregation function, and traverses the task templates in the corresponding relationship between the task type identifier and the task template saved in the state aggregation function to obtain a list of task templates corresponding to the computing node group. For each task template in the task template list, the computing power information of the computing node group and the first task information in the task template are used as inputs to the first model corresponding to the computing node group, and the measured value of the task processing time output by the first model is obtained, thereby obtaining a five-tuple list consisting of the computing power information, resource parameters, task parameters, input data information, and the measured value of the task processing time of the computing node group. Each five-tuple in the five-tuple list is used to construct a third training data, such as the computing power information, task parameters, input data information, and the measured value of the task processing time of the computing node group in the five-tuple as features in the third training data, and the resource parameters in the five-tuple as labels in the third training data.
[0298] Optionally, the computing power evaluation function may save the corresponding relationship between the identifier of the computing node and the fourth model to the state aggregation function.
[0299] It should be noted that the example shown in Figure 7 can be combined with any of the above embodiments. For example, the incremental learning process in the embodiment shown in Figure 4 can be applied to this embodiment; the training process of the second model in the embodiment shown in Figure 5 can also be applied to this embodiment.
[0300] For example, the computing power evaluation function can update the fourth model while updating the first model (or incrementally learning the first model). For example, the fourth model can be retrained based on the updated first model, or training data for the updated fourth model can be generated based on the training data of the updated first model, thereby incrementally learning the fourth model.
[0301] Figure 8 is a schematic block diagram of a communication device provided in an embodiment of the present application. The communication device 800 may be the first device in the above-described method embodiment. The communication device may be a communication device in a RAN or core network, or may be a device within a communication device, or may be a device capable of being used in conjunction with a communication device. As shown in Figure 8 , the device 800 may include a transceiver module 810 and a processing module 820.
[0302] When executing any of the method embodiments shown in Figures 3 to 5, the transceiver module 810 can be used to receive a prediction request for the task processing duration, which carries the first task information; the processing module 820 can be used to predict the task processing duration of the computing node group based on the computing power information and the first task information of the computing node group through a first model, where the computing node group includes M computing nodes, where M is a positive integer, and the first model is determined based on the computing node group.
[0303] When executing any of the method embodiments shown in Figures 6 to 7, the transceiver module 810 can be used to receive a prediction request for task resource information, which carries the second task information; the processing module 820 can be used to predict the task resource information of the computing node group based on the computing power information and the second task information of the computing node group through a fourth model, where the computing node group includes M computing nodes, where M is a positive integer, and the fourth model is determined based on the computing node group.
[0304] It should be understood that the specific process executed by each module has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0305] The transceiver module 810 in the communication device 800 can be implemented by a transceiver, for example, it can correspond to the transceiver 920 in the communication device 900 shown in Figure 9. The processing module 820 in the communication device 800 can be implemented by at least one processor, for example, it can correspond to the processor 910 in the communication device 900 shown in Figure 9.
[0306] When the communication device 800 is a chip or chip system configured in a communication device, the transceiver module 810 in the communication device 800 can be implemented through an input / output interface, circuit, etc., and the processing module 820 in the communication device 800 can be implemented through a processor, microprocessor or integrated circuit integrated on the chip or chip system.
[0307] Figure 9 is another schematic block diagram of a communication device provided in an embodiment of the present application. As shown in Figure 9, the communication device 900 may include: a processor 910. The processor 910 may be configured to execute the method executed by the first device in the above method embodiment.
[0308] In some possible implementations, the communication device 900 may include a transceiver 920. The transceiver 920 may communicate with the processor 910 via an internal connection path. The processor 910 may control the transceiver 920 to send and / or receive signals.
[0309] In some possible implementations, the communication device 900 may include a memory 930. The memory 930 may communicate with the processor 910 via an internal connection path. The memory 930 and the processor 910 may be integrated or provided separately. The memory 930 may also be a memory external to the device. The memory 930 is used to store instructions, and the processor 910 is used to execute the instructions stored in the memory 930 to perform the method in the above method embodiment.
[0310] It should be understood that the communication device 900 may correspond to the first device in the above-mentioned method embodiment, and may be used to execute the various steps and / or processes performed by the first device in the above-mentioned method embodiment. Optionally, the memory 930 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. The memory 930 may be a separate device or integrated into the processor 910. The processor 910 may be used to execute the instructions stored in the memory 930, and when the processor 910 executes the instructions stored in the memory, the processor 910 is used to execute the various steps and / or processes of the above-mentioned method embodiment corresponding to the first device.
[0311] The transceiver 920 may include a transmitter and a receiver. The transceiver 920 may further include an antenna, which may be one or more. The processor 910, memory 930, and transceiver 920 may be integrated on different chips. For example, the processor 910 and memory 930 may be integrated in a baseband chip, and the transceiver 920 may be integrated in a radio frequency chip. The processor 910, memory 930, and transceiver 920 may also be integrated on the same chip. This application does not limit this.
[0312] Alternatively, the transceiver 920 may also be a communication interface, such as an input / output interface, a circuit, etc. The transceiver 920 , the processor 910 , and the memory 930 may all be integrated into the same chip, such as a baseband chip.
[0313] The present application also provides a communication device, comprising at least one processor, wherein the at least one processor executes a computer program or logic circuit to cause the processing device to execute the method executed by the first device in the above method embodiment. The above communication device may also include a memory for storing the above computer program.
[0314] An embodiment of the present application further provides a communication device comprising a processor and an input / output interface. The input / output interface is coupled to the processor. The input / output interface is used to input and / or output information. The information includes at least one of instructions and data. The processor is configured to execute a computer program to cause the processing device to perform the method performed by the first device in the above method embodiment.
[0315] The present application also provides a communication device including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the processing device executes the method executed by the first device in the above method embodiment.
[0316] It should be understood that the processing device may be one or more chips. For example, the processing device may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0317] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.
[0318] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0319] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0320] According to the method provided in the embodiment of the present application, the present application also provides a computer program product, which includes: a computer program or a set of instructions, which, when the computer program or a set of instructions is run on a computer, enables the computer to execute the method executed by the first device in the above method embodiment.
[0321] According to the method provided in the embodiment of the present application, the present application also provides a computer-readable storage medium, which stores a program. When the program is run on a computer, the computer executes the method executed by the first device in the above method embodiment.
[0322] According to the method provided in the embodiment of the present application, the present application also provides a communication system, which may include the aforementioned first device. Optionally, the communication system may also include the aforementioned second device, a computer node, etc.
[0323] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0324] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0325] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for predicting the task processing duration, characterized in that Including: Receiving a prediction request for the task processing duration, where the prediction request carries task information; Using a first model to predict the task processing duration of the computing node group based on the computing power information of the computing node group and the task information. The computing node group includes M computing nodes, where M is a positive integer, and the first model is determined based on the computing node group.
2. The method according to claim 1, wherein The task information includes resource parameters, and the resource parameters indicate the type and quantity of resources for performing task processing.
3. The method according to claim 1 or 2, characterized in that, The task information includes task parameters, and the task parameters are used to indicate the parameters adopted for task processing, and the task parameters are associated with the task processing duration.
4. The method according to claim 3, wherein The task information includes information on input data, and the information on input data indicates the size of the input data required for task processing.
5. The method according to any one of claims 1 to 4, characterized in that At least one computing power parameter of the M computing nodes is the same.
6. The method according to any one of claims 1 to 5, characterized in that The first model is determined based on a second model in at least one first computing node. The second model is used to predict the task processing duration of the corresponding first computing node, and the M computing nodes include the at least one first computing node.
7. The method according to claim 6, characterized in that Further including: Sending a first request, where the first request is used to request the sending of information on the second model; Receiving a first response, where the first response includes the information on the second model.
8. The method according to claim 6 or 7, characterized in that Further including: Receiving first information from each of the N computing nodes. The first information includes information indicating whether the computing node includes the second model. The N computing nodes include the M computing nodes; Determining the at least one first computing node among the M computing nodes according to the N pieces of first information.
9. The method according to any one of claims 6 to 8, characterized in that, Further including: Receiving the computing power information from each of the N computing nodes. The computing power information includes at least one computing power parameter of the corresponding computing node; Determining the computing node group and the computing power information of the computing node group according to the computing power information of each of the N computing nodes.
10. The method according to any one of claims 6 to 9, characterized in that Further including: Sending a second request, where the second request carries first indication information for training the second model; Receiving a second response; where When the first indication information indicates training to obtain the second model, the second response includes the information on the second model, Or, when the first indication information indicates obtaining a first training data set, the second response includes the first training data set used to train to obtain the second model. Each first training data in the first training data set includes features and labels, and the label is a measured value of the task processing duration corresponding to the feature.
11. The method according to any one of claims 1 to 10, characterized in that, Further including: Sending a first prediction response, where the first prediction response carries the information on the task processing duration corresponding to the computing node group.
12. The method according to claim 11, wherein The first prediction response includes the identifier of the computing node group and the information on the task processing duration.
13. The method according to any one of claims 1 to 12, characterized in that Further including: Sending third indication information; When the third indication information indicates updating the second model, receiving the information on the updated second model sent by the first computing node; Updating the first model according to the information on the updated second model; Or, When the third indication information indicates not to update the second model, receive the measured value of the task processing duration corresponding to the task information sent by the first computing node; Train the first model based on the measured value of the task processing duration corresponding to the task information to obtain an updated first model; Wherein, the first computing node belongs to the M computing nodes.
14. The method according to claim 13, wherein The first request carries the third indication information, and the first request is used to request the information of the second model.
15. A communication device, characterized in that, Including: A transceiver module, configured to receive a prediction request for task processing duration, where the prediction request carries task information; A processing module, configured to predict the task processing duration of the computing node group based on the computing power information of the computing node group and the task information through the first model, where the computing node group includes M computing nodes, M is a positive integer, and the first model is determined based on the computing node group.
16. The device according to claim 15, characterized in that, The task information includes resource parameters, and the resource parameters indicate the type and quantity of resources for performing task processing.
17. The device according to claim 15 or 16, characterized in that, The task information includes task parameters, and the task parameters are used to indicate the parameters adopted for task processing, and the task parameters are associated with the task processing duration.
18. The device according to claim 17, characterized in that, The task information includes information of input data, and the information of input data indicates the size of the input data required for task processing.
19. The device according to any one of claims 15 to 18, characterized in that, At least one computing power parameter of the M computing nodes is the same.
20. The device according to any one of claims 15 to 19, characterized in that, The first model is determined based on a second model in at least one first computing node, and the second model is used to predict the task processing duration of the corresponding first computing node, and the M computing nodes include the at least one first computing node.
21. The device according to claim 20, wherein, The transceiver module is further configured to: Send a first request, where the first request is used to request the information of the second model; Receive a first response, and the first response includes the information of the second model.
22. The device according to claim 20 or 21, wherein The transceiver module is further configured to receive first information from each of the N computing nodes, where the first information includes information indicating whether the computing node includes the second model, and the N computing nodes include the M computing nodes; The processing module is further configured to determine the at least one first computing node among the M computing nodes according to the N pieces of first information.
23. The device according to any one of claims 20 to 22, wherein The transceiver module is further configured to receive the computing power information from each of the N computing nodes, and the computing power information includes at least one computing power parameter of the corresponding computing node; The processing module is further configured to determine the computing node group and the computing power information of the computing node group according to the computing power information of each of the N computing nodes.
24. The device according to any one of claims 20 to 23, characterized in that, The transceiver module is further configured to: Send a second request, where the second request carries first indication information for training the second model; Receive a second response; wherein, When the first indication information indicates training to obtain the second model, the second response includes the information of the second model. Alternatively, when the first indication information indicates obtaining a first training data set, the second response includes the first training data set for training the second model. Each first training data in the first training data set includes features and labels, and the label is a measured value of the task processing duration corresponding to the features.
25. The device according to any one of claims 15 to 24, characterized in that, The transceiver module is further configured to: Send a first prediction response, where the first prediction response carries information about the task processing duration corresponding to the computing node group.
26. The device according to claim 25, characterized in that, The first prediction response includes the identifier of the computing node group and the information about the task processing duration.
27. The device according to any one of claims 15 to 26, characterized in that, The transceiver module is further configured to: Send third indication information; When the third indication information indicates updating the second model, receive the information of the updated second model sent by the first computing node; and update the first model according to the information of the updated second model; Or, When the third indication information indicates not to update the second model, receive the measured value of the task processing duration corresponding to the task information sent by the first computing node; Train the first model according to the measured value of the task processing duration corresponding to the task information to obtain an updated first model; Wherein, the first computing node belongs to the M computing nodes.
28. The device according to claim 27, characterized in that, The first request carries the third indication information, and the first request is used to request the sending of the information of the second model.
29. The device according to any one of claims 15 to 28, characterized in that The processing module is deployed in a centralized unit CU or a distributed unit DU.
30. A communication device, characterized in that, It includes: A processor, which is configured to execute the method according to any one of claims 1 to 14 by running a computer program or through logic circuits.
31. The device according to claim 30, characterized in that, It further includes a memory, which is configured to store the computer program.
32. The device according to claim 30 or 31, characterized in that, It further includes a communication interface, which is configured to input and output signals.
33. A computer-readable storage medium, characterized in that, For storing computer program instructions, the computer program causes the computer to execute the method according to any one of claims 1 to 14.
34. A computer program product, characterized in that, It includes computer program instructions, and the computer program instructions cause the computer to execute the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Task processing duration prediction method and device, equipment and storage medium
CN120234219A
Task processing method and device, computer equipment and storage medium
CN111104222A
Mass point cloud data processing method, device and system and server
CN113934535A
Resource scheduling method and device, cloud platform, equipment and storage medium
CN115794337A
Optimization method and optimization device for computing power resource allocation, electronic equipment and medium
CN116541176A