Control method and device of heterogeneous computing system, equipment, medium and product

By monitoring and migrating expert networks in heterogeneous computing systems, the inefficiency problem caused by slow computing nodes is solved, and more efficient task execution and resource utilization are achieved.

CN120704839APending Publication Date: 2025-09-26SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510883789.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When executing reasoning tasks based on hybrid expert models in heterogeneous computing systems, they are limited by computing nodes with slower computing speeds, resulting in low efficiency in reasoning task execution.

Method used

By monitoring the activation state parameters and resource status information of the expert network on the computing nodes of the heterogeneous computing system, time consumption prediction is performed, and the expert network reasoning tasks to be migrated are migrated to the computing nodes with smaller time consumption prediction results.

Benefits of technology

It improves the task execution efficiency and resource utilization of heterogeneous computing systems, dynamically responds to computing bottlenecks, and ensures the efficient execution of inference computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704839A_ABST
    Figure CN120704839A_ABST
Patent Text Reader

Abstract

The invention discloses a control method, device and equipment of a heterogeneous computing system, a medium and a product, and relates to the technical field of computers. Activation state parameters of an expert network are monitored through the expert network of a hybrid expert model deployed for computing nodes of the heterogeneous computing system; according to the activation state parameter of the expert network, time consumption prediction of the calculation node for executing the reasoning task is carried out, and accurate prediction of time consumption of the calculation node for executing expert network calculation in the heterogeneous calculation system is realized; and executing migration of the expert network according to the time consumption prediction result so as to migrate the expert network from the node to the computing node with a smaller time consumption prediction result, so that a hybrid expert model deployment scheme for enabling the heterogeneous computing system to execute more reasoning computing tasks can be obtained, the resource utilization rate is improved, and the computing efficiency is improved. And the calculation bottleneck can be responded in advance in the operation process of the heterogeneous calculation system, and the expert network reasoning task is dynamically migrated to ensure the efficient execution of reasoning calculation based on the hybrid expert model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a control method, device, equipment, medium and product for a heterogeneous computing system. Background Art

[0002] With the development of artificial intelligence (AI) technology, the scale of neural network models used has continued to increase, leading to the need to deploy large-scale neural network models using heterogeneous computing systems. The Mixture of Experts (MoE) is a machine learning approach that partitions an AI model into separate subnetworks (also called "expert networks"), each specializing in a subset of the input data to jointly perform a task. In a heterogeneous computing system, the different expert networks of the MoE are deployed on multiple computing devices, with each device responsible for running a portion of the expert model. This fully utilizes the computing resources of the heterogeneous computing system and improves the efficiency of inference tasks through parallelization. In practical applications, when heterogeneous computing systems perform inference tasks based on the MoE, they are limited by the slower computing nodes, resulting in low inference efficiency. Summary of the Invention

[0003] The present invention provides a control method, device, equipment, medium and product for a heterogeneous computing system, so as to at least solve the problem in the related art that when a heterogeneous computing system performs reasoning tasks based on a hybrid expert model, it is affected by computing nodes with slower computing speed, resulting in low execution efficiency of the reasoning tasks.

[0004] The present invention provides a control method for a heterogeneous computing system, comprising: Determine an expert network of hybrid expert models deployed on computing nodes of a heterogeneous computing system; monitoring activation state parameters of the expert network; Calculating a time-consuming prediction result of the computing node executing the expert network reasoning task based on the activation state parameters of the expert network on the computing node and the resource state information of the computing node; Determining the expert network reasoning task to be migrated based on the time consumption prediction results of the plurality of computing nodes; Migrate the expert network reasoning task to be migrated from the computing node where it is located to the computing node with a smaller prediction result in terms of time consumption.

[0005] The present invention also provides a control device for a heterogeneous computing system, comprising: A task information collection module, used to determine the expert network of the hybrid expert model deployed by the computing nodes of the heterogeneous computing system; A system information collection module, configured to obtain resource status information of the computing node; A system status monitoring module, used for monitoring activation status parameters of the expert network; A calculation module is configured to calculate a time consumption prediction result of the computing node executing the expert network reasoning task based on the activation state parameters of the expert network on the computing node and the resource state information of the computing node; and determine the expert network reasoning task to be migrated based on the time consumption prediction results of the plurality of computing nodes; The task migration module is used to migrate the expert network reasoning task to be migrated from the computing node where it is located to the computing node with a smaller prediction result.

[0006] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned control methods for heterogeneous computing systems when executing the computer program.

[0007] The present invention also provides a non-volatile storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned control methods for heterogeneous computing systems are implemented.

[0008] The present invention also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned control methods for heterogeneous computing systems when executed by a processor.

[0009] Through the present invention, the expert network of the hybrid expert model deployed on the computing nodes of the heterogeneous computing system monitors the activation state parameters of the expert network, and predicts the time consumption of the computing nodes to execute reasoning tasks based on the activation state parameters of the expert network, thereby achieving accurate prediction of the time consumption of the computing nodes in the heterogeneous computing system to execute expert network calculations; the migration of expert network reasoning tasks is performed according to the time consumption prediction results, so as to migrate the expert network from the node where it is located to the computing node with a smaller time consumption prediction result. A deployment scheme of the hybrid expert model can be implemented to enable the heterogeneous computing system to execute more reasoning computing tasks, improve task execution efficiency and resource utilization, and also respond to computing bottlenecks in advance according to the distribution changes of reasoning computing tasks during the operation of the heterogeneous computing system, and dynamically migrate expert network reasoning tasks to ensure efficient execution of reasoning computing based on the hybrid expert model. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A flowchart of a first method for controlling a heterogeneous computing system provided by an embodiment of the present invention; Figure 2 A schematic diagram of the deployment architecture of the first hybrid expert model provided by an embodiment of the present invention; Figure 3 A schematic diagram of the deployment architecture of the second hybrid expert model provided by an embodiment of the present invention; Figure 4 This is a flowchart of a second method for controlling a heterogeneous computing system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0013] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0014] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0015] Here, some key terms used in the embodiments of the present invention are explained.

[0016] Mixture of experts (MoE) is a machine learning approach that partitions an artificial intelligence (AI) model into separate subnetworks (or "experts"), each specializing in a subset of the input data to collectively perform a task. MoE enables large-scale models, even those containing billions of parameters, to significantly reduce computational costs during pre-training and achieve faster performance during inference time. Generally speaking, it achieves this efficiency by selectively activating the specific experts needed for a particular task, rather than activating the entire neural network for each task.

[0017] The hybrid expert model mainly consists of two parts: the gating network and the expert network.

[0018] Each expert network is an independent neural network that has been trained to excel at specific reasoning and computational tasks. The expert network is typically a neural network whose algorithmic structure can be a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM) network, etc. It can also be a model built on an attention network, such as a transformer model, a bidirectional encoder representation from transformers (BERT) model, or a contrastive language-image pre-training (CLIP) model, etc., but the present invention is not limited to this.

[0019] The gating network is a key component of the hybrid expert model. Its role is to select the appropriate expert network for each input and assign weights to the output of each expert network, ensuring that only the expert model most suitable for the current task participates in the final prediction.

[0020] When using a hybrid expert model to perform inference computing tasks, multiple expert networks share a large model space. The input data for each inference computing task only activates a subset of the expert networks for computation. The gating network processes the input feature vector, calculating the weight of each expert network through linear transformation and the softmax function. The top k expert networks are then selected based on the weights. These selected expert networks process the assigned input data, and the outputs of all activated expert networks are weighted and summed according to the weights assigned by the gating network to produce the final output.

[0021] Expert parallel computing is a type of model parallel computing. By deploying different expert networks of a hybrid expert model on different computing nodes, parallel computing between expert networks can be achieved. This parallelization can significantly improve the inference speed and scalability of the model.

[0022] Hybrid expert models can handle distributed data with more complex distributions. However, depending on the requirements of the inference computing task, the expert networks activated on different computing nodes may be different, and the hardware resource status of the computing nodes themselves may also vary. This is especially true in recent years, with the gradual application of multi-heterogeneous computing systems. In these systems, heterogeneous computing power with different computing performance (such as different types of computing chips or computing cards) is integrated into the same distributed computing environment and collaborates to complete the distributed inference tasks of the hybrid expert model. This results in different computational times for different computing nodes when executing inference computing tasks using expert parallel computing. This in turn leads to the system being limited by slower computing nodes, which delays the time it takes to execute an inference computing task as a whole.

[0023] In order to solve the problem that when a heterogeneous computing system performs reasoning tasks based on a hybrid expert model, the execution efficiency of the reasoning tasks is low due to the influence of computing nodes with slower computing speed. The present invention provides a control scheme for a heterogeneous computing system. By monitoring the activation state parameters of the expert network of the hybrid expert model deployed on the computing nodes of the heterogeneous computing system, and predicting the time consumption of the computing nodes performing reasoning tasks based on the activation state parameters of the expert network, the time consumption of the computing nodes performing reasoning tasks in the heterogeneous computing system is accurately predicted; the migration of the expert network reasoning tasks is performed based on the time consumption prediction results, so as to migrate the expert network from the node where it is located to the computing node with a smaller time consumption prediction result. A hybrid expert model deployment scheme can be implemented to enable the heterogeneous computing system to perform more reasoning computing tasks, thereby improving task execution efficiency and resource utilization. It can also respond to computing bottlenecks in advance according to the distribution changes of reasoning computing tasks during the operation of the heterogeneous computing system, and dynamically migrate the expert network reasoning tasks to ensure the efficient execution of reasoning computing based on the hybrid expert model.

[0024] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the control solution of the heterogeneous computing system depends, the specific application environment architecture or specific hardware architecture is described here.

[0025] The control solution for heterogeneous computing systems provided by the present invention can be deployed based on heterogeneous computing systems. Heterogeneous computing systems include multiple computing devices (computing nodes), and different computing nodes may have different parameters such as computing core type, computing power, memory size, and communication bandwidth.

[0026] In the heterogeneous computing system targeted by the present invention, the computing core type of the computing node may include, but is not limited to, one or more of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a neural network processor (NPU), a microcontroller unit (MCU), and an application-specific integrated circuit (ASIC).

[0027] In the heterogeneous computing system targeted by the present invention, different computing nodes can be interconnected through a network, such as Ethernet; or they can be interconnected through a bus, such as a high-speed serial computer expansion bus (Peripheral Component Interconnect Express, PCIe) or an NVIDIA high-speed interconnect bus.

[0028] The control method for a heterogeneous computing system provided by the embodiment of the present invention can be applied to one or more computing nodes in the heterogeneous computing system, and can also be applied to a control node outside the heterogeneous computing system.

[0029] An embodiment of the present invention provides a control method for a heterogeneous computing system. The method is described in detail below in conjunction with the execution flow of the control method for a heterogeneous computing system.

[0030] Figure 1 A flowchart of a first method for controlling a heterogeneous computing system provided by an embodiment of the present invention; Figure 2 A schematic diagram of the deployment architecture of the first hybrid expert model provided by an embodiment of the present invention; Figure 3 A schematic diagram of the deployment architecture of the second hybrid expert model provided in an embodiment of the present invention.

[0031] like Figure 1 As shown, the control method of the heterogeneous computing system provided by the embodiment of the present invention may include: S101: determining an expert network of a hybrid expert model deployed by computing nodes of the heterogeneous computing system.

[0032] S102: Monitoring activation state parameters of the expert network.

[0033] S103: Calculate and obtain a time-consuming prediction result of the computing node executing the expert network reasoning task based on the activation state parameters of the expert network on the computing node and the resource state information of the computing node.

[0034] S104: Determine the expert network reasoning task to be migrated based on the time consumption prediction results of multiple computing nodes.

[0035] S105: Migrate the expert network reasoning task to be migrated from the computing node where it is located to a computing node with a shorter prediction result.

[0036] The heterogeneous computing system on which the embodiments of the present invention are based may be a multi-heterogeneous computing system, that is, different computing nodes in the heterogeneous computing system may have different resource status information and may be based on different computing cores.

[0037] In terms of the deployment of the gated network, the embodiments of the present invention can be oriented towards two deployment modes of the gated network.

[0038] In some optional implementations of the present invention, a gated network can be deployed on each computing node. The gated network on each computing node will dynamically decide which input data to assign to the local expert network based on the local input data, or forward the input data to the expert network on other computing nodes. Figure 2 As shown, expert network 1, expert network 2, and gating network are deployed on computing nodes 1 and 2. After input data 1 enters computing node 1, the gating network decides to activate the expert network 1 of the hybrid expert model, and then distributes input data 1 to the local expert model 1 and the expert model 1 of computing node 2. The outputs of the two expert networks 1 are then aggregated to obtain output data 1. After input data 2 enters computing node 2, the gating network decides to activate the expert network 2 of the hybrid expert model, and then distributes input data 2 to the local expert model 2 and the expert network 2 of computing node 1. The outputs of the two expert networks 2 are then aggregated to obtain output data 2.

[0039] In some other optional implementations of the present invention, the gated network can also be deployed on a centralized node. All input data of the hybrid expert model are first sent to the centralized node, and the gated network makes a unified routing decision, and then distributes the input data to the computing nodes where each expert network is located. Figure 3As shown in the figure, expert network 1 and expert network 2 are deployed on computing node 1, and expert network 3 and expert network 4 are deployed on computing node 2. The gating network makes decisions on the input data. After making a decision on input number 1, it determines to activate expert network 1, and then sends input data 1 to computing node 1 to activate expert network 1, obtaining output data 1. After making a decision on input number 2, it determines to activate expert network 4, and then sends input data 2 to computing node 2 to activate expert network 4, obtaining output data 2.

[0040] It should be noted that the above Figure 2 、 Figure 3 The deployment architecture shown is only used to illustrate the deployment method of the gated network. In the embodiment of the present invention, one computing node can deploy one or more expert networks, and the same expert network can be deployed on different computing nodes.

[0041] Embodiments of the present invention aim to proactively identify performance bottlenecks in heterogeneous computing systems by monitoring the activation state parameters of each expert network when a heterogeneous computing system executes an inference computing task based on a hybrid expert model. Based on the activation state parameters, the time required for a computing node to execute the next inference computing task is predicted, thereby enabling early detection of performance bottlenecks in the heterogeneous computing system. It is understood that when a heterogeneous computing system executes an inference computing task, the weights of each expert network are calculated based on the output of the gating network, and the expert network is activated accordingly. If the activation frequency of the expert network on a computing node is high while the node has fewer resources, the node will require more computation time than other nodes during the inference computing task, thereby reducing the efficiency of the heterogeneous computing system in executing the inference computing task. Therefore, in embodiments of the present invention, the expert network inference task on the computing node with the higher predicted time consumption is migrated to the computing node with the lower predicted time consumption. This ensures that the computation time consumed by each computing node in the inference computing task is consistent, preventing performance bottleneck computing nodes from affecting the inference computing efficiency of the heterogeneous computing system.

[0042] For S101 , the expert network of the hybrid expert model deployed by the computing nodes of the heterogeneous computing system is determined, that is, the expert network deployed by each computing node is determined.

[0043] In S102, monitoring the activation state parameter of the expert network may include monitoring the number of activations of the expert network on the computing node per unit time as the activation state parameter of the expert network. By monitoring the number of activations of the expert network per unit time, it is determined how many times each expert network is activated per unit time when the hybrid expert model performs the actual inference computing task.

[0044] The control method for a heterogeneous computing system provided in an embodiment of the present invention can be applied before deploying a hybrid expert model in a heterogeneous computing system. At this point, a deployment plan for the hybrid expert model in the heterogeneous computing system can be initialized, for example by evenly distributing each expert network to each computing node in the heterogeneous computing system without duplication. Monitoring the activation state parameters of the expert networks in S102 can be performed to pre-deploy the hybrid expert model, input the required inference computing tasks, and monitor the activation state parameters of the expert networks.

[0045] The control method for a heterogeneous computing system provided by an embodiment of the present invention can also be applied to the operation process after the hybrid expert model is deployed in the heterogeneous computing system. With the changes in the distribution of inference computing tasks and the changes in the resources of the heterogeneous computing system, the deployment scheme of the hybrid expert model previously determined may again encounter a computing performance bottleneck. At this time, the control method for a heterogeneous computing system provided by an embodiment of the present invention can be applied to dynamically adjust the deployment of the hybrid expert model in the heterogeneous computing system. At this time, monitoring the activation state parameters of the expert network in S102 can be to monitor the activation state parameters of the expert network at multiple historical moments adjacent to the current moment, so as to determine the computing performance bottleneck in the heterogeneous computing system based on the activation state parameters of the expert network.

[0046] Since the gating network determines which expert network is activated, monitoring the activation state parameters of the expert network in S102 may further include: monitoring the weight of the expert network output by the gating network; and determining the activation state parameters of the expert network according to the weight of the expert network.

[0047] For S103, resource status information of the computing nodes can be obtained by accessing the cluster management node of the heterogeneous computing system or by accessing each computing node separately. The resource status information of the computing nodes refers to the resources provided by the computing nodes to the hybrid expert model. The resource status information may include one or more of computing power resource parameters, memory resource parameters, and network resource parameters. The computing power resource parameter may be represented by the number of floating point operations per second (FLOPS) of the computing node. The network resource parameter may be represented by the communication bandwidth of the computing node's external links. The memory resource parameter may be represented by the memory capacity allocated by the computing node to the hybrid expert model.

[0048] Based on the resource status information of the computing node and the activation status parameters of the expert network deployed on the computing node, the computational time consumed by the expert network activated on the computing node can be predicted. If the number of activations of the expert network on the computing node per unit time is used as the activation status parameter of the expert network, then in S103, the computational time consumption prediction result of the computing node executing the expert network reasoning task based on the activation status parameter of the expert network on the computing node and the resource status information of the computing node may include: computing the computational time consumption prediction result of the computing node executing the expert network reasoning task of the expert network activated per unit time based on the number of activations of the expert network on the computing node per unit time and the resource status information of the computing node.

[0049] For S104, the expert network reasoning tasks to be migrated are determined based on the time prediction results of some or all computing nodes in the heterogeneous computing system, with the aim of migrating the expert network reasoning tasks on the computing nodes with larger time prediction results to the computing nodes with smaller time prediction results. The determined expert network reasoning tasks to be migrated can be the expert network to be migrated, that is, the network parameters of the expert network on the computing node with larger time prediction results are migrated as a whole to another computing node. The expert network reasoning tasks to be migrated can also be the activation requirements of the expert network to be migrated, that is, the activation frequency of the expert network on the computing node with larger time prediction results can be reduced, or the activation requirements of the expert network on the computing node with larger time prediction results can be shared by deploying the expert network on other computing nodes to share the activation requirements of the expert network on the computing node with larger time prediction results, thereby reducing the activation frequency of the expert network on the computing node with larger time prediction results, and realizing the calculation speed-up of the computing node with larger time prediction results.

[0050] In some optional implementations of the embodiments of the present invention, determining the expert network reasoning task to be migrated based on the time prediction results of multiple computing nodes in S104 may include: determining the computing node with the largest time prediction result and the computing node with the smallest time prediction result in the heterogeneous computing system; judging whether the determined computing node is the same as the previous moment; if not, determining the expert network with the smallest time activation state parameter of the computing node with the largest time prediction result as the expert network reasoning task to be migrated, and determining the computing node with the smallest time prediction result as the target computing node of the expert network reasoning task to be migrated.

[0051] That is to say, by continuously monitoring the activation state parameters of the expert network and the resource status information of the computing nodes, and predicting the time consumption of the computing nodes in executing the expert network reasoning tasks, the computing nodes that cause the performance bottleneck of the heterogeneous computing system, that is, the computing nodes with the largest time consumption prediction results, are determined. By repeatedly determining the computing nodes with the largest time consumption prediction results and the computing nodes with the smallest time consumption prediction results in the heterogeneous computing system and comparing them with the corresponding computing nodes determined at the previous moment, the result of controlling the convergence of the time consumption prediction results of each computing node in a reasoning computing task is finally achieved.

[0052] In other optional implementations of the embodiments of the present invention, determining the expert network reasoning task to be migrated based on the time consumption prediction results of multiple computing nodes may also include: determining the sum of the activation state parameters of the expert network on the computing node; if it is determined that there is a computing node that meets the threshold based on the sum of the activation state parameters of each computing node, then the expert network with the smallest activation state parameter in time of the computing node with the largest time consumption prediction result is determined as the expert network reasoning task to be migrated, and the computing node with the smallest time consumption prediction result is determined as the target computing node of the expert network reasoning task to be migrated. For example, if the sum of the largest sum of the activation state parameters and the smallest sum of the activation state parameters in the computing node is greater than the first threshold, it is necessary to determine the expert network reasoning task to be migrated in order to perform the migration of the expert network reasoning task.

[0053] In some other optional implementations of the embodiments of the present invention, determining the expert network reasoning task to be migrated based on the time prediction results of multiple computing nodes may also include: judging whether there are at least two computing nodes, where the difference in time prediction results of different computing nodes is greater than a first preset difference; if so, determining the expert network with the smallest activation state parameter on the computing node with the larger time prediction result among the two computing nodes as the expert network reasoning task to be migrated, and determining the computing node with the smaller time prediction result as the target computing node of the expert network reasoning task to be migrated.

[0054] For S105, after determining the expert network reasoning task to be migrated, the expert network reasoning task to be migrated is migrated from the computing node where it is located to the computing node with a smaller time consumption prediction result, so as to achieve the convergence of the time consumption prediction results of each computing node in the heterogeneous computing system executing the expert network reasoning task.

[0055] In an embodiment of the present invention, two computing nodes can be determined at one time, and the expert network reasoning task on the computing node with the larger time-consuming prediction result can be migrated; or more than two computing nodes can be determined at one time, and the expert network reasoning task on the computing node with the larger time-consuming prediction result can be migrated.

[0056] The control method of the heterogeneous computing system provided by the embodiment of the present invention monitors the activation state parameters of the expert network of the hybrid expert model deployed on the computing nodes of the heterogeneous computing system, and predicts the time consumption of the computing nodes to execute reasoning tasks based on the activation state parameters of the expert network, thereby achieving accurate prediction of the time consumption of the computing nodes in the heterogeneous computing system to execute expert network calculations; according to the time consumption prediction results, the migration of expert network reasoning tasks is performed to migrate the expert network from the node where it is located to the computing node with a smaller time consumption prediction result. A hybrid expert model deployment scheme can be implemented to enable the heterogeneous computing system to execute more reasoning computing tasks, improve task execution efficiency and resource utilization, and can also respond to computing bottlenecks in advance according to the distribution changes of reasoning computing tasks during the operation of the heterogeneous computing system, and dynamically migrate expert network reasoning tasks to ensure efficient execution of reasoning computing based on the hybrid expert model.

[0057] Based on the above embodiment, the embodiment of the present invention continues to describe the step of predicting the time consumed by the computing node to execute the expert network reasoning task according to the activation state parameters of the expert network.

[0058] In an embodiment of the present invention, S103 calculates the time consumption prediction result of the computing node executing the expert network reasoning task based on the activation state parameters of the expert network on the computing node and the resource state information of the computing node, which may include: calculating the computing time prediction result of the computing node based on the activation state parameters of the expert network on the computing node and the computing power resource parameters of the computing node; calculating the communication time prediction result of the computing node based on the activation state parameters of the expert network on the computing node and the network resource parameters of the computing node; and calculating the time consumption prediction result of the computing node based on the computing time prediction result and the communication time prediction result.

[0059] Since the calculation of the expert network, the reception of the expert network input data, and the output of the expert network output data cannot be performed synchronously, the time consumed by the computing node to execute the expert network reasoning task can be determined based on the sum of the calculation time prediction result and the communication time prediction result.

[0060] In an embodiment of the present invention, the computing time prediction result of the computing node is calculated based on the activation state parameters of the expert network on the computing node and the computing power resource parameters of the computing node, which may include: calculating the total computing amount of the expert network when the computing node is activated based on the activation state parameters of the expert network on the computing node and the computing amount of the expert network; and calculating the computing time prediction result of the computing node based on the total computing amount and the computing power resource parameters provided by the computing node to the expert network.

[0061] Specifically, if the number of expert network activations per unit time on a compute node is used as the expert network activation state parameter, then the total computational effort of the expert network activated on the compute node can be calculated based on the unit number of expert network activations and the expert network's computational effort. Furthermore, based on the total computational effort of the expert network and the computing power resource parameters provided by the compute node to the expert network, the computational time prediction result of the compute node can be calculated.

[0062] In an embodiment of the present invention, calculating the total computational load of the expert network activated on the computing node based on the activation state parameters of the expert network on the computing node and the computational load of the expert network may include: calculating the total number of floating-point operations of the expert network activated on the computing node based on the activation state parameters of the expert network on the computing node and the number of floating-point operations included in the expert network's execution of one inference computation of the expert network. Calculating the computational time prediction result of the computing node based on the total computational load and the computing power resource parameters provided by the computing node to the expert network may include: calculating a first ratio of the total number of floating-point operations to the number of floating-point operations per second provided by the computing node to the expert network to obtain the computational time prediction result of the computing node.

[0063] That is to say, the computational amount of the expert network is expressed by the total number of floating-point operations. The total computational amount obtained is the total number of floating-point operations that the expert network activated on the computing node needs to perform. Dividing it by the number of floating-point operations per second of the computing node, we can get the total computing time required for the expert network activated on the computing node to perform an expert network inference task.

[0064] When multiple expert networks are deployed on a computing node, the total computing amount of the expert network activated on the computing node is calculated based on the activation state parameters of the expert network on the computing node and the computing amount of the expert network. This may include: taking the activation state parameters of the expert network as weights to perform weighted summation on the computing amount of the expert network on the computing node to obtain the total computing amount.

[0065] In an embodiment of the present invention, the communication time prediction result of the computing node is calculated based on the activation state parameters of the expert network on the computing node and the network resource parameters of the computing node, which may include: calculating the communication time prediction result of the computing node based on the activation state parameters of the expert network on the computing node, the communication time prediction result of the activated expert network, and the expected number of communications between the computing node and other computing nodes when the expert network is activated.

[0066] The step of determining the communication time prediction result of the activated expert network may include: calculating the communication time prediction result of the activated expert network based on the input and output data volume of an inference calculation of the activated expert network and the network resource parameters of the computing node.

[0067] Specifically, if the number of activations of the expert network on the computing node per unit time is used as the activation state parameter of the expert network, then the communication time prediction result of the activated expert network can be calculated based on the unit activation number of the expert network on the computing node and the input and output data volume of the expert network performing an expert network inference task.

[0068] Furthermore, based on the amount of input and output data of an activated expert network's inference calculation and the network resource parameters of the computing node, calculating the communication time prediction result of the activated expert network can include: for an activated expert network on a computing node, the sum of the input data and the output data transmitted between the activated expert network and the activated expert networks of other computing nodes is the input and output data amount; based on the input and output data amount and the network resource parameters of the computing node, calculating the communication time prediction result of the activated expert network.

[0069] That is to say, for each expert network on the computing node, its calculation process can be approximated as requiring two communication processes: the first is to receive the expert input data once from each activated expert network except itself, and the second is to receive the expert output data once from each activated expert network except itself.

[0070] In an embodiment of the present invention, calculating the communication time prediction result of the activated expert network based on the input and output data volume and the network resource parameters of the computing node can include: calculating a second ratio of the input and output data volume to the network resource parameters of the computing node; calculating the sum of the second ratio, the communication link delay of the computing node performing a data input, and the communication link delay of the computing node performing a data output, to obtain the communication time prediction result of the expert network.

[0071] For two activated expert networks, the communication time can be obtained by dividing the input and output data volume by the external communication bandwidth of the computing node to obtain the transmission time of the input and output data, and adding the communication link delay of the computing node twice to the input and output to obtain the communication time prediction result of the expert network.

[0072] In an embodiment of the present invention, the step of determining the expected number of communications between a computing node and other computing nodes when the expert network is activated may include: calculating the expected number of communications between the computing node and other computing nodes when the expert network is activated based on the first number of times the computing node calls the expert network per unit time and the second number of times the expert network is called in the heterogeneous computing system in executing an inference calculation of a hybrid expert model.

[0073] Then, based on the first number of times the computing node calls the expert network per unit time and the second number of times the expert network is called in the heterogeneous computing system in executing an inference calculation of a hybrid expert model, the expected number of communications between the computing node and other computing nodes when the expert network is activated is calculated, which can include: calculating a first difference between the second number and the first number; calculating a second difference between the number of expert networks activated in the heterogeneous computing system in executing an inference calculation of the hybrid expert model minus one; calculating the ratio of the product of the first difference and the second difference to the second number, to obtain the expected number of communications between the computing node and other computing nodes when the expert network is activated.

[0074] As for the expected number of communications, the expected number (estimated value) of communications between an expert network on a computing node and expert networks on other computing nodes is obtained by subtracting the activation frequency of the expert network on the computing node from the total activation frequency of all expert networks in the hybrid expert model and dividing it by the total frequency.

[0075] When multiple expert networks are deployed on a computing node, the communication time prediction result of the computing node is calculated based on the activation state parameters of the expert network on the computing node, the communication time prediction result of the activated expert network, and the expected number of communications between the computing node and other computing nodes when the expert network is activated. This can include: for the expert network on the computing node, calculating the communication time prediction result of the expert network based on the activation state parameters of the expert network, the communication time prediction result of the activated expert network, and the expected number of communications between the computing node and other computing nodes when the expert network is activated; and summing the communication time prediction results of each expert network on the computing node to obtain the communication time prediction result of the computing node.

[0076] Based on the above embodiments, the embodiments of the present invention further introduce the overall process of the control method of the heterogeneous computing system.

[0077] Figure 4 This is a flowchart of a second method for controlling a heterogeneous computing system provided by an embodiment of the present invention.

[0078] like Figure 4 As shown, the control method of the heterogeneous computing system provided by the embodiment of the present invention can be implemented based on five modules, namely: a task information collection module, a system information collection module, a calculation module, a task migration module, and a system status monitoring module.

[0079] Among them, the task information collection module is used to parse or calculate the necessary input information based on the information of the hybrid expert distributed reasoning task input by the user.

[0080] The task information collection module is used to collect the computational complexity of the forward propagation of the expert network in the hybrid expert model: this can be expressed in floating point operations (FLOPs). The computational complexity of the forward propagation can be estimated based on the neural layer composition of the neural network, and the computational complexity is equal to the sum of the computational complexity of all neural layers. The computational complexity of a neural layer can be estimated using mathematical methods. For example, for the forward computation of a fully connected layer, assuming the input data dimensions are (N, D), the hidden layer weight dimensions are (D, out), and the output is (N, out), the computational complexity is FLOPs = N*(2×D-1)*out. Different neural layers have different calculation methods, which are not listed here. Generally, the network shape of the expert networks is the same, so the computational complexity of only one expert network needs to be calculated.

[0081] The task information collection module is used to collect the input and output data volume of the expert network: the input data volume can be estimated by multiplying the batch size by the tensor size of the expert network input. For example, if the batch size is 10 and the input size of the neural network layer is [10,10], then the input data volume is 10*10*10*data precision (such as 32-bit floating point number fp32 is 4 bytes). The above information can also be counted using existing open source tools such as torchstat. In an embodiment of the present invention, the batch size of the inference computing task is set to 1. The output data volume can also be estimated by multiplying the output tensor size of the expert network by the batch size. The above information can all be counted using existing open source tools such as torchstat.

[0082] The task information collection module is used to collect the total number of expert networks in the hybrid expert model: this information can be directly obtained through the model structure.

[0083] The task information collection module is used to collect the number of activated expert neural networks in an inference calculation task. This parameter can be obtained by analyzing the output of the gating network.

[0084] The task information collection module is used to collect the computing nodes that participate in the inference computing task specified by the user.

[0085] The system information collection module is used to collect resource status information for compute nodes in a heterogeneous computing system. This information can include: Each compute node's computing power parameters. If the compute node is used only for inference computing tasks within a hybrid expert model, the node's peak computing power parameters (in floating-point operations per second) can be used. This parameter can be obtained from the compute node's product manual or actual testing. Each compute node's network resource parameters, specifically the bandwidth and latency of its external links, can be obtained using standard benchmark tools such as iperf.

[0086] The system status monitoring module is used to monitor the activation status parameters of the expert network on each computing node, specifically the number of expert network activations per unit time. For example, a time threshold, such as 30 minutes, can be set, and the number of times each expert network is activated per second can be calculated.

[0087] Based on the information collected by the task information collection module, the system information collection module and the system status monitoring module, the calculation module performs time consumption prediction of each computing node in an inference calculation task according to the deployment method of the hybrid expert model in the heterogeneous computing system, and outputs the information of the expert network inference task to be migrated.

[0088] The task migration module performs task migration based on the information of the expert network reasoning task to be migrated output by the calculation module.

[0089] The calculation steps of the calculation module are further introduced below.

[0090] Based on the information collected by the task information collection module, system information collection module and system status monitoring module, the input parameters of the calculation module may include the total number of computing devices participating in the inference calculation task. , the total number of expert networks , compute nodes The number of expert networks assigned to , compute nodes The bandwidth of the external communication link , compute nodes The delay of the external communication link , the computational cost of each expert network (unit: floating point operations FLOPs), computing nodes Computing resource parameters (unit: floating point operations per second FLOPS), the amount of input data for each expert network , the output data volume of each expert network , expert network Number of activations per unit time (i.e., the average number of times it is activated per second), the number of expert networks activated in a heterogeneous computing system performing an inference computing task .

[0091] Among them, for expert networks Number of activations per unit time If it is the first calculation, the calculation module can set the value to the same default value, such as 1, when it does not receive information from the system status monitoring module.

[0092] When deploying a hybrid expert model in a heterogeneous computing system, you can evenly distribute expert networks across all compute nodes. For example, if the hybrid expert model includes 20 expert networks and there are three compute nodes participating in the inference task, you can allocate 7, 7, and 6 expert networks to each compute node, respectively. At this point, each compute node is assigned a different expert network, and this initial deployment is recorded.

[0093] In actual calculation, for computing nodes , the time-consuming prediction result of the expert network activated in unit time (seconds) can be calculated : ; in, Represents a compute node The predicted result of the computation time of the expert network activated on it in processing unit time. The summation formula represents the computation node All expert networks on . This formula first calculates the nodes The total amount of computation of all expert networks in a unit time is summed up, and then the computation nodes are used to calculate the Computing resource parameters Estimate the overall computation time.

[0094] Represents a compute node The communication time prediction result of the expert network activated on it for reasoning per unit time is processed. Regarding the communication time, in the distributed reasoning process, the operation process of each expert network can be approximated as two communication processes as follows: the first is to receive the expert input data from each activated expert network except itself, and the second is to receive the expert output data from each activated expert network except itself. Represents the time taken for two communications between the two expert networks during the inference process. Representative expert network When activated once, the expected (estimated) number of times the expert network communicates with the expert network on other computing nodes, where the denominator represents the total frequency of the entire system calling the expert network per unit time, and the numerator is the total frequency of the system minus the expert network Compute node The total frequency of calling the expert network per unit time.

[0095] For all computing nodes, the above time consumption prediction results are calculated .

[0096] The time consumption prediction results of each computing node obtained from the calculation Take the largest and the smallest , the corresponding computing node is recorded as computing node and compute nodes , from the computing node Take out the number of activations per unit time Minimum expert network and migrate its expert network reasoning tasks to computing nodes , while updating the compute nodes The number of expert networks 、 , update the compute nodes The number of expert networks 、 .

[0097] Record 、 If the selected 、 With the last selected 、 Same (the same computing node is repeatedly selected and the expert network is repeatedly deployed to the other party. This shows that the time consumption prediction results of all computing nodes in the current system are Has reached as close as possible, that is, the time consumed by each computing node to process the activated expert network per unit time has reached as close as possible, the load of the heterogeneous computing system has been relatively balanced, and the number of inference requests that the heterogeneous computing system can process per unit time has reached as large as possible), then the output result of the expert network corresponding to each computing node is output. Otherwise, after the above migration work is performed, the output of each computing node is calculated under the new deployment mode. and determine the maximum and the smallest , re-enter the judgment of this step.

[0098] Therefore, the control method of the heterogeneous computing system provided by the embodiment of the present invention can dynamically update the deployment method of the hybrid expert model in the heterogeneous computing system composed of multiple heterogeneous computing systems with different performance and computing power from different manufacturers, taking into account the influence of multiple variables such as computing power and network, so as to ensure that the heterogeneous computing system can process as many inference requests as possible, thereby improving the efficiency of inference computing tasks and improving the utilization of equipment resources.

[0099] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0100] An embodiment of the present invention also provides a control device for a heterogeneous computing system, which may include: a task information collection module for determining the expert network of the hybrid expert model deployed by the computing nodes of the heterogeneous computing system; a system information collection module for obtaining the resource status information of the computing nodes; a system status monitoring module for monitoring the activation status parameters of the expert network; a calculation module for calculating the time consumption prediction result of the computing node executing the expert network reasoning task based on the activation status parameters of the expert network on the computing node and the resource status information of the computing node; determining the expert network reasoning task to be migrated based on the time consumption prediction results of multiple computing nodes; and a task migration module for migrating the expert network reasoning task to be migrated from the computing node where it is located to a computing node with a smaller time consumption prediction result.

[0101] For the description of the features in the embodiments corresponding to the control device of the heterogeneous computing system, reference may be made to the relevant description of the embodiments corresponding to the control method of the heterogeneous computing system, which will not be repeated here.

[0102] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the control method for a heterogeneous computing system.

[0103] An embodiment of the present invention further provides a non-volatile storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned embodiments of the control method for a heterogeneous computing system when running.

[0104] In an exemplary embodiment, the non-volatile storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0105] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned embodiments of the control method for a heterogeneous computing system are implemented.

[0106] An embodiment of the present invention further provides another computer program product, including a non-volatile storage medium, wherein the non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned control method embodiments for heterogeneous computing systems are implemented.

[0107] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0108] The above is a detailed introduction to the control method, device, equipment, medium and product of a heterogeneous computing system provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A control method for a heterogeneous computing system, characterized in that: include: Determine an expert network of hybrid expert models deployed on computing nodes of a heterogeneous computing system; monitoring activation state parameters of the expert network; Calculating a time-consuming prediction result of the computing node executing the expert network reasoning task based on the activation state parameters of the expert network on the computing node and the resource state information of the computing node; Determining the expert network reasoning task to be migrated based on the time consumption prediction results of the plurality of computing nodes; Migrate the expert network reasoning task to be migrated from the computing node where it is located to the computing node with a smaller prediction result in terms of time consumption.

2. The control method of a heterogeneous computing system according to claim 1, characterized in that: Monitoring activation state parameters of the expert network includes: The number of activations of the expert network on the computing node within a unit time is monitored as an activation state parameter of the expert network.

3. The control method of a heterogeneous computing system according to claim 2, wherein: The method further comprises calculating, based on the activation state parameters of the expert network on the computing node and the resource state information of the computing node, a time consumption prediction result of the computing node executing the expert network reasoning task, including: According to the number of activations of the expert network on the computing node within a unit time and the resource status information of the computing node, a time-consuming prediction result of the expert network reasoning task of the computing node activated within a unit time is calculated.

4. The control method of a heterogeneous computing system according to claim 1, wherein: The method further comprises calculating, based on the activation state parameters of the expert network on the computing node and the resource state information of the computing node, a time consumption prediction result of the computing node executing the expert network reasoning task, including: Calculating a prediction result of the computing time of the computing node according to the activation state parameter of the expert network on the computing node and the computing power resource parameter of the computing node; Calculating a communication time prediction result of the computing node according to the activation state parameters of the expert network on the computing node and the network resource parameters of the computing node; The time consumption prediction result of the computing node is calculated based on the computing time consumption prediction result and the communication time consumption prediction result.

5. The control method of a heterogeneous computing system according to claim 4, characterized in that: The method further comprises calculating, based on the activation state parameters of the expert network on the computing node and the computing power resource parameters of the computing node, a computing time prediction result of the computing node, including: Calculating the total computational load of the expert network when the computing node is activated according to the activation state parameter of the expert network on the computing node and the computational load of the expert network; The calculation time prediction result of the calculation node is calculated based on the total calculation amount and the calculation resource parameters provided by the calculation node to the expert network.

6. The control method of a heterogeneous computing system according to claim 5, characterized in that: Calculating the total computational load of the expert network when the computing node is activated according to the activation state parameter of the expert network on the computing node and the computational load of the expert network includes: Calculate the total floating-point operation number of the expert network when the computing node is activated according to the activation state parameter of the expert network on the computing node and the number of floating-point operations included in the expert network performing one inference calculation of the expert network; The calculation time prediction result of the computing node is calculated based on the total computing amount and the computing power resource parameters provided by the computing node to the expert network, including: A first ratio of the total number of floating-point operations to the number of floating-point operations per second provided by the computing node to the expert network is calculated to obtain a computing time prediction result of the computing node.

7. The control method of a heterogeneous computing system according to claim 5, characterized in that: Calculating the total computational load of the expert network when the computing node is activated according to the activation state parameter of the expert network on the computing node and the computational load of the expert network includes: The activation state parameter of the expert network is used as a weight to perform weighted summation on the computation amount of the expert network on the computation nodes to obtain the total computation amount.

8. The control method of a heterogeneous computing system according to claim 4, characterized in that: The method further comprises calculating, based on the activation state parameters of the expert network on the computing node and the network resource parameters of the computing node, a communication time prediction result of the computing node, including: The communication time prediction result of the computing node is calculated based on the activation state parameters of the expert network on the computing node, the communication time prediction result of the activated expert network, and the expected number of communications between the computing node and other computing nodes when the expert network is activated.

9. The control method of a heterogeneous computing system according to claim 8, characterized in that: The step of determining the communication time prediction result of the activated expert network includes: The communication time prediction result of the activated expert network is calculated according to the input and output data volume of the inference calculation performed once by the activated expert network and the network resource parameters of the computing node.

10. The control method of a heterogeneous computing system according to claim 9, characterized in that: Calculating a communication time prediction result of the activated expert network according to the input and output data volume of the inference calculation of the activated expert network and the network resource parameters of the computing node, including: For one of the activated expert networks on the computing nodes, the sum of one input data and one output data transmitted between the activated expert network and the activated expert networks of other computing nodes is the input and output data amount; The communication time prediction result of the activated expert network is calculated according to the input and output data volume and the network resource parameters of the computing node.

11. The control method of a heterogeneous computing system according to claim 10, characterized in that: Calculating a communication time prediction result of the activated expert network according to the input and output data volume and the network resource parameters of the computing node, including: Calculating a second ratio of the input and output data volume to a network resource parameter of the computing node; The sum of the second ratio, the communication link delay of the computing node executing one data input, and the communication link delay of the computing node executing one data output is calculated to obtain a communication time prediction result of the expert network.

12. The control method of a heterogeneous computing system according to claim 8, characterized in that: The step of determining the expected number of communications between the computing node and other computing nodes when the expert network is activated includes: Based on the first number of times the computing node calls the expert network per unit time and the second number of times the expert network is called in the heterogeneous computing system during an execution of an inference calculation of the hybrid expert model, the expected number of communications between the computing node and the other computing nodes when the expert network is activated is calculated.

13. The control method of a heterogeneous computing system according to claim 12, characterized in that: Calculating an expected number of communications between the computing node and other computing nodes when the expert network is activated based on a first number of times the computing node calls the expert network per unit time and a second number of times the expert network is called in an inference calculation of the hybrid expert model in the heterogeneous computing system, including: Calculate a first difference between the second number of times and the first number of times; Calculating a second difference value of the number of expert networks activated in executing a reasoning calculation of the hybrid expert model in the heterogeneous computing system minus one; The ratio of the product of the first difference and the second difference to the second number of times is calculated to obtain the expected number of communications between the computing node and other computing nodes when the expert network is activated.

14. The control method of a heterogeneous computing system according to claim 7, wherein: The communication time prediction result of the computing node is calculated based on the activation state parameters of the expert network on the computing node, the communication time prediction result of the activated expert network, and the expected number of communications between the computing node and other computing nodes when the expert network is activated, including: For the expert network on the computing node, calculating a communication time prediction result of the expert network based on the activation state parameter of the expert network, the communication time prediction result of the activated expert network, and the expected number of communications between the computing node and other computing nodes when the expert network is activated; The communication time prediction results of each of the expert networks on the computing node are summed to obtain the communication time prediction result of the computing node.

15. The control method of a heterogeneous computing system according to claim 1, wherein: Determining the expert network reasoning task to be migrated based on the time consumption prediction results of the plurality of computing nodes includes: Determine the computing node with the largest predicted time consumption result and the computing node with the smallest predicted time consumption result in the heterogeneous computing system; Determining whether the determined computing node is the same as that at the previous moment; If not, the expert network with the smallest time activation state parameter of the computing node with the largest time consumption prediction result is determined as the expert network reasoning task to be migrated, and the computing node with the smallest time consumption prediction result is determined as the target computing node of the expert network reasoning task to be migrated.

16. The control method of a heterogeneous computing system according to claim 1, wherein: Determining the expert network reasoning task to be migrated based on the time consumption prediction results of the plurality of computing nodes includes: determining a sum of activation state parameters of the expert network on the computing node; If it is determined based on the sum of the activation state parameters of each computing node that there is a computing node that meets the threshold, then the computing node with the largest time consumption prediction result and the expert network with the smallest time activation state parameter is determined as the expert network reasoning task to be migrated, and the computing node with the smallest time consumption prediction result is determined as the target computing node of the expert network reasoning task to be migrated.

17. A control device for a heterogeneous computing system, characterized in that: include: A task information collection module, used to determine the expert network of the hybrid expert model deployed by the computing nodes of the heterogeneous computing system; A system information collection module, configured to obtain resource status information of the computing node; A system status monitoring module, used for monitoring activation status parameters of the expert network; A calculation module is configured to calculate a time consumption prediction result of the computing node executing the expert network reasoning task based on the activation state parameters of the expert network on the computing node and the resource state information of the computing node; and determine the expert network reasoning task to be migrated based on the time consumption prediction results of the plurality of computing nodes; The task migration module is used to migrate the expert network reasoning task to be migrated from the computing node where it is located to the computing node with a smaller prediction result.

18. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for controlling a heterogeneous computing system according to any one of claims 1 to 16 when executing the computer program.

19. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the control method of the heterogeneous computing system according to any one of claims 1 to 16.

20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the control method of the heterogeneous computing system according to any one of claims 1 to 16 are implemented.

Citation Information

Cited By

  • Heterogeneous computing system, fault processing method and device, equipment, medium and product

    CN121233404A