Hybrid expert model deployment method and device based on heterogeneous hardware equipment cluster, equipment and medium
By abstracting and updating data in real time for heterogeneous hardware device clusters and generating hardware-aware routing weights, the storage and communication problems of hybrid expert models deployed on heterogeneous devices are solved, achieving efficient resource utilization and inference efficiency.
Patent Information
- Application Number
- CN202510903475.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies are unable to efficiently coordinate heterogeneous hardware device clusters, resulting in exponential growth in storage overhead, increased communication latency, and load imbalance when hybrid expert models are deployed on different hardware devices.
By performing data abstraction processing on the target heterogeneous hardware device cluster, generating an initial device description vector, and compiling the hybrid expert model file into a binary intermediate representation, hardware probes are generated based on the hardware type for real-time updates, hardware-aware routing weights are determined, and appropriate expert sub-models are split and deployed on each device.
It achieves efficient deployment of hybrid expert models on heterogeneous hardware device clusters, optimizes resource utilization, reduces service costs and improves inference efficiency.
Smart Images

Figure CN120704714A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a hybrid expert model deployment method, apparatus, equipment and medium based on a heterogeneous hardware device cluster. Background Art
[0002] With the widespread application of large models with hundreds of billions of parameters, hybrid expert models, due to their dynamic sparse activation characteristics, have become a key technology for reducing inference costs. However, industrial-level deployments face many technical bottlenecks. For example, existing deployment solutions are mainly designed for homogeneous GPU (Graphics Processing Unit) clusters and cannot efficiently coordinate multiple heterogeneous device clusters. The barriers between different hardware instruction sets require expert models to be compiled for different hardware. Static compilation solutions will exponentially increase the storage overhead of expert replicas and cannot dynamically adapt to changes in device topology. There is generally no fast communication channel between different devices. When activated experts are distributed across devices with different architectures, data needs to be serialized and deserialized multiple times, increasing communication latency. Traditional hybrid expert model routing algorithms select experts based solely on input semantics, ignoring the real-time status of the hardware, which can easily exacerbate load imbalance.
[0003] Therefore, how to achieve heterogeneous deployment of hybrid expert models needs to be solved. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, apparatus, device, and medium for deploying a hybrid expert model based on a heterogeneous hardware device cluster, which can realize the heterogeneous deployment of the hybrid expert model. The specific solution is as follows:
[0005] In a first aspect, the present application discloses a hybrid expert model deployment method based on a heterogeneous hardware device cluster, comprising:
[0006] Performing data abstraction processing on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector, and compiling the target expert model file in the hybrid expert model into a binary intermediate representation; the target expert model file is an initial expert file deployed on each device and not activated;
[0007] generating a hardware probe for the target device based on the hardware type of the target device in the target heterogeneous hardware device cluster and the binary intermediate representation, and updating the initial device description vector in real time based on the hardware probe to obtain a real-time device description vector;
[0008] Determining a hardware-aware routing weight of each device in the target heterogeneous hardware device cluster based on the real-time device description vector, and determining a to-be-activated expert model file corresponding to the target heterogeneous hardware device cluster from the target expert model file based on the hardware-aware routing weight;
[0009] The expert model file to be activated is split into corresponding expert sub-models based on operator types, and the expert sub-models are distributed to the target heterogeneous hardware device cluster for deployment and installation.
[0010] Optionally, performing data abstraction processing on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector includes:
[0011] The device-related information of each device in the target heterogeneous hardware device cluster is obtained, and the device-related information is converted into a unified device descriptor through data abstraction processing to obtain an initial device description vector.
[0012] Optionally, compiling the target expert model file in the hybrid expert model into a binary intermediate representation includes:
[0013] Performing operator normalization processing on the target expert model file in the hybrid expert model to decompose the target expert model file into a basic operator set; the basic operator set is a computing node element in the target expert model file that is unrelated to the hardware device;
[0014] Constructing a model calculation graph corresponding to the target expert model file based on each calculation node element in the basic operator set, and inserting corresponding shape placeholders into the model calculation graph to obtain corresponding memory descriptors;
[0015] Corresponding cross-device sharing attributes are defined in the memory descriptor to obtain a binary intermediate representation corresponding to the target expert model file.
[0016] Optionally, generating a hardware probe for the target device based on the hardware type of the target device in the target heterogeneous hardware device cluster and the binary intermediate representation, and updating the initial device description vector in real time based on the hardware probe to obtain a real-time device description vector includes:
[0017] Detecting the hardware type of the target device in the target heterogeneous hardware device cluster to obtain an instruction set architecture type, and determining a target optimization strategy from a device optimization strategy mapping table in the binary intermediate representation;
[0018] Compiling the binary intermediate representation into a device native instruction set based on the target optimization strategy and the instruction set architecture type;
[0019] Injecting probe code into the native instruction set of the device to obtain a hardware probe for the target device, and monitoring the real-time status of the target device based on the hardware probe to obtain the real-time status of the device;
[0020] The initial device description vector is updated in real time based on the real-time status of the device to obtain a real-time device description vector.
[0021] Optionally, determining the hardware-aware routing weight of each device in the target heterogeneous hardware device cluster based on the real-time device description vector, and determining the to-be-activated expert model file corresponding to the target heterogeneous hardware device cluster from the target expert model file based on the hardware-aware routing weight, includes:
[0022] Substituting the real-time device description vector of each device in the target heterogeneous hardware device cluster into a preset routing weight determination function to obtain the hardware-aware routing weight corresponding to each device;
[0023] Top-K experts are determined from the target expert model file based on the hardware-aware routing weights to obtain an expert model file to be activated corresponding to the target heterogeneous hardware device cluster.
[0024] Optionally, splitting the to-be-activated expert model file into corresponding expert sub-models based on operator type, and allocating the expert sub-models to the target heterogeneous hardware device cluster for deployment and installation, includes:
[0025] Splitting the expert model file to be activated into corresponding expert sub-models based on operator type;
[0026] The to-be-deployed device corresponding to the expert sub-model is determined based on the operator characteristics used by each device in the target heterogeneous hardware device cluster, and the expert sub-model is allocated to the to-be-deployed device for deployment and installation; data is transmitted between the to-be-deployed devices via a shared virtual address.
[0027] In a second aspect, the present application discloses a hybrid expert model deployment device based on a heterogeneous hardware device cluster, comprising:
[0028] A device data abstraction module is used to perform data abstraction processing on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector and compile the target expert model file in the hybrid expert model into a binary intermediate representation; the target expert model file is the initial expert file deployed on each device and not activated;
[0029] a device data updating module, configured to generate a hardware probe for the target device in the target heterogeneous hardware device cluster based on the hardware type of the target device and the binary intermediate representation, and to update the initial device description vector in real time based on the hardware probe to obtain a real-time device description vector;
[0030] an expert model determination module, configured to determine a hardware-aware routing weight of each device in the target heterogeneous hardware device cluster based on the real-time device description vector, and determine a to-be-activated expert model file corresponding to the target heterogeneous hardware device cluster from the target expert model file based on the hardware-aware routing weight;
[0031] The model deployment module is used to split the expert model file to be activated into corresponding expert sub-models based on operator types, and distribute the expert sub-models to the target heterogeneous hardware device cluster for deployment and installation.
[0032] Optionally, the device data abstraction module includes:
[0033] An operator decomposition unit is used to perform operator normalization processing on a target expert model file in a hybrid expert model to decompose the target expert model file into a basic operator set; the basic operator set is a computing node element in the target expert model file that is not related to a hardware device;
[0034] A model calculation graph determining unit, configured to construct a model calculation graph corresponding to the target expert model file based on each calculation node element in the basic operator set, and insert corresponding shape placeholders into the model calculation graph to obtain corresponding memory descriptors;
[0035] The binary intermediate representation determination module is used to define corresponding cross-device sharing attributes in the memory descriptor to obtain the binary intermediate representation corresponding to the target expert model file.
[0036] In a third aspect, the present application discloses an electronic device, comprising:
[0037] Memory, used to store computer programs;
[0038] A processor is used to execute the computer program to implement the aforementioned hybrid expert model deployment method based on a heterogeneous hardware device cluster.
[0039] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the aforementioned hybrid expert model deployment method based on a heterogeneous hardware device cluster.
[0040] It can be seen that in the present application, data abstraction processing is performed on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector, and the target expert model file in the hybrid expert model is compiled into a binary intermediate representation; the target expert model file is an initial expert file that is deployed on each device and has not been activated; a hardware probe for the target device is generated based on the hardware type of the target device in the target heterogeneous hardware device cluster and the binary intermediate representation, and the initial device description vector is updated in real time based on the hardware probe to obtain a real-time device description vector; the hardware-aware routing weight of each device in the target heterogeneous hardware device cluster is determined based on the real-time device description vector, and the expert model file to be activated corresponding to the target heterogeneous hardware device cluster is determined from the target expert model file based on the perceived routing weight; the expert model file to be activated is split into corresponding expert sub-models based on the operator type, and the expert sub-model is allocated to the target heterogeneous hardware device cluster for deployment and installation. Specifically, the target expert files are compiled into a unified binary intermediate representation, which then generates hardware probes for each device. The final expert model file to be activated is determined based on the hardware-aware routing weights corresponding to each device. Based on the operator type corresponding to each device, the appropriate expert sub-model is selected from the expert model file to be activated and deployed, completing the deployment and installation of the hybrid expert model on the target heterogeneous hardware device cluster. This ensures efficient inference on hybrid devices through a unified cross-architecture orchestration system and adaptive expert-hardware matching technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0042] Figure 1 This is a flow chart of a hybrid expert model deployment method based on a heterogeneous hardware device cluster disclosed in this application;
[0043] Figure 2 This is a flow chart of a specific hybrid expert model deployment method based on heterogeneous hardware device clusters disclosed in this application;
[0044] Figure 3 This is a schematic diagram of the structure of a hybrid expert model deployment device based on a heterogeneous hardware device cluster disclosed in this application;
[0045] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0047] Deploying hybrid expert models on heterogeneous hardware clusters can fully leverage the unique advantages of different hardware, such as CPUs (Central Processing Units), GPUs, various accelerators like NPUs (Neural Processing Units) and TPUs (Tensor Processing Units), edge devices, and cloud servers, addressing requirements that are difficult to meet with a single hardware solution. Hybrid expert models can be applied in e-commerce, content platforms, and advertising platforms that need to process massive amounts of user and item characteristics to deliver personalized recommendations based on user needs. Alternatively, autonomous vehicles must process multi-sensor data (cameras, LiDAR, millimeter-wave radar) for real-time perception, prediction, and decision-making, collaborating with roadside units and the cloud. Another application is factories where numerous sensors monitor equipment status and require rapid local response to anomalies while integrating global data for in-depth analysis and prediction.
[0048] Traditional hybrid expert model routing algorithms select experts based solely on input semantics, ignoring the real-time state of hardware, which can easily exacerbate load imbalance. This application will specifically introduce a hybrid expert model deployment method based on a heterogeneous hardware cluster. This method ensures efficient inference on a mix of devices such as CPUs, GPUs, TPUs, and NPUs through a unified cross-architecture orchestration system and adaptive expert-hardware matching technology.
[0049] See also Figure 1 As shown, the embodiment of the present application discloses a hybrid expert model deployment method based on a heterogeneous hardware device cluster, including:
[0050] Step S11: Perform data abstraction processing on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector, and compile the target expert model file in the hybrid expert model into a binary intermediate representation; the target expert model file is an initial expert file deployed on the devices and not activated.
[0051] In this embodiment, the data abstraction processing is performed on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector, including: obtaining device-related information of each device in the target heterogeneous hardware device cluster, and converting the device-related information into a unified device descriptor through data abstraction processing to obtain an initial device description vector. Specifically, each device in the heterogeneous hardware device cluster is abstracted into a unified device descriptor UDID (Unique Device Identifier), and the UDID contains a computing power score, a memory bandwidth value, and a delay sensitivity coefficient. In the specific operation, the computing power, memory bandwidth, delay sensitivity, energy efficiency ratio, device architecture type, memory capacity, and device connection information of the device are obtained. Taking a device as an example, the unified device descriptor UDID looks like the following:
[0052] {
[0053] "Calculation ability score": 0.86,
[0054] "Memory Bandwidth": 408,
[0055] "Delay Sensitivity": 1.1,
[0056] "Energy Efficiency Ratio": 42.3,
[0057] "Schema Type": "****",
[0058] "Memory Capacity": 32768,
[0059] "Connect device": [
[0060] {"target_device_id": 0, "bandwidth": 64, "latency": 2.8} ]
[0062] }
[0063] The computing power score is calculated by multiplying the weighted sum of the computing power corresponding to the supported data accuracy by the ratio of the actual frequency to the nominal frequency.
[0064] In this embodiment, compiling the target expert model file in the hybrid expert model into a binary intermediate representation includes: performing operator normalization on the target expert model file in the hybrid expert model to decompose the target expert model file into a basic operator set; the basic operator set is hardware-independent computational node elements in the target expert model file; constructing a model computation graph corresponding to the target expert model file based on each computational node element in the basic operator set, inserting corresponding shape placeholders into the model computation graph to obtain a corresponding memory descriptor; and defining corresponding cross-device sharing attributes in the memory descriptor to obtain a binary intermediate representation corresponding to the target expert model file. Specifically, the target expert model file in the hybrid expert model is received and compiled into a hardware-independent binary intermediate representation (BIR), wherein the BIR includes: a dynamic operator directed acyclic graph (DAG); explicitly declared memory alignment requirements and byte order specifications; and a device optimization strategy mapping table. Specifically, the original expert model undergoes operator normalization, decomposing device-specific operators into a basic operator set. Dynamically shaped placeholders are inserted into the computation graph, supporting runtime dimensionality reshaping. Cross-device sharing properties for tensors are defined in the memory descriptor, including rules for generating the Virtual Address Translation (VAT). A data consistency maintenance protocol is also implemented. It should be noted that the basic operator set is the element that constitutes the computation graph nodes. This hardware-independent operator set removes certain hardware-specific features, facilitating the subsequent generation of a unified BIR. The data consistency maintenance protocol is essential for ensuring cross-device data correctness and integrity. This protocol is designed to avoid read and write conflicts, ensure service continuity during device failures, and guarantee data correctness across devices. First, after decomposing the target expert model file into a basic operator set, the model computation graph is constructed. A memory descriptor is generated through computation graph compilation. Finally, the corresponding cross-device sharing properties are defined in the memory descriptor to obtain the binary intermediate representation corresponding to the target expert model file.
[0065] Step S12: generating a hardware probe for the target device based on the hardware type of the target device in the target heterogeneous hardware device cluster and the binary intermediate representation, and updating the initial device description vector in real time based on the hardware probe to obtain a real-time device description vector.
[0066] In this embodiment, the hardware probe for the target device is generated based on the hardware type of the target device in the target heterogeneous hardware device cluster and the binary intermediate representation, and the initial device description vector is updated in real time based on the hardware probe to obtain a real-time device description vector, including: detecting the hardware type of the target device in the target heterogeneous hardware device cluster to obtain an instruction set architecture type, and determining a target optimization strategy from a device optimization strategy mapping table in the binary intermediate representation; compiling the binary intermediate representation into a device native instruction set based on the target optimization strategy and the instruction set architecture type; injecting probe code into the device native instruction set to obtain a hardware probe for the target device, and monitoring the real-time status of the target device based on the hardware probe to obtain the real-time status of the device; and updating the initial device description vector in real time based on the real-time status of the device to obtain a real-time device description vector. First, at runtime, the binary intermediate representation is compiled into a device native instruction set based on the hardware type of the target device. Specifically, the instruction set architecture type of the target device is detected; the matching optimization strategy is selected from the device optimization strategy mapping table of the BIR; the hardware probe code is injected into the generated native instruction, and the probe code is used to collect computing unit utilization, memory occupancy, latency and instruction throughput to obtain a real-time device description vector. The real-time device description vector can also be fed back to the expert-device affinity matrix. It should be noted here that the expert-device affinity matrix reflects the matching degree between a specific expert and the target device, whether the expert is suitable for deployment on the target device, and is used to predict the operation of a specific expert on the target device. If a device failure occurs later on the device running a specific expert, the matrix can quickly generate an expert migration path to migrate the specific expert to the most matching device at that time.
[0067] Step S13: determining the hardware-aware routing weight of each device in the target heterogeneous hardware device cluster based on the real-time device description vector, and determining the expert model file to be activated corresponding to the target heterogeneous hardware device cluster from the target expert model file based on the hardware-aware routing weight.
[0068] In this embodiment, the hardware-aware routing weight of each device in the target heterogeneous hardware device cluster is determined based on the real-time device description vector, and the expert model file to be activated corresponding to the target heterogeneous hardware device cluster is determined from the target expert model file based on the perceived routing weight, including: substituting the real-time device description vector of each device in the target heterogeneous hardware device cluster into the preset routing weight determination function to obtain the hardware-aware routing weight corresponding to each device; determining the Top-K experts from the target expert model file based on the hardware-aware routing weight to obtain the expert model file to be activated corresponding to the target heterogeneous hardware device cluster. In other words, the initial expert probability distribution of the input data x is obtained through the gating network. That is, the user asks questions to the hybrid expert model and inputs the corresponding text. Then the initial expert probability distribution of the initial expert file that has not been activated is obtained through the gating network. Then, the real-time collected device status data is integrated to update the initial expert probability distribution to generate hardware-aware routing weights:
[0069] ;
[0070] in, Represents semantically driven raw routing decisions, Real-time abstract descriptor for devices, Score the health of the device. is a learnable weight parameter. Select Top-K experts to trigger expert preloading and cross-device pipeline execution.
[0071] Step S14: splitting the expert model file to be activated into corresponding expert sub-models based on operator types, and distributing the expert sub-models to the target heterogeneous hardware device cluster for deployment and installation.
[0072] In this embodiment, the method of splitting the expert model file to be activated into corresponding expert sub-models based on operator type and allocating the expert sub-models to the target heterogeneous hardware device cluster for deployment and installation includes: splitting the expert model file to be activated into corresponding expert sub-models based on operator type; determining the to-be-deployed device corresponding to the expert sub-model based on the operator characteristics used by each device in the target heterogeneous hardware device cluster, and allocating the expert sub-model to the to-be-deployed device for deployment and installation; and data transmission between the to-be-deployed devices via shared virtual addresses. That is, the computation graph of a single expert is split into sub-task groups based on operator type; sub-tasks are allocated to the heterogeneous device cluster based on operator characteristics: convolution operators are allocated to NPU devices; matrix multiplication operators are allocated to GPU devices; and control logic operators are allocated to CPU devices; and zero-copy data transmission between sub-tasks is achieved via a shared virtual address space. Specifically, for scenarios such as intelligent customer service, content creation, and multimodal search that require AI (artificial intelligence) services supporting multimodal input and output, such as text, images, audio, and video, experts in text processing, image recognition, speech recognition, and cross-modal alignment can be deployed on hardware optimized for their respective computing modes. For example, CPUs / general-purpose GPUs can be used for text, Tensor Core-rich GPUs or dedicated AI accelerator cards can be used for images and video, and devices with DSP (digital signal processing) can be used for voice. This significantly improves the throughput and response speed of multimodal services, optimizes overall hardware resource utilization, and reduces service costs.
[0073] It can be seen that in this embodiment, Figure 2As shown, data abstraction processing is performed on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector, and the target expert model file in the hybrid expert model is compiled into a binary intermediate representation; the target expert model file is an initial expert file deployed on each device and has not been activated; a hardware probe for the target device in the target heterogeneous hardware device cluster is generated based on the hardware type of the target device in the target heterogeneous hardware device cluster and the binary intermediate representation, and the initial device description vector is updated in real time based on the hardware probe to obtain a real-time device description vector; the hardware-aware routing weight of each device in the target heterogeneous hardware device cluster is determined based on the real-time device description vector, and the expert model file to be activated corresponding to the target heterogeneous hardware device cluster is determined from the target expert model file based on the perceived routing weight; the expert model file to be activated is split into corresponding expert sub-models based on the operator type, and the expert sub-models are allocated to the target heterogeneous hardware device cluster for deployment and installation. Specifically, the target expert files are compiled into a unified binary intermediate representation, which then generates hardware probes for each device. The final expert model file to be activated is determined based on the hardware-aware routing weights corresponding to each device. Based on the operator type corresponding to each device, the appropriate expert sub-model is selected from the expert model file to be activated and deployed, completing the deployment and installation of the hybrid expert model on the target heterogeneous hardware device cluster. This ensures efficient inference on hybrid devices through a unified cross-architecture orchestration system and adaptive expert-hardware matching technology.
[0074] refer to Figure 3 The embodiment of the present application further discloses a hybrid expert model deployment device based on a heterogeneous hardware device cluster, including:
[0075] The device data abstraction module 11 is used to perform data abstraction processing on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector and compile the target expert model file in the hybrid expert model into a binary intermediate representation; the target expert model file is the initial expert file deployed on each device and not activated;
[0076] a device data updating module 12, configured to generate a hardware probe for the target device in the target heterogeneous hardware device cluster based on the hardware type of the target device and the binary intermediate representation, and to update the initial device description vector in real time based on the hardware probe to obtain a real-time device description vector;
[0077] an expert model determination module 13, configured to determine a hardware-aware routing weight of each device in the target heterogeneous hardware device cluster based on the real-time device description vector, and determine a to-be-activated expert model file corresponding to the target heterogeneous hardware device cluster from the target expert model file based on the hardware-aware routing weight;
[0078] The model deployment module 14 is configured to split the expert model file to be activated into corresponding expert sub-models based on operator types, and distribute the expert sub-models to the target heterogeneous hardware device cluster for deployment and installation.
[0079] As can be seen, in this embodiment, the target expert file is compiled into a unified binary intermediate representation, which then generates hardware probes for each device. The final expert model file to be activated is determined based on the hardware-aware routing weight corresponding to each device. Based on the operator type corresponding to each device, the appropriate expert sub-model is determined from the expert model file to be activated and deployed, completing the deployment and installation of the hybrid expert model on the target heterogeneous hardware device cluster. In this way, through a unified cross-architecture orchestration system and expert-hardware adaptive matching technology, efficient inference is ensured on hybrid devices.
[0080] In some specific embodiments, the device data abstraction module 11 includes:
[0081] An operator decomposition unit is used to perform operator normalization processing on a target expert model file in a hybrid expert model to decompose the target expert model file into a basic operator set; the basic operator set is a computing node element in the target expert model file that is not related to a hardware device;
[0082] A model calculation graph determining unit, configured to construct a model calculation graph corresponding to the target expert model file based on each calculation node element in the basic operator set, and insert corresponding shape placeholders into the model calculation graph to obtain corresponding memory descriptors;
[0083] The binary intermediate representation determination module is used to define corresponding cross-device sharing attributes in the memory descriptor to obtain the binary intermediate representation corresponding to the target expert model file.
[0084] In some specific embodiments, the device data abstraction module 11 may specifically include:
[0085] The information conversion unit is used to obtain device-related information of each device in the target heterogeneous hardware device cluster, and convert the device-related information into a unified device descriptor through data abstraction processing to obtain an initial device description vector.
[0086] In some specific embodiments, the device data updating module 12 includes:
[0087] a target strategy determining unit, configured to detect a hardware type of a target device in the target heterogeneous hardware device cluster to obtain an instruction set architecture type, and determine a target optimization strategy from a device optimization strategy mapping table in the binary intermediate representation;
[0088] a binary conversion unit, configured to compile the binary intermediate representation into a device native instruction set based on the target optimization strategy and the instruction set architecture type;
[0089] A probe generating unit is configured to inject a probe code into the native instruction set of the device to obtain a hardware probe for the target device, and monitor the real-time status of the target device based on the hardware probe to obtain the real-time status of the device;
[0090] The vector determining unit is configured to update the initial device description vector in real time based on the real-time status of the device to obtain a real-time device description vector.
[0091] In some specific embodiments, the expert model determination module 13 includes:
[0092] A weight determination unit, configured to substitute the real-time device description vector of each device in the target heterogeneous hardware device cluster into a preset routing weight determination function to obtain a hardware-aware routing weight corresponding to each device;
[0093] The target model determining unit is configured to determine Top-K experts from the target expert model file based on the hardware-aware routing weights, so as to obtain an expert model file to be activated corresponding to the target heterogeneous hardware device cluster.
[0094] In some specific embodiments, the expert model determination module 13 includes:
[0095] A sub-model determining unit, configured to split the expert model file to be activated into corresponding expert sub-models based on operator types;
[0096] A data transmission unit is used to determine the to-be-deployed device corresponding to the expert sub-model based on the operator characteristics used by each device in the target heterogeneous hardware device cluster, and to allocate the expert sub-model to the to-be-deployed device for deployment and installation; data transmission is performed between the to-be-deployed devices through a shared virtual address.
[0097] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0098] Figure 4 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the hybrid expert model deployment method based on a heterogeneous hardware device cluster disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0099] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0100] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0101] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to implement the hybrid expert model deployment method based on a heterogeneous hardware device cluster performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.
[0102] Furthermore, this application discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned method for deploying a hybrid expert model based on a heterogeneous hardware device cluster. The specific steps of this method can be found in the corresponding content disclosed in the aforementioned embodiments and will not be repeated here.
[0103] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0104] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0105] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0106] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0107] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A hybrid expert model deployment method based on heterogeneous hardware device clusters, characterized in that: include: Perform data abstraction processing on each device in the target heterogeneous hardware device cluster to obtain an initial device description vector, and compile the target expert model file in the hybrid expert model into a binary intermediate representation; The target expert model file is an initial expert file deployed on each of the devices and not activated; generating a hardware probe for the target device based on the hardware type of the target device in the target heterogeneous hardware device cluster and the binary intermediate representation, and updating the initial device description vector in real time based on the hardware probe to obtain a real-time device description vector; Determining a hardware-aware routing weight of each device in the target heterogeneous hardware device cluster based on the real-time device description vector, and determining a to-be-activated expert model file corresponding to the target heterogeneous hardware device cluster from the target expert model file based on the hardware-aware routing weight; The expert model file to be activated is split into corresponding expert sub-models based on operator types, and the expert sub-models are distributed to the target heterogeneous hardware device cluster for deployment and installation.
2. The hybrid expert model deployment method based on heterogeneous hardware device cluster according to claim 1 is characterized in that: The data abstraction processing of each device in the target heterogeneous hardware device cluster to obtain an initial device description vector includes: The device-related information of each device in the target heterogeneous hardware device cluster is obtained, and the device-related information is converted into a unified device descriptor through data abstraction processing to obtain an initial device description vector.
3. The hybrid expert model deployment method based on heterogeneous hardware device cluster according to claim 1 is characterized in that: Compiling the target expert model file in the hybrid expert model into a binary intermediate representation includes: Performing operator normalization processing on the target expert model file in the hybrid expert model to decompose the target expert model file into a basic operator set; the basic operator set is a computing node element in the target expert model file that is unrelated to the hardware device; Constructing a model calculation graph corresponding to the target expert model file based on each calculation node element in the basic operator set, and inserting corresponding shape placeholders into the model calculation graph to obtain corresponding memory descriptors; Corresponding cross-device sharing attributes are defined in the memory descriptor to obtain a binary intermediate representation corresponding to the target expert model file.
4. The hybrid expert model deployment method based on heterogeneous hardware device cluster according to claim 1 is characterized in that: The generating of a hardware probe for the target device based on the hardware type of the target device in the target heterogeneous hardware device cluster and the binary intermediate representation, and updating the initial device description vector in real time based on the hardware probe to obtain a real-time device description vector, includes: Detecting the hardware type of the target device in the target heterogeneous hardware device cluster to obtain an instruction set architecture type, and determining a target optimization strategy from a device optimization strategy mapping table in the binary intermediate representation; Compiling the binary intermediate representation into a device native instruction set based on the target optimization strategy and the instruction set architecture type; Injecting probe code into the native instruction set of the device to obtain a hardware probe for the target device, and monitoring the real-time status of the target device based on the hardware probe to obtain the real-time status of the device; The initial device description vector is updated in real time based on the real-time status of the device to obtain a real-time device description vector.
5. The hybrid expert model deployment method based on heterogeneous hardware device cluster according to claim 1 is characterized in that: The determining, based on the real-time device description vector, a hardware-aware routing weight of each device in the target heterogeneous hardware device cluster, and determining, from the target expert model file, a to-be-activated expert model file corresponding to the target heterogeneous hardware device cluster, based on the hardware-aware routing weight, includes: Substituting the real-time device description vector of each device in the target heterogeneous hardware device cluster into a preset routing weight determination function to obtain the hardware-aware routing weight corresponding to each device; Top-K experts are determined from the target expert model file based on the hardware-aware routing weights to obtain an expert model file to be activated corresponding to the target heterogeneous hardware device cluster.
6. The hybrid expert model deployment method based on heterogeneous hardware device cluster according to any one of claims 1 to 5, characterized in that: The step of splitting the expert model file to be activated into corresponding expert sub-models based on operator types, and allocating the expert sub-models to the target heterogeneous hardware device cluster for deployment and installation includes: Splitting the expert model file to be activated into corresponding expert sub-models based on operator type; The to-be-deployed device corresponding to the expert sub-model is determined based on the operator characteristics used by each device in the target heterogeneous hardware device cluster, and the expert sub-model is allocated to the to-be-deployed device for deployment and installation; data is transmitted between the to-be-deployed devices via a shared virtual address.
7. A hybrid expert model deployment device based on a heterogeneous hardware device cluster, characterized in that: include: The device data abstraction module is used to perform data abstraction processing on each device in the target heterogeneous hardware device cluster to obtain the initial device description vector and compile the target expert model file in the hybrid expert model into a binary intermediate representation; The target expert model file is an initial expert file deployed on each of the devices and not activated; a device data updating module, configured to generate a hardware probe for the target device in the target heterogeneous hardware device cluster based on the hardware type of the target device and the binary intermediate representation, and to update the initial device description vector in real time based on the hardware probe to obtain a real-time device description vector; an expert model determination module, configured to determine a hardware-aware routing weight of each device in the target heterogeneous hardware device cluster based on the real-time device description vector, and determine a to-be-activated expert model file corresponding to the target heterogeneous hardware device cluster from the target expert model file based on the hardware-aware routing weight; The model deployment module is used to split the expert model file to be activated into corresponding expert sub-models based on operator types, and distribute the expert sub-models to the target heterogeneous hardware device cluster for deployment and installation.
8. The hybrid expert model deployment device based on heterogeneous hardware device cluster according to claim 7, characterized in that: The device data abstraction module includes: An operator decomposition unit is used to perform operator normalization processing on a target expert model file in a hybrid expert model to decompose the target expert model file into a basic operator set; the basic operator set is a computing node element in the target expert model file that is not related to a hardware device; A model calculation graph determining unit, configured to construct a model calculation graph corresponding to the target expert model file based on each calculation node element in the basic operator set, and insert corresponding shape placeholders into the model calculation graph to obtain corresponding memory descriptors; The binary intermediate representation determination module is used to define corresponding cross-device sharing attributes in the memory descriptor to obtain the binary intermediate representation corresponding to the target expert model file.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the hybrid expert model deployment method based on a heterogeneous hardware device cluster as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that Used to store a computer program, which, when executed by a processor, implements the hybrid expert model deployment method based on a heterogeneous hardware device cluster as described in any one of claims 1 to 6.
Citation Information
Cited By
Construction method of cross-platform high-performance tensor program generation model and related device
CN121278388A