Model deployment method and device combined with hardware deployment
By building hardware topology diagrams and model calculation diagrams, the matching values of device nodes and target operators are calculated automatically to generate the optimal deployment solution, which solves the problems of hardware topology mismatch and slow response to resource changes in model deployment, and realizes real-time adaptation and automated deployment of AI systems.
Patent Information
- Application Number
- CN202510703640.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
In the existing technology, model deployment relies on the experience of technicians, resulting in the model parallelism strategy that does not match the hardware topology, cannot respond to changes in hardware resources in real time, lack of automation mechanisms, and limits the elastic expansion of AI systems.
By building hardware topology diagrams and model calculation diagrams, using the evaluation algorithm to calculate the matching values between the device nodes and the target operators, automatically generate the optimal deployment plan according to the hierarchical strategy, perceive the hardware state in real time and trigger adaptive optimization.
Real-time collaborative adaptation between model segmentation strategies and hardware environments is realized, the deployment automation level and system flexibility are improved, and bottleneck problems in static deployment mode are solved.
Smart Images

Figure CN120234014A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a model deployment method and device combined with hardware deployment. Background Art
[0002] The purpose of artificial intelligence model deployment is to transform large model technology into actual productivity. Model deployment promotes the technology from laboratory to industrial application by optimizing computing power resources, improving data processing efficiency, and enhancing scene adaptability.
[0003] In current model deployment practices, technicians mostly adopt an experience-driven offline model segmentation strategy, and deploy the model to the target device at one time through manual prediction. This method has three limitations: First, it overly relies on technicians' understanding of the hardware topology, which easily leads to a deviation between the model parallel strategy and the actual hardware architecture; Second, when the computing node configuration or network topology changes dynamically, manual re-evaluation of resources and strategy formulation are required, lacking real-time response capabilities; Third, the entire deployment process lacks an automated mechanism, making it difficult to achieve dynamic adaptation between the model segmentation strategy and the hardware environment. This static deployment mode has become a key bottleneck restricting the elastic expansion of AI systems. Summary of the Invention
[0004] This application provides a model deployment method and device combined with hardware deployment to solve the problems of relying on technicians' experience, which easily leads to a mismatch between the model parallel strategy and the physical connection topology, and after the hardware resources change, manual re-evaluation and adoption of a new model segmentation strategy are required, and the entire process cannot be automated.
[0005] In a first aspect, this application provides a model deployment method combined with hardware deployment, including: Construct a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each device node; Analyze the dependency relationships of the target operators in the target model to generate a model computation graph corresponding to the target operators; According to the hardware topology graph and the model computation graph, use a preset evaluation algorithm to calculate the target matching value between the device nodes and the target operators; Determine the deployment plan corresponding to the target model according to the target matching value and a preset grading strategy.
[0006] In a second aspect, this application provides a model deployment device combined with hardware deployment, including: A hardware topology graph determination module configured to construct a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each device node; A model calculation graph determination module, configured to parse the dependency relationship of target operators in a target model to generate a model calculation graph corresponding to the target operators; A target matching value determination module, configured to calculate a target matching value between a device node and a target operator according to a hardware topology graph and a model calculation graph by using a preset evaluation algorithm; A deployment plan determination module, configured to determine a deployment plan corresponding to the target model according to the target matching value and a preset classification strategy.
[0007] In a third aspect, the present application provides a readable medium, including execution instructions, when a processor of an electronic device executes the execution instructions, the electronic device executes the method described in any one of the first aspects.
[0008] In a fourth aspect, the present application provides an electronic device, including a processor and a memory storing execution instructions, when the processor executes the execution instructions stored in the memory, the processor executes the method described in any one of the first aspects.
[0009] The present application provides a model deployment method combined with hardware deployment. According to device nodes in a target model deployment environment and the corresponding connection relationships of each device node, a hardware topology graph corresponding to the device nodes is constructed; the dependency relationship of target operators in the target model is parsed to generate a model calculation graph corresponding to the target operators; according to the hardware topology graph and the model calculation graph, a target matching value between a device node and a target operator is calculated by using a preset evaluation algorithm; according to the target matching value and a preset classification strategy, a deployment plan corresponding to the target model is determined. Through a dynamic matching mechanism that perceives the hardware state and the model calculation graph in real time, an optimal deployment strategy is automatically generated, and adaptive optimization is triggered when the hardware topology or resource conditions change, realizing real-time collaborative adaptation of the model segmentation strategy and the heterogeneous environment, and significantly improving the deployment automation level and system elasticity.
[0010] The further effects of the above non-conventional preferred methods will be described below in conjunction with specific embodiments. Description of the Drawings
[0011] In order to more clearly illustrate the embodiments of the present application or the existing technical solutions, the following will briefly introduce the drawings required for use in the description of the embodiments or the existing technical solutions. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic flowchart of a model deployment method combined with hardware deployment provided by an embodiment of the present application; Figure 2A schematic flowchart of another model deployment method combined with hardware deployment provided by an embodiment of the present application; Figure 3 A schematic flowchart of another model deployment method combined with hardware deployment provided by an embodiment of the present application; Figure 4 A schematic flowchart of another model deployment method combined with hardware deployment provided by an embodiment of the present application; Figure 5 A schematic structural diagram of a model deployment device combined with hardware deployment provided by an embodiment of the present application; Figure 6 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0013] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0014] The purpose of artificial intelligence model deployment is to transform large model technology into actual productivity. Model deployment promotes the technology from laboratory to industrial application by optimizing computing power resources, improving data processing efficiency, and enhancing scenario adaptability.
[0015] In current model deployment practices, technicians mostly adopt an experience-driven offline model segmentation strategy, and deploy the model to the target device at one time through manual prediction. This method has three limitations: First, it overly relies on technicians' understanding of the hardware topology, which easily leads to deviations between the model parallel strategy and the actual hardware architecture; Second, when the computing node configuration or network topology changes dynamically, manual re-evaluation of resources and strategy formulation are required, lacking real-time response capabilities; Third, the entire deployment process lacks an automated mechanism and it is difficult to achieve dynamic adaptation between the model segmentation strategy and the hardware environment. This static deployment mode has become the key bottleneck restricting the elastic expansion of AI systems.
[0016] To solve this problem, the embodiments of the present application propose a model deployment method combined with hardware deployment, aiming to solve the problem of relying on technicians' experience as described above, which easily leads to mismatches between the model parallel strategy and the physical connection topology. Moreover, after the hardware resources change, manual re-evaluation and adoption of a new model segmentation strategy are required, and the entire process cannot be automated. In this embodiment, a model deployment method combined with hardware deployment includes: Step 101: Construct a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each device node.
[0017] In the model deployment environment, a device node represents a physical hardware unit with independent computing and storage capabilities, such as a GPU, CPU, TPU, or edge computing device. The core attributes of each device node include static performance indicators such as theoretical computing power, video memory capacity, and input / output bandwidth upper limits, as well as dynamic state information such as real-time available video memory, current computing power load, and actual bandwidth utilization rate. For example, in a GPU cluster, the device nodes need to clarify the number of CUDA cores, video memory specifications, and interconnection topology, while in an edge computing scenario, the device nodes may cover heterogeneous CPUs and low-power accelerators, and their energy consumption constraints and other characteristics need to be synchronously collected.
[0018] The connection relationship between devices is quantitatively modeled through communication costs, and its core parameters include physical transmission delay and effective bandwidth. For example, if two GPUs are connected through PCIe 4.0 x16, their theoretical bidirectional bandwidth is 32 GB / s, but the actual effective bandwidth may be affected by protocol overhead or link contention, and the minimum bidirectional measured value needs to be taken; the delay includes the startup time and propagation time of data transmission. The communication cost can be further converted into the time cost of transmitting a unit data volume to characterize the efficiency difference of data exchange between devices.
[0019] Based on the above information, the hardware topology graph is abstractly expressed in the form of a graph structure, where the vertex set corresponds to the device nodes, and the edge weights map the communication costs. For example, in a multi-GPU training cluster, the edge weights between devices interconnected by NVLink are significantly lower than those connected by PCIe. Such high-bandwidth and low-latency links will be preferentially used for high-frequency parameter synchronization.
[0020] Through this structured modeling, the hardware topology graph not only intuitively presents the resource distribution and connection characteristics, but also provides a global decision-making basis for dynamic optimization algorithms. For example, when resources fluctuate, it automatically adjusts the task allocation path to achieve the global optimal balance of computing power and communication efficiency.
[0021] Step 102: Analyze the dependency relationship of the target operator in the target model to generate a model computation graph corresponding to the target operator.
[0022] To analyze the dependency relationships of target operators in a target model, we need to start from the computational logic of the model and deconstruct its execution process and data flow path layer by layer. Taking a deep learning model as an example, first, by statically analyzing the model structure (such as the connection order of neural network layers, branch conditions, or loop control flows), we extract the input-output relationships of all operators. For example, the output tensor of a convolutional layer may be used as the input of an activation function layer, and the addition operation in a residual connection needs to synchronously aggregate the outputs of multiple branches. Such dependency relationships directly define the execution order of operators and the data transfer rules.
[0023] Based on this, the model computational graph is abstractly represented in the form of a directed graph, where vertices represent operators and edges represent data flow dependencies. Each operator vertex needs to record its computational cost (such as the number of floating-point operations) and resource requirement characteristics (such as video memory occupancy, bandwidth requirements). The data flow edges, on the other hand, quantify the data transfer volume between adjacent operators (such as tensor dimensions and data types). For example, in a Transformer model, the multi-head computation of the self-attention layer generates an intermediate key-value pair matrix, and its transfer volume needs to be dynamically calculated according to the number of heads and sequence length. Such information will directly affect the evaluation of subsequent inter-device communication costs.
[0024] For complex models (such as graph structures containing dynamic control flows or conditional branches), it is necessary to further introduce symbolic execution or dynamic tracing techniques to ensure the integrity and accuracy of dependency relationships. For example, in a recurrent neural network, the hidden state transfer between time steps may form a cross-iteration data dependency chain, and its computational path needs to be clarified by unfolding along the time dimension. The finally generated model computational graph not only needs to reflect the explicit data flow between operators but also capture implicit resource competition relationships (such as the mutual exclusion between operators sharing video memory buffers), thus providing an optimization basis for dynamic deployment.
[0025] Step 103: According to the hardware topology graph and the model computational graph, use a preset evaluation algorithm to calculate the target matching value between the device node and the target operator.
[0026] During the dynamic deployment process, the calculation of the target matching value first extracts the multi-dimensional resource requirement characteristics of each operator from the model computational graph, including dimensions such as computing power requirements (such as the number of floating-point operations), video memory occupancy, bandwidth requirements, and quantization support requirements. For example, a convolutional operator may have high requirements for video memory capacity, while a fully connected layer is more dependent on the parallel processing ability of high-computing-power devices.
[0027] At the same time, the hardware topology graph provides the real-time capability characteristics of the device node, covering parameters such as the theoretical computing power limit, current available video memory, measured bandwidth, and supported quantization precision. For example, the available video memory of a certain GPU node may be dynamically reduced due to the occupation of other tasks, and its actual bandwidth may also be lower than the theoretical value due to network congestion.
[0028] The evaluation algorithm quantifies the model through multi-dimensional matching degrees, and analyzes the differences between operator requirements and device capabilities. Specifically, for each resource dimension (such as video memory), the normalized difference between operator requirements and device capabilities can be calculated. First, the mean-variance normalization can be performed on the parameters of each dimension to eliminate the dimensional differences (for example, converting the video memory requirement from GB to the same standardized unit as the device's video memory). Subsequently, the differences are mapped to a single-dimensional matching degree through an exponential function. Finally, the matching degrees of all dimensions are combined into a global target matching value through a product form. For example, if an operator has high requirements for video memory and bandwidth, and a device has sufficient video memory but limited bandwidth, its comprehensive matching value will be significantly lower than that of a device with both suitable video memory and bandwidth.
[0029] Step 104: Determine the deployment plan corresponding to the target model according to the target matching value and the preset grading strategy.
[0030] Determine the target matching level corresponding to the target matching value according to the grading threshold corresponding to the grading strategy; the target matching levels include high matching level, low matching level, and medium matching level; determine the deployment plan based on the target matching level.
[0031] In the formulation of the dynamic deployment plan, the grading strategy is based on the quantization result of the target matching value, divides the device-operator adaptability into three levels: high, medium, and low according to the grading threshold, and adopts different optimization means for different levels. The design of the grading threshold needs to balance the continuity of resource adaptation and the discreteness of policy execution, and the threshold setting is usually based on historical deployment data statistics and scenario-based experimental verification.
[0032] For example, in the video memory-intensive task scenario, if the target matching values of 90% of the efficient deployment cases are concentrated above 0.8, the threshold for the high matching level can be initially set to 0.8, and the target matching value below 0.3 may correspond to the low-level scenario with serious resource shortage.
[0033] When the target matching value belongs to the high matching level, based on the numerical size of the target matching value, prioritize the target operators to generate a deployment queue corresponding to the target operators; deploy the target operators to the corresponding device nodes according to the deployment queue.
[0034] When the target matching value belongs to the high matching level, the system realizes the optimal allocation of resources through the priority sorting mechanism. First, globally sort the highly adaptable operators based on the numerical size of the matching values. For example, prioritize the convolutional operator with a matching value of 0.95 over the fully connected layer operator with a matching value of 0.88 to form a deployment queue.
[0035] This sorting not only reflects the absolute level of device-operator adaptation, but also needs to be dynamically adjusted in combination with the dependency relationship of the model computation graph. If a high-matching operator has a predecessor dependency (such as the output of the data preprocessing layer being the input of the subsequent layer), then its deployment order needs to be optimized under the premise of satisfying the dependency chain to avoid resource idleness caused by task blocking. For example, in an image recognition model, although the matching value of the feature extraction operator is slightly lower than that of the classification layer, it still needs to be deployed first as the starting point of the computation process to ensure the coherence of the overall task pipeline.
[0036] After generating the deployment queue, the system performs dynamic task allocation according to the real-time load status of the device nodes. For example, if a certain GPU is currently carrying high-computing-power tasks, even if it has the highest matching value with the target operator, the system may choose a sub-optimal but less-loaded device to avoid performance degradation caused by local resource overload. Such decisions need to comprehensively evaluate multi-dimensional metrics such as the available video memory of the device, bandwidth utilization, and task queuing duration. For example, the deployment order is dynamically adjusted through a weighted scoring model.
[0037] The execution of the deployment queue needs to be deeply coupled with the communication cost of the hardware topology. For example, in distributed training, the high-matching parameter server is preferentially deployed to the device group with high interconnection bandwidth and low latency, while ensuring the shortest data exchange path between adjacent computing nodes. If a sudden change in device status (such as a sudden decrease in video memory or network congestion) is detected during the deployment process, the system will trigger a dynamic reconstruction of the queue - pause the current deployment, recalculate the matching value, and generate an updated queue to maintain the global resource utilization efficiency.
[0038] This closed-loop optimization mechanism enables the deployment strategy with a high matching level to not only have static priorities but also respond to environmental changes in real time. For example, in an autonomous driving scenario, when the on-vehicle computing unit experiences a sudden increase in load due to a sudden task, the system automatically migrates some high-matching tasks to the roadside edge nodes to ensure that the decision-making tasks with the highest real-time requirements continuously obtain optimal resource support.
[0039] When the target matching value belongs to the low matching level, determine the original subgraph corresponding to the target operator in the model computation graph; use the computation path recombination algorithm to recombine the paths of the original subgraph to generate an optimized subgraph; replace the original subgraph with the optimized subgraph to generate a corresponding optimized computation graph; calculate the first matching value corresponding to the optimized computation graph; when the first matching value does not belong to the low matching level, update the model computation graph with the optimized subgraph.
[0040] When the target matching value falls into the low matching level, the system will activate the calculation path reorganization mechanism. First, locate the original subgraph associated with the low-matching operator in the model computational graph, including the data flow dependencies of its predecessor and successor nodes. For example, if the self-attention layer in a Transformer model is determined to be a low match due to exceeding the video memory requirement limit, it is necessary to extract the subgraph composed of this layer and associated operators such as normalization and residual connection, and clarify the flow path of the input and output tensors and the resource occupancy characteristics.
[0041] The path reorganization algorithm optimizes the structure of the original subgraph through various strategies. For example, in a scenario with limited video memory, operator fusion technology can be used to merge multiple consecutive small operators into a composite operator, reducing the temporary storage requirement of intermediate results in the video memory. For extremely large operators in a distributed scenario, they can be disassembled into multiple subtasks, and the load can be shared through collaborative computing between devices. For example, decompose large-scale matrix multiplication into block sub-matrices, deploy them to multiple GPUs for parallel execution respectively, and ensure the consistency of calculation results through communication synchronization.
[0042] After generating the optimized subgraph, the system replaces the original subgraph with it to form a new optimized computational graph, and recalculates the first matching value of the device-operator. If the new matching value exits the low matching level (such as being promoted to the medium or high level), it triggers the dynamic update of the model computational graph, and incorporates the optimized subgraph into the deployment process. For example, an edge device could not execute the original convolution operator due to insufficient computing power. After path reorganization, it is replaced with a fused subgraph of depthwise separable convolution and grouped convolution, and its matching value is promoted from 0.3 (low level) to 0.8 (high level), successfully exiting the low matching range.
[0043] At this time, the system synchronizes the updated computational graph to all relevant devices and reallocates tasks. If the optimized matching value still remains at the low level, it activates the iterative optimization process, tries other reorganization strategies (such as introducing sparse computing or dynamic pruning), until the deployment conditions are met or the degradation fault tolerance mechanism is triggered.
[0044] When the target matching value belongs to the medium matching level, analyze the resource requirement characteristics to determine the bottleneck dimension characteristics that result in the target matching value being at the medium matching level; according to the bottleneck dimension characteristics, adjust the connection relationship from the initial connection to the optimized connection; determine the second matching value corresponding to the optimized connection; when the second matching value belongs to the high matching level, use the optimized connection to update the connection relationship in the hardware topology graph.
[0045] When the target matching value is at the medium matching level, the system will activate the bottleneck dimension diagnosis and topology optimization mechanism to break through local resource constraints. By analyzing the multi-dimensional resource requirement characteristics of the operator (such as video memory, bandwidth, computing power), locate the core bottleneck that causes the insufficient matching value.
[0046] For example, the matching value of a certain TPU node is classified as medium level because the video memory occupancy rate is close to the threshold (such as 90%), and the computing power and bandwidth dimensions are well adapted. At this time, the video memory becomes the limiting dimension; or the bandwidth utilization rate of a certain edge device continuously exceeds 80%, resulting in a sharp increase in data transmission delay and dragging down the overall matching value. This kind of diagnosis needs to combine real-time monitoring data with historical trend analysis. For example, by tracking the time correlation of the bandwidth utilization rate, the accuracy of bottleneck determination can be ensured.
[0047] For the identified bottleneck dimension, the system dynamically adjusts the connection relationship of the hardware topology. For example, in a bandwidth-limited scenario, if the original connection bandwidth between device A and device B is insufficient, a relay device C can be introduced to build a transmission path of A→C→B, and the single-path pressure can be shared through multiple links; or switch to an alternative high-bandwidth link.
[0048] After completing the connection optimization, the system recalculates the second matching value of the device-operator to verify the adjustment effect. If the new matching value is upgraded to a high level, it triggers the real-time update of the hardware topology graph - synchronize the optimized connection relationship to the global view, and regenerate the deployment strategy based on the new topology.
[0049] For example, the original deployment of a certain video encoding operator has a matching value of 0.7 (medium level) due to insufficient bandwidth of the edge device. After switching to a 5G millimeter-wave link, the bandwidth is increased from 50 Mbps to 1 Gbps, and the second matching value jumps to 0.92. The system immediately marks the new link as the preferred path and migrates the operator to the edge device for execution, making full use of local computing power to reduce the dependence on the cloud. If the matching value after optimization is still at the medium level or downgraded, the bottleneck diagnosis result is traced back, and other optimization directions are tried (such as adjusting the task allocation ratio or introducing redundant computing nodes).
[0050] From the above technical solutions, the beneficial effects of this embodiment are as follows: Construct a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each device node; analyze the dependency relationships of the target operators in the target model to generate a model calculation graph corresponding to the target operators; according to the hardware topology graph and the model calculation graph, use a preset evaluation algorithm to calculate the target matching value between the device nodes and the target operators; according to the target matching value and the preset classification strategy, determine the deployment plan corresponding to the target model. Through the dynamic matching mechanism of real-time perception of the hardware state and the model calculation graph, an optimal deployment strategy is automatically generated, and adaptive optimization is triggered when the hardware topology or resource conditions change, realizing the real-time collaborative adaptation of the model segmentation strategy and the heterogeneous environment, and significantly improving the deployment automation level and system elasticity.
[0051] Figure 1The following only shows a basic embodiment of a model deployment method combined with hardware deployment in this application. Based on this, with certain optimizations and expansions, other preferred embodiments of a model deployment method combined with hardware deployment can also be obtained.
[0052] As Figure 2 shown, it is another specific embodiment of a model deployment method combined with hardware deployment in this application.
[0053] In this embodiment, a model deployment method combined with hardware deployment includes the following steps: Step 201: Construct a hardware topology diagram corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each device node.
[0054] Step 202: Determine the device information corresponding to each device node.
[0055] Device information is a core data set that describes the performance and status of physical hardware units in a computing environment, covering their inherent capabilities, real-time operating status, and connection characteristics. In the model deployment environment, determining the device information of each device node requires systematic collection and modeling from two dimensions: static capabilities and dynamic status.
[0056] Static capability characteristics reflect the inherent performance parameters of the hardware. For example, the number of CUDA cores, theoretical floating-point computing power, video memory capacity, supported interconnection protocols, and their theoretical bandwidth limits of the GPU. For heterogeneous devices (such as TPU or FPGA), it is also necessary to record the characteristics of dedicated computing units (such as the matrix multiplication acceleration ability of the TPU) or the scale of programmable logic resources. This information is usually obtained through hardware specification documents or system-level query tools, and a device capability database is established to provide a benchmark basis for resource allocation.
[0057] Dynamic status information, on the other hand, requires real-time monitoring of the resource usage of the device during operation, including the current available video memory ratio, computing power load, actual bandwidth utilization rate, device temperature, and energy consumption. In addition, the connection status between devices (such as link failures or bandwidth fluctuations) also needs to be dynamically updated. For example, the availability of the NVLink channel is verified through a heartbeat detection mechanism, or the delay and packet loss rate between nodes are measured through network probes.
[0058] Step 203: Determine the communication cost corresponding to the device node according to the device information and the connection relationship.
[0059] Communication cost is a core metric for measuring the data transmission efficiency between devices, and its calculation requires integrating the physical characteristics and real-time status of the device connection relationship. Specifically, the communication cost is jointly determined by the transmission delay and the effective bandwidth. The effective bandwidth takes the smaller value of the bidirectional bandwidth between devices to reflect the upper limit of the transmission capacity of the actual link. For example, if the bandwidth from device A to device B is 10 Gbps and the reverse bandwidth is 8 Gbps, the effective bandwidth is 8 Gbps to avoid misjudging the transmission efficiency due to link asymmetry.
[0060] In a dynamic environment, the communication cost needs to be updated in real time to adapt to network state fluctuations. For example, in distributed training, if the actual bandwidth of a certain NVLink link decreases due to a sudden increase in load, the system needs to re-measure its effective bandwidth and update the communication cost; in an edge computing scenario, the bandwidth of a 5G link may suddenly drop from 1 Gbps to 200 Mbps due to signal strength changes. At this time, the cost calculation needs to be dynamically adjusted to avoid allocating high-data-volume tasks to inefficient links. In addition, changes in the physical topology (such as adding relay devices or switching faulty links) will also trigger recalculation of the communication cost. For example, by introducing cache nodes to build multi-hop paths and splitting a single high-latency link into multiple low-latency sub-paths, the overall transmission cost can be reduced.
[0061] In practical applications, optimizing the communication cost directly affects the task deployment strategy. For example, in a GPU cluster, NVLink direct-connected devices with high bandwidth and low latency are preferentially used for adjacent computing nodes that frequently synchronize parameters; in a cloud-edge collaborative inference scenario, preprocessing tasks are deployed to edge devices to reduce the need to transmit raw data to the cloud, thus avoiding the impact of high public network latency. By dynamically perceiving and calculating the communication cost in real time, the system can achieve optimal allocation of data streams and computing tasks in a complex heterogeneous environment, maximizing the overall utilization rate of hardware resources.
[0062] Step 204: Construct a hardware topology graph with device nodes as vertices and communication cost as edge weights.
[0063] In the specific process, with device nodes as vertices, their attributes need to completely record static capabilities (such as computing power, video memory capacity, protocol support) and dynamic states (such as real-time available video memory, bandwidth utilization, temperature); while the connection relationship between devices is quantified by edge weights, and the weight values are dynamically calculated by the communication cost formula.
[0064] For example, in a GPU cluster, if GPU0 and GPU1 are directly connected via NVLink 3.0 (bandwidth 600 GB / s, latency 0.1 μs), their edge weights are significantly lower than those of GPU2 and GPU3 connected via PCIe 4.0 x16 (bandwidth 32 GB / s, latency 1 μs). This difference directly affects the task allocation strategy - operators with high-frequency data exchange (such as parameter synchronization) will be preferentially deployed to the NVLink direct connection device group to avoid communication bottlenecks.
[0065] The construction of the hardware topology graph needs to take into account both static initialization and dynamic update. In the initialization phase, the basic attributes of the devices are obtained through system detection tools, and the initial latency and bandwidth are measured through network performance tests to generate a weighted graph structure.
[0066] In a dynamic scenario, the topology graph needs to be updated periodically: if the actual bandwidth of a device decreases due to a sharp increase in load (e.g., from 10 Gbps to 5 Gbps), or a link failure triggers a standby path switch (e.g., from fiber optic to 5G redundant link), the system will recalculate the edge weights and update the graph structure to ensure the real-time nature of the resource view.
[0067] Step 205: Analyze the dependency relationship of the target operator in the target model to generate the model calculation corresponding to the target operator.
[0068] Step 206: According to the hardware topology graph and the model calculation graph, use a preset evaluation algorithm to calculate the target matching value between the device node and the target operator.
[0069] Step 207: Determine the deployment plan corresponding to the target model according to the target matching value and the preset grading strategy.
[0070] From the above technical solutions, it can be seen that the beneficial effects of this embodiment are: the dynamic update mechanism of the hardware topology graph ensures its real-time nature and accuracy in a heterogeneous environment, providing underlying support for the elasticity and self-adaptability of model deployment.
[0071] As Figure 3 shown, this is another specific embodiment of a model deployment method combining hardware deployment in this application. This embodiment is further described on the basis of the foregoing embodiment.
[0072] In this embodiment, a model deployment method combining hardware deployment includes the following steps: Step 301: Construct a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each device node.
[0073] Step 302: Analyze the dependency relationship of the target operator in the target model to generate a model calculation graph corresponding to the target operator.
[0074] Step 303: According to the calculation process corresponding to the target model, analyze the dependency relationships among the target operators to generate data flow edges corresponding to the target operators.
[0075] During the model deployment process, analyzing the dependency relationships of the target operators needs to start from the computational logic of the model. By combining static analysis and dynamic tracing, clarify the input-output dependencies and execution order among the operators. Taking a typical deep learning model as an example, first extract the connection topology of the operators by parsing the model definition. For example, the output tensor of the convolutional layer serves as the input to the activation function layer, and the addition operation in the residual connection needs to synchronously receive the output results of the main path and the bypass branch. Such explicit dependencies directly determine the execution timing of the operators.
[0076] For dynamic graph models or structures with conditional branches (such as the dynamic masking mechanism in Transformer), it is necessary to combine symbolic execution or runtime instrumentation techniques to capture the data flow paths during the actual inference process. For example, in a reinforcement learning policy network, action selection may depend on the real-time feedback of the environmental state. It is necessary to convert implicit dependencies into explicit data flow edges through multiple samplings or dynamic graph unfolding.
[0077] In complex model scenarios, analyzing the dependency relationships needs to handle multi-modal inputs and heterogeneous data streams. For example, in a multi-task learning model, the output of the same feature extraction layer may flow to both the classification head and the regression head simultaneously, forming branch dependencies; while in a time series model, the cross-time-step transmission of hidden states introduces cyclic dependencies. For such scenarios, it is necessary to introduce virtual operators or time step markers to enhance the expressive power of the computational graph.
[0078] For example, in a recurrent neural network, by unfolding the time steps into independent computational nodes, explicitly construct the data flow edges between adjacent time steps to ensure the accuracy of the gradient propagation path. At the same time, for control flow, it is necessary to represent the conditional judgment logic through placeholder operators and dynamically activate the corresponding sub-graph paths according to runtime decisions. For example, in the model compression scenario, select different sub-network branches according to the input complexity.
[0079] The generation of data flow edges needs to quantify the data transmission scale and type between operators. For example, in an object detection model, the feature pyramid network needs to transfer feature maps of different scales to the fusion layer. Each data flow edge needs to record the dimension, data type, and batch size of the feature map to accurately calculate the communication cost.
[0080] Step 304: Using the target operators as vertices and the data flow edges as directed edges, construct the model computational graph.
[0081] The core of constructing a model computational graph lies in abstracting target operators as vertices and characterizing the dependency relationships and data flow logics between them through data flow edges. First, the computational process of the model needs to be decomposed into atomic operator units, such as convolution, matrix multiplication, or activation functions. Each operator serves as a vertex in the graph, and its attributes need to record key parameters in detail, such as computational cost, peak VRAM occupancy, input and output tensor dimensions, and bandwidth requirements.
[0082] For example, in the Transformer model, the vertices of the self-attention layer need to label the VRAM capacity required for its multi-head computation and quantify the transmission scale of the key-value matrix, while the vertices of the fully connected layer need to clarify the dimensions and computational intensity of the weight matrix.
[0083] As a directed edge connecting operators, the data flow edge needs to precisely describe the data transfer relationship and transmission characteristics between operators. The direction of the edge is determined by the computational dependency relationship: if the output tensor of operator A is the input of operator B, there is a directed edge from A to B. The weight of the edge needs to quantify the amount of data transmitted (such as the byte size of the tensor), data type (such as float32 or int8), and bandwidth sensitivity (such as high throughput requirement or low latency priority).
[0084] For dynamic models (such as recurrent neural networks), the transfer of hidden states between time steps needs to generate strided data flow edges through time unfolding. Each edge corresponds to the output of a specific time step and the input dependency of the next time step, while recording the dimensions and update frequency of the hidden states.
[0085] The construction of the model computational graph also needs to handle complex control flows and heterogeneous data paths. For example, in a model containing conditional branches (such as a multi-task learning network), the data flow edges may fork into different subgraph branches, and a virtual decision operator needs to be introduced to represent the branch selection logic. In real-time stream processing scenarios, if the model supports dynamic batching, the data flow edges need to be attached with the range of batch sizes and elastic scaling strategies. For example, when the resources of edge devices fluctuate, the batch size can be automatically adjusted to adapt to the available VRAM.
[0086] In addition, for parameter synchronization in distributed training, the computational graph needs to explicitly label the gradient aggregation edges, and the transmission volume of its data flow edges is determined by the parameter scale and update frequency and is associated with the high-bandwidth links in the hardware topology graph.
[0087] Step 305: According to the hardware topology graph and the model computational graph, use a preset evaluation algorithm to calculate the target matching value between the device node and the target operator.
[0088] Step 306: According to the target matching value and the preset grading strategy, determine the deployment plan corresponding to the target model.
[0089] As can be seen from the above technical solutions, the beneficial effects of this embodiment are as follows: For dynamic or unstructured models, introducing virtual operators and dynamic edges can enhance the expressive power of the computational graph. This refined modeling enables the model's computational graph to comprehensively reflect the computational logic, resource requirements, and performance bottlenecks, providing an accurate optimization basis for subsequent dynamic deployment.
[0090] As Figure 4 shown, this is another specific embodiment of a model deployment method combined with hardware deployment in this application. This embodiment is further described based on the foregoing embodiment.
[0091] Step 401: Construct a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each device node.
[0092] Step 402: Analyze the dependency relationships of the target operators in the target model to generate a model computational graph corresponding to the target operators.
[0093] Step 403: According to the hardware topology graph and the model computational graph, use a preset evaluation algorithm to calculate the target matching value between the device nodes and the target operators.
[0094] Step 404: Determine the resource requirement characteristics corresponding to each target operator according to the model computational graph.
[0095] The resource requirement characteristics are multi-dimensional quantization indicators that describe the hardware resource requirements of model operators (such as convolutional layers, fully connected layers, etc.) during operation. The resource requirement characteristics are obtained by analyzing the computational logic, data flow, and hardware adaptability of the operators.
[0096] First, analyze its computational characteristics based on the operator type (such as convolution, matrix multiplication, activation function): Convolutional layers usually require high parallel computing power (such as the CUDA cores of GPUs) and video memory capacity to store weights and feature maps; attention mechanism operators are sensitive to bandwidth because they need to frequently transmit key-value matrices; while quantization operators (such as INT8 inference) may rely on specific hardware acceleration units (such as Tensor Cores). For example, the self-attention layer in the Transformer model needs to quantify the video memory occupancy of its multi-head calculations and clarify the bandwidth requirements for cross-device transmission of key-value matrices.
[0097] The transmission characteristics recorded by the data flow edges directly affect the bandwidth requirements. For example, in the feature pyramid network, the data volume of the cross-layer fusion edges is determined by the feature map resolution and the number of channels. If sparsification or quantization compression is adopted, the bandwidth requirements need to be adjusted to 25% of the original value.
[0098] For dynamic models (such as recurrent neural networks), the amount of data transferred between hidden states at each time step needs to be dynamically calculated based on the sequence length and the hidden layer dimension. The resource requirement characteristics need to be aligned with the hardware capabilities to form multi-dimensional quantization metrics.
[0099] Step 405: Determine the device capability characteristics corresponding to each device node according to the hardware topology map.
[0100] Device capability characteristics are a set of multi-dimensional attributes that describe the comprehensive performance and real-time status of hardware units in a computing environment, covering their inherent capabilities, runtime resource status, and connection efficiency. In a heterogeneous computing environment, the determination of device capability characteristics needs to be globally modeled from three dimensions: static performance baseline, dynamic resource status, and connection efficiency. These capability characteristics are structured and stored through a unified data model and integrated into the vertex attributes of the hardware topology map.
[0101] The static performance baseline reflects the inherent hardware attributes of the device. For example, the number of CUDA cores, theoretical floating-point computing power, video memory capacity, and supported interconnection protocols and their theoretical bandwidth limits of a GPU. For dedicated accelerators (such as TPU or NPU), it is necessary to record their dedicated computing unit characteristics (such as the throughput of the matrix multiplication engine of a TPU) or instruction set support (such as INT8 quantization acceleration ability).
[0102] The dynamic resource status is periodically collected through a real-time monitoring system, including the current available video memory ratio (such as 18GB remaining), computing power load (such as 70% occupancy of streaming multiprocessors), actual bandwidth utilization (such as the NVLink bandwidth dropping to 400 GB / s due to network congestion), device temperature, and power consumption, etc.
[0103] For example, in distributed training, due to long-term operation, the video memory of a certain GPU becomes fragmented, and the available video memory drops from 30GB to 22GB. At this time, its dynamic video memory capacity needs to be updated in real time to avoid triggering a memory overflow during subsequent task allocation. The collection frequency of the dynamic state needs to balance accuracy and overhead. For example, the computing power load is sampled once per second in a real-time inference scenario, while the video memory occupancy is updated every 5 seconds in offline training.
[0104] The connection efficiency is indirectly characterized by the communication cost (such as latency and effective bandwidth) between devices. For example, the communication cost (latency + data volume / bandwidth * latency + bandwidth * data volume) of two GPU devices directly connected through NVLink is significantly lower than that of a device group connected through PCIe; in an edge-cloud collaboration scenario, the public network link between the edge node and the cloud TPU may introduce additional latency (such as increasing from 1 ms to 10 ms) due to an increase in the number of routing hops, which needs to be dynamically measured and updated to the topology map.
[0105] Step 406: Calculate the difference value between the resource requirement characteristics and the device capability characteristics.
[0106] In a dynamic deployment system, the calculation of the difference value aims to quantify the degree of fit between the resource requirements of model operators and the device capabilities. Through multi-dimensional normalization and non-linear mapping, the differences in heterogeneous resource dimensions are transformed into comparable numerical values.
[0107] Specifically, first, it is necessary to align the dimensionality of the resource requirement characteristics of the operator and the device capabilities. For example, for the video memory dimension, if a certain convolutional operator requires 20GB of video memory and the available video memory of the target device is 25GB, the original difference is 20 - 25 = -5GB. The negative value indicates that the device capabilities can cover the requirements; however, if the available video memory of the device is only 15GB, the difference is +5GB, indicating a shortage of resources.
[0108] Step 407: Input the difference value into a preset mathematical model so that the mathematical model outputs a target matching value.
[0109] The process of inputting the difference value into a preset mathematical model to generate a target matching value is achieved through non-linear mapping of multi-dimensional resource adaptability. The formula is: , where represents the matching degree between device node i and operator j. By quantifying the degree of fit between device capabilities and operator resource requirements, it reflects the optimization potential of their collaboration. The higher the value, the more suitable device i is for deploying operator j; the lower the value, the worse the adaptability.
[0110] k represents different resource dimensions. The parameters of each dimension are normalized using the mean and variance of their own dimension to avoid calculation biases caused by dimensional differences.
[0111] is the sensitivity adjustment factor, which is used to control the influence intensity of the difference in the k th resource dimension on the matching degree.
[0112] is the difference value between the resource requirement characteristics and the device capability characteristics.
[0113] is the difference value between the operator requirements and the device capabilities in dimension k on.
[0114] The parameters of each dimension are normalized using the mean and variance of their own dimension to avoid calculation biases caused by dimensional differences.
[0115] The target matching value is calculated and updated dynamically in real time to drive the generation of the deployment strategy. For example, when the video memory of a certain GPU drops from 25GB to 10GB, the normalized difference in the video memory dimension changes from -1.25 to +2.5, and the matching degree drops sharply from 0.8 to 0.05. The system immediately marks the device as a low matching level.
[0116] Step 408: Determine the deployment plan corresponding to the target model according to the target matching value and the preset classification strategy.
[0117] As can be seen from the above technical solutions, the beneficial effects of this embodiment are as follows: The mathematical model not only quantifies the global optimality of resource adaptation, but also ensures elasticity and robustness in heterogeneous environments through a dynamic response mechanism.
[0118] As Figure 5 shown, it is a specific embodiment of a model deployment device combined with hardware deployment in the present application. A model deployment device combined with hardware deployment in this embodiment, that is, an entity device for executing Figures 1 - 4 the model deployment method combined with hardware deployment. Its technical solution is essentially the same as that of the above embodiment, and the corresponding descriptions in the above embodiment also apply to this embodiment. A model deployment device combined with hardware deployment in this embodiment includes: A hardware topology diagram determination module 501, configured to construct a hardware topology diagram corresponding to the device node according to the device nodes in the target model deployment environment and the corresponding connection relationships of each device node; A model calculation diagram determination module 502, configured to analyze the dependency relationship of the target operator in the target model to generate a model calculation diagram corresponding to the target operator; A target matching value determination module 503, configured to calculate the target matching value between the device node and the target operator according to the hardware topology diagram and the model calculation diagram by using a preset evaluation algorithm; A deployment plan determination module 504, configured to determine the deployment plan corresponding to the target model according to the target matching value and the preset classification strategy.
[0119] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0120] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a bidirectional arrow is used in Figure 6 , but it does not mean that there is only one bus or one type of bus.
[0121] Memory, which is used to store executable instructions. Specifically, executable instructions are computer programs that can be executed. The memory can include a memory and a non-volatile memory, and provide executable instructions and data to the processor.
[0122] In a possible implementation, the processor reads the corresponding executable instructions from the non-volatile memory into the memory and then runs them, or can also obtain the corresponding executable instructions from other devices to form a model deployment device that combines hardware deployment at the logical level. The processor executes the executable instructions stored in the memory to implement a model deployment method that combines hardware deployment provided in any embodiment of this application through the executed executable instructions.
[0123] As described above in this application Figure 5 The method executed by a model deployment device that combines hardware deployment provided in the embodiments shown can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0124] The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0125] The embodiments of the present application also propose a readable medium. When the execution instructions stored in the readable storage medium are executed by the processor of the electronic device, the electronic device can execute a model deployment method combined with hardware deployment provided in any embodiment of the present application, and is specifically used to execute as Figure 1 or Figure 2 or Figure 3 or Figure 4 the method shown.
[0126] The electronic device in each of the foregoing embodiments may be a computer.
[0127] Those skilled in the art should understand that the embodiments of the present application may be provided as a method or a computer program product. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or a combination of software and hardware.
[0128] The embodiments in the present application are all described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can be referred to the partial description of the method embodiments.
[0129] It should also be noted that the term "including", "comprising" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element.
[0130] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A model deployment method combined with hardware deployment, characterized in that, Including: Construct a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each of the device nodes; Analyze the dependency relationships of the target operators in the target model to generate a model computation graph corresponding to the target operators; According to the hardware topology graph and the model computation graph, use a preset evaluation algorithm to calculate a target matching value between the device nodes and the target operators; Determine a deployment plan corresponding to the target model according to the target matching value and a preset grading strategy.
2. The method according to claim 1, wherein The constructing a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the corresponding connection relationships of each of the device nodes includes: Determine the device information corresponding to each of the device nodes; Determine the communication cost corresponding to the device nodes according to the device information and the connection relationships; Construct the hardware topology graph with the device nodes as vertices and the communication cost as edge weights.
3. The method according to claim 1, wherein The analyzing the dependency relationships of the target operators in the target model and the resource requirements corresponding to each of the target operators to generate a model computation graph corresponding to the target operators includes: According to the computation process corresponding to the target model, analyze the dependency relationships between the target operators to generate data flow edges corresponding to the target operators; Construct the model computation graph with the target operators as vertices and the data flow edges as directed edges.
4. The method according to claim 1, characterized in that, The calculating a target matching value between the device nodes and the target operators according to the hardware topology graph and the model computation graph, using a preset evaluation algorithm includes: Determine the resource requirement characteristics corresponding to each of the target operators according to the model computation graph; Determine the device capability characteristics corresponding to each of the device nodes according to the hardware topology graph; Calculate the difference value between the resource requirement characteristics and the device capability characteristics; Input the difference value into a preset mathematical model so that the mathematical model outputs the target matching value.
5. The method according to claim 4, characterized in that, The determining a deployment plan corresponding to the target model according to the target matching value and a preset grading strategy includes: Determine the target matching level corresponding to the target matching value according to the grading threshold corresponding to the grading strategy; the target matching levels include a high matching level, a low matching level, and a medium matching level; Determine the deployment plan based on the target matching level.
6. The method according to claim 5, characterized in that, The determining the deployment plan based on the target matching level includes: When the target matching value belongs to the high matching level, perform a priority sorting on the target operators based on the numerical magnitude of the target matching value to generate a deployment queue corresponding to the target operators; Deploy the target operators to the corresponding device nodes according to the deployment queue; And / or When the target matching value belongs to the low matching level, determine the original sub-graph corresponding to the target operators in the model computation graph; Use a computation path recombination algorithm to perform path recombination on the original sub-graph to generate an optimized sub-graph; Replace the original sub-graph with the optimized sub-graph to generate a corresponding optimized computation graph; Calculate a first matching value corresponding to the optimized computation graph; When the first matching value does not belong to the low matching level, update the model computation graph by using the optimized sub-graph.
7. The method according to claim 5, wherein The determining of the deployment plan based on the target matching level includes: When the target matching value belongs to the medium matching level, parse the resource requirement characteristics to determine the bottleneck dimension characteristics that cause the target matching value to be at the medium matching level; According to the bottleneck dimension characteristics, adjust the connection relationship from the initial connection to an optimized connection; Determine the second matching value corresponding to the optimized connection; When the second matching value belongs to the high matching level, update the connection relationship in the hardware topology graph by using the optimized connection.
8. A model deployment device combined with hardware deployment, characterized in that, Including: A hardware topology graph determination module configured to construct a hardware topology graph corresponding to the device nodes according to the device nodes in the target model deployment environment and the connection relationships corresponding to the device nodes; A model computation graph determination module configured to parse the dependency relationship of the target operator in the target model to generate a model computation graph corresponding to the target operator; A target matching value determination module configured to calculate a target matching value between the device nodes and the target operator by using a preset evaluation algorithm according to the hardware topology graph and the model computation graph; A deployment plan determination module configured to determine a deployment plan corresponding to the target model according to the target matching value and a preset grading strategy.
9. A computer-readable storage medium storing a computer program, characterized in that, The computer program is used to execute a model deployment method combined with hardware deployment according to any one of claims 1-7 above.
10. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement a model deployment method combined with hardware deployment according to any one of claims 1-7 above.
Citation Information
Patent Citations
Model deployment method and device, model operation method and device, offline analysis tool and electronic equipment
CN115525436A
Neural network model deployment method, system and device and storage medium
CN115796041A
Deep learning model deployment method and device based on search
CN116306856A
Deep learning model terminal optimization deployment method based on storage and calculation integration
CN117436546A
Model deployment method and device, electronic equipment and computer readable medium
CN118394717A
Cited By
Method and device for determining equipment
CN120849223A
Large language model pairing method and device based on hardware and model features, terminal, medium and product
CN121009379A
Method and apparatus for pairing large language model based on hardware and model characteristics, terminal, medium and product
CN121009379B