AIOT board card-oriented inference model seamless migration and adaptive optimization method

Through the combination of hardware performance analysis and deep learning models, the transformation strategy of the inference model and real-time adjustment of parameters is optimized, the problem of limited AIOT hardware resources is solved, seamless migration and adaptive optimization of the inference model are achieved, and the efficiency and performance of the model are improved.

CN119987998AInactive Publication Date: 2025-05-13CHENGDU SHUXI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411885470.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to achieve seamless migration and adaptive optimization of inference models on limited AIOT hardware resources, and cannot fully utilize hardware performance and meet the needs of real-time data processing.

Method used

The computing power of AIOT devices is evaluated through hardware performance analysis tools, the computational volume, memory requirements and complexity of the original inference model are layered, the conversion strategy of the model is optimized, the seamless migration of the model is achieved, and the accuracy, batch size and execution path of the inference are dynamically adjusted using deep learning models and real-time load monitoring.

Benefits of technology

The efficient operation of the inference model on AIOT hardware is achieved, the model migration and optimization problems under resource constraints are solved, and the efficiency and performance of the model under limited resources are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987998A_ABST
    Figure CN119987998A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of inference model migration and optimization, and relates to an AIOT board card-oriented inference model seamless migration and adaptive optimization method. A hardware performance analysis tool is used for comprehensively evaluating the computing power of AIOT equipment to obtain hardware performance indexes, an original reasoning model is layered, the analysis tool is used for evaluating performance data of each layer, a conversion strategy of the model is optimized on the basis of the hardware performance indexes and model performance data, and the conversion strategy of the AIOT equipment is optimized. Hardware limitation and model performance requirements can be considered in the conversion process of the model. Converting the original reasoning model based on the determined conversion strategy, and deploying the original reasoning model to the AIOT board card, thereby achieving the purpose of seamless migration; and secondly, a deep learning model and real-time load monitoring are adopted to dynamically adjust the reasoning precision, the batch size and the execution path, so that the reasoning model can be adaptively optimized according to the actual operation condition, and the efficiency and the performance of the model under limited resources are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of inference model migration and optimization, and more specifically, to a method for seamless migration and adaptive optimization of inference models for AIOT boards. Background Art

[0002] With the rapid development of Internet of Things (IoT) technology, AIOT (AI+IoT) devices are increasingly used in smart homes, smart industries, smart cities and other fields. These devices usually need to integrate artificial intelligence reasoning models to achieve real-time data processing and decision-making. However, the hardware resources of AIOT devices are relatively limited, including processing power, memory bandwidth, storage space and power consumption, which puts higher requirements on the deployment and operation of reasoning models.

[0003] Existing technologies usually include methods such as compressing, quantizing, and optimizing inference models to adapt to the hardware limitations of AIOT devices. For example, the parameters and computational complexity of the model are reduced through technologies such as model pruning, weight quantization, and knowledge distillation, thereby reducing the model's demand for hardware resources. However, these methods often fail to achieve seamless migration and adaptive optimization of inference models, that is, existing technologies often ignore the matching problem between the hardware characteristics of AIOT devices and the performance of inference models. First, the hardware performance indicators of AIOT devices (such as the processing power of the computing unit, memory bandwidth, latency, etc.) have a direct impact on the operating efficiency of the model. Secondly, the original inference model has different computational complexity, memory requirements, and complexity at each layer, which requires optimization for specific hardware. Therefore, how to achieve seamless migration and adaptive optimization of inference models on limited AIOT hardware resources to fully utilize hardware performance and meet the needs of real-time data processing is a technical problem that needs to be solved at present. Summary of the invention

[0004] The present invention provides a method for seamless migration and adaptive optimization of reasoning models for AIOT boards, which aims to solve the technical problem of achieving seamless migration and adaptive optimization of reasoning models on limited AIOT hardware resources.

[0005] The seamless migration and adaptive optimization method of the inference model for AIOT boards includes the following steps:

[0006] Step 1: Evaluate the computing power of the target AIOT device based on the hardware performance analysis tool, clarify the processing power, memory bandwidth, latency, storage space, and power consumption of the computing unit of the AIOT device, and obtain the performance indicators of each computing unit;

[0007] Step 2: Layer the original inference model and use analytical tools to evaluate the computational workload, memory requirements, and complexity of each layer of the original inference model.

[0008] Step 3: Optimize the model conversion strategy based on the performance indicators of each computing unit and the computational workload, memory requirements, and complexity of each layer of the original inference model to obtain the conversion strategy of the original inference model;

[0009] Step 4: Convert the original reasoning model based on the determined conversion strategy to obtain a converted reasoning model, and deploy the converted reasoning model to the AIOT board;

[0010] Step 5: Use the deep learning model in combination with the actual load of the device to dynamically adjust the inference accuracy, batch size, and execution path.

[0011] The present invention uses a hardware performance analysis tool to comprehensively evaluate the computing power of the AIOT device, ensuring that the hardware performance indicators of the device, including processing power, memory bandwidth, latency, storage space and power consumption, etc., can be accurately understood, thereby realizing the basis for matching the model with the hardware; then, the original reasoning model is layered, and the analysis tool is used to evaluate the computing amount, memory requirement and complexity of each layer, providing detailed model performance data for subsequent optimization; then, the conversion strategy of the model is optimized based on the hardware performance indicators and the model performance data, ensuring that the model can take into account the hardware limitations and model performance requirements during the conversion process; based on the determined conversion strategy, the original reasoning model is converted to obtain a reasoning model that is both hardware-adaptive and performance-maintaining, and is deployed on the AIOT board, thereby achieving the purpose of seamless migration; secondly, a deep learning model and real-time load monitoring are used to dynamically adjust the reasoning accuracy, batch size and execution path, so that the reasoning model can be adaptively optimized according to the actual operation conditions, further improving the efficiency and performance of the model under limited resources; therefore, the present invention ensures the efficient operation of the reasoning model on the AIOT hardware, and solves the model migration and optimization problems under resource-constrained conditions.

[0012] Preferably, the original inference model in step 2 is preliminarily layered into an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer, wherein the computational complexity and memory requirements of each layer are obtained according to the following steps:

[0013] Input layer: The computational complexity of the input layer is O(H×W×C), where H represents the height of the input feature, W represents the width, and C represents the number of channels. The memory requirement of the input layer is H×W×C×B, where B represents the batch size.

[0014] Convolutional Layer:

[0015]

[0016] Where: and Respectively represent the height and width of the convolution kernel; Represents the depth of the output feature map of the convolutional layer, that is, the number of convolution kernels; Represents the height of the output feature map of the convolutional layer; Represents the width of the output feature map of the convolutional layer; Indicates the number of channels of the input image of the convolutional layer; Represents the computational effort of the convolutional layer;

[0017] The memory requirement of the convolution layer is the memory requirement of each convolution kernel. Plus the memory requirements of the output feature map

[0018] Pooling layer:

[0019]

[0020] Where: Represents the height of the feature map output by the pooling layer; Indicates the width of the feature map output by the pooling layer; Indicates the number of channels of the input image of the pooling layer; Indicates the window size of the pooling layer; Represents the computational effort of the pooling layer;

[0021] The memory requirement of the pooling layer is

[0022] Fully connected layer:

[0023]

[0024] Where: Represents the input vector dimension of the fully connected layer; Represents the output vector dimension of the fully connected layer; Represents the computational effort of the fully connected layer;

[0025] The memory requirement of the full connection is the sum of the size of the weight matrix and the size of the output vector

[0026] Output layer:

[0027]

[0028] Where: Represents the output dimension of the fully connected layer in the output layer; Indicates the number of categories; Represents the computational complexity of softmax, which is the exponential calculation for each category;

[0029] The memory requirement of the output layer is the sum of the memory requirement of the fully connected layer in the output layer and the memory requirement of the softmax, that is,

[0030] Preferably, the complexity of each layer of the original inference model is calculated based on the number of floating-point operations and memory bandwidth of each layer:

[0031]

[0032] Where: C i represents the complexity of the i-th layer; F ops represents the floating point operation number of the i-th layer; T model M represents the processing power of the inference model executed on each computing unit, which is the time required for each floating-point operation; i represents the memory requirement of the i-th layer; B model Indicates the amount of data that can be transmitted per second.

[0033] Preferably, step 3 comprises the following steps:

[0034] Initialization: Initialize the computing unit of each AIOT board, randomly assign each inference model layer to the computing unit, and obtain the initial deployment plan;

[0035] Layer scheduling: For each inference model layer, the latency and memory bottleneck of each inference model layer on the current computing unit are calculated, and the layer allocation scheme is dynamically adjusted according to the remaining load of each computing unit;

[0036] Optimization target calculation: Optimization calculation target function for each level:

[0037] Total_Cost i =w delay Delay i +w memory Memory_Load i +w power Power_Load i ;

[0038] Where: Delay i Represents the inference model layer L i In the calculation unit H j Delay on Memory_Load i Represents the inference model layer L i In the calculation unit H j Memory bottleneck on Power_Load i Represents the computing unit H j Run the L i The power consumption of the layer model inference layer; w delay 、w memory and w powerRepresents the weight coefficient; Total_Cost i Represents the total cost, and the optimization goal is to minimize the total cost;

[0039] Quantization and pruning optimization judgment: judge whether quantization optimization or pruning optimization is needed based on the set objective function threshold and constraint conditions;

[0040] The amount of calculation after quantization optimization becomes:

[0041]

[0042] Where: b quant Indicates the number of quantization bits; F i represents the computational amount of the i-th layer; Indicates the amount of quantized computation;

[0043] The memory requirements are:

[0044]

[0045] Where: M i Indicates the memory requirement of the i-th layer; Indicates the quantized memory requirements;

[0046] The pruning optimization analyzes the computational graph structure of each layer and uses a structured pruning method to remove redundant connections and neurons. The amount of computation after pruning is as follows:

[0047]

[0048] Where: pruning_ratio represents the pruning ratio;

[0049] Accuracy evaluation: During the quantization and pruning process, the accuracy change of each layer is evaluated in real time to ensure that the accuracy loss does not exceed the preset threshold;

[0050] A genetic algorithm is used to dynamically adjust the hierarchical allocation and optimization strategy of each computing unit through selection, crossover, and mutation operations to minimize the overall objective function and output the final conversion strategy, where the conversion strategy includes the inference model layer allocated to each processing unit on the AIOT board and the pruning and / or quantization operations required for each inference model layer.

[0051] Preferably, the constraints include processor performance constraints, memory limit constraints, power consumption limit constraints, delay requirement constraints, and precision loss limit constraints.

[0052] The processor performance constraint is that each inference model layer, when executed on the computing unit, does not exceed the difference between the maximum processing capacity of the board and the preset redundant processing capacity;

[0053] The memory limit constraint is that the memory requirement required for each inference model layer does not exceed the difference between the total board memory and the preset memory redundancy requirement;

[0054] The power consumption limit constraint is that the power consumption required for each inference model layer does not exceed the difference between the total power consumption budget of the board and the preset power consumption redundancy requirement;

[0055] The delay requirement constraint is that the total delay of the inference model does not exceed a preset maximum acceptable delay;

[0056] The precision loss limit constraint is that the precision loss of the model after quantization or pruning does not exceed a preset precision loss threshold.

[0057] Preferably, the deep learning model includes an input layer, a shared hidden layer, and a task-specific output layer;

[0058] The input layer is used to input the load data of the device and the characteristics of the reasoning task;

[0059] The shared hidden layer receives the features of the input layer and uses a deep neural network to map the input features to a higher-dimensional hidden layer representation, where the hidden layer structure is as follows:

[0060] H1=ReLU(W1·X1+b1);

[0061] H2=ReLU(W2·X2+b2);

[0062] H3=ReLU(W3·X3+b3);

[0063] Where: W i and b i represents the weight and bias of the i-th layer, i is an integer from 1 to 3; X represents the input feature; X1 represents the input feature of the first hidden layer; X2 represents the input feature of the second hidden layer; X3 represents the input feature of the third hidden layer; ReLU represents the activation function; H1 represents the output of the first hidden layer; H2 represents the output of the second hidden layer; H3 represents the output of the third hidden layer;

[0064] The task-specific output layer is used to receive the output features of the shared hidden layer, wherein the task-specific output layer includes an accuracy prediction unit, a batch size prediction unit, and an execution path prediction unit;

[0065] The accuracy prediction unit is a neuron that predicts the accuracy level of the inference model based on the output features of the shared hidden layer:

[0066] Precision = Softmax(W prec H3+b prec );

[0067] Where: W prec Represents the weight matrix of accuracy prediction; b prec It represents the bias term of precision prediction; Precision represents the result of precision prediction;

[0068] The batch size prediction unit is a single neuron that predicts the batch size of the inference model based on the output features of the shared hidden layer:

[0069] Batch Size = Round (W batch H3+b batch );

[0070] Where: Batch Size represents the prediction result of the batch size of the inference model; Round represents the integer function; W batch b represents the weight matrix for batch size prediction; batch Represents the bias term for batch size prediction, with dimension 1;

[0071] The execution path prediction unit is a multi-classification neuron that predicts the appropriate hardware execution path of the inference model based on the output features of the shared hidden layer:

[0072] Path=Softmax(W path H3+b path );

[0073] Where: Path represents the execution path prediction result of the inference model; W path Represents the weight matrix of path prediction; b path Represents the bias term for path prediction.

[0074] Preferably, the loss function of the deep learning model is as follows:

[0075] L total =λ prec L prec +λ batch L batch +λ path L path ;

[0076] Where: L total represents the total loss function; λ prec , batch and path Represents the weight coefficient; L prec Represents the accuracy prediction loss, using the cross entropy loss function; L batch represents the loss of batch size prediction, using the mean square error loss function; L path Represents the loss of execution path prediction, using the cross entropy loss function.

[0077] Preferably, an adaptive weighting mechanism is introduced into the deep learning model to weight the output H3 of the shared hidden layer:

[0078] H task,i =A task,i H3;

[0079] Where: H task,i A represents the weighted output for the i-th prediction task, which is the input of the i-th prediction unit; task,i Represents the weight coefficient of the i-th prediction task, which is learned through training.

[0080] The beneficial effects of the present invention include:

[0081] The present invention uses a hardware performance analysis tool to comprehensively evaluate the computing power of the AIOT device, ensuring that the hardware performance indicators of the device, including processing power, memory bandwidth, latency, storage space and power consumption, etc., can be accurately understood, thereby realizing the basis for matching the model with the hardware; then, the original reasoning model is layered, and the analysis tool is used to evaluate the computing amount, memory requirement and complexity of each layer, providing detailed model performance data for subsequent optimization; then, the conversion strategy of the model is optimized based on the hardware performance indicators and the model performance data, ensuring that the model can take into account the hardware limitations and model performance requirements during the conversion process; based on the determined conversion strategy, the original reasoning model is converted to obtain a reasoning model that is both hardware-adaptive and performance-maintaining, and is deployed on the AIOT board, thereby achieving the purpose of seamless migration; secondly, a deep learning model and real-time load monitoring are used to dynamically adjust the reasoning accuracy, batch size and execution path, so that the reasoning model can be adaptively optimized according to the actual operation conditions, further improving the efficiency and performance of the model under limited resources; therefore, the present invention ensures the efficient operation of the reasoning model on the AIOT hardware, and solves the model migration and optimization problems under resource-constrained conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0083] Figure 1 An overall step block diagram provided for an embodiment of the present invention.

[0084] Figure 2 A deep learning model framework diagram provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0085] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0086] See also Figure 1 As shown in the figure, the seamless migration and adaptive optimization method of the inference model for the AIOT board includes the following steps:

[0087] Step 1: Evaluate the computing power of the target AIOT device based on the hardware performance analysis tool, clarify the processing power, memory bandwidth, latency, storage space, and power consumption of the computing unit of the AIOT device, and obtain the performance indicators of each computing unit;

[0088] Step 2: Layer the original inference model and use analytical tools to evaluate the computational workload, memory requirements, and complexity of each layer of the original inference model.

[0089] As a possible implementation of this embodiment, the original reasoning model in step 2 is preliminarily layered into an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer, wherein the computational complexity and memory requirements of each layer are obtained according to the following steps:

[0090] Input layer: The computational complexity of the input layer is O(H×W×C), where H represents the height of the input feature, W represents the width, and C represents the number of channels. The memory requirement of the input layer is H×W×C×B, where B represents the batch size.

[0091] Convolutional Layer:

[0092]

[0093] Where: and Respectively represent the height and width of the convolution kernel; Represents the depth of the output feature map of the convolutional layer, that is, the number of convolution kernels; Represents the height of the output feature map of the convolutional layer; Represents the width of the output feature map of the convolutional layer; Indicates the number of channels of the input image of the convolutional layer; Represents the computational effort of the convolutional layer;

[0094] in:

[0095]

[0096] Where: Represents the height of the input feature map in the convolution layer; S represents the stride, that is, the number of pixels that the convolution kernel slides on the input feature map each time; P represents padding, that is, the number of layers of zero padding added to the edge of the input feature map; Represents the width of the input feature map in the convolutional layer;

[0097] The memory requirement of the convolution layer is the memory requirement of each convolution kernel. Plus the memory requirements of the output feature map

[0098] Pooling layer:

[0099]

[0100] Where: Represents the height of the feature map output by the pooling layer; Indicates the width of the feature map output by the pooling layer; Indicates the number of channels of the input image of the pooling layer; Indicates the window size of the pooling layer; Represents the computational effort of the pooling layer;

[0101] in:

[0102]

[0103] Where: Represents the height of the feature map input to the pooling layer; Represents the width of the feature map input to the pooling layer; S represents the stride, that is, the step length of the window movement in the pooling operation;

[0104] The memory requirement of the pooling layer is

[0105] Fully connected layer:

[0106]

[0107] Where: Represents the input vector dimension of the fully connected layer; Represents the output vector dimension of the fully connected layer; Represents the computational effort of the fully connected layer;

[0108] The memory requirement of the full connection is the sum of the size of the weight matrix and the size of the output vector

[0109] The calculation formula for the above fully connected computation and memory requirements is as follows:

[0110] Since the input vector x is a The row vector of the weight matrix W is The output vector y is a The actual memory requirement is therefore composed of the following components:

[0111] Weight matrix: dimension is The memory required is

[0112] Bias vector: Each output node has a bias term, and the dimension of the bias vector is The memory requirement is therefore

[0113] Output vector: The dimension of the output vector is The memory requirement is

[0114] Based on the above, the actual memory requirement is Since the output vector is the result of the calculation, not the part stored in the model parameters, the calculation memory requirement refers to the model parameters and persistent data, not the intermediate calculation results, so the memory requirement of the output vector is not included;

[0115] Output layer:

[0116]

[0117] Where: Represents the output dimension of the fully connected layer in the output layer; Indicates the number of categories; Represents the computational complexity of softmax, which is the exponential calculation for each category;

[0118] The memory requirement of the output layer is the sum of the memory requirement of the fully connected layer in the output layer and the memory requirement of the softmax, that is,

[0119] The complexity is calculated based on the number of floating point operations and memory bandwidth of each layer:

[0120]

[0121] Where: C i represents the complexity of the i-th layer; F ops represents the floating point operation number of the i-th layer; T model M represents the processing power of the inference model executed on each computing unit, which is the time required for each floating-point operation; i represents the memory requirement of the i-th layer; B model Indicates the amount of data that can be transmitted per second.

[0122] Step 3: Optimize the model conversion strategy based on the performance indicators of each computing unit and the computational workload, memory requirements, and complexity of each layer of the original inference model to obtain the conversion strategy of the original inference model;

[0123] As a possible implementation of this embodiment, step 3 includes the following steps:

[0124] Initialization: Initialize the computing unit of each AIOT board, randomly assign each inference model layer to the computing unit, and obtain the initial deployment plan;

[0125] Layer scheduling: For each inference model layer, the latency and memory bottleneck of each inference model layer on the current computing unit are calculated, and the layer allocation scheme is dynamically adjusted according to the remaining load of each computing unit;

[0126] Optimization target calculation: Optimization calculation target function for each level:

[0127] Total_Cost i =w delay Delay i +w memory Memory_Load i +w power Power_Load i ;

[0128] Where: Delay i Represents the inference model layer L i In the calculation unit H j Delay on Memory_Load i Represents the inference model layer L i In the calculation unit H j Memory bottleneck on Power_Load i Represents the computing unit H j Run the L i The power consumption of the layer model inference layer; w delay 、w memory and w power Represents the weight coefficient; Total_Cost i Represents the total cost, and the optimization goal is to minimize the total cost;

[0129] Quantization and pruning optimization judgment: judge whether quantization optimization or pruning optimization is needed based on the set objective function threshold and constraint conditions;

[0130] The amount of calculation after quantization optimization becomes:

[0131]

[0132] Where: b quant Indicates the number of quantization bits (such as 8 bits, 4 bits); F i represents the computational amount of the i-th layer; Indicates the amount of quantized computation;

[0133] The memory requirements are:

[0134]

[0135] Where: M i Indicates the memory requirement of the i-th layer; Indicates the quantized memory requirements;

[0136] The pruning optimization analyzes the computational graph structure of each layer and uses a structured pruning method to remove redundant connections and neurons. The amount of computation after pruning is as follows:

[0137]

[0138] Where: pruning_ratio represents the pruning ratio;

[0139] Accuracy evaluation: During the quantization and pruning process, the accuracy change of each layer is evaluated in real time to ensure that the accuracy loss does not exceed the preset threshold;

[0140] A genetic algorithm is used to dynamically adjust the hierarchical allocation and optimization strategy of each computing unit through selection, crossover, and mutation operations to minimize the overall objective function and output the final conversion strategy, where the conversion strategy includes the inference model layer allocated to each computing unit on the AIOT board and the pruning and / or quantization operations required for each inference model layer.

[0141] Regarding the description of computing units, there are multiple computing units on the AIOT board. For example, there may be multiple CPUs and multiple GPUs. We collectively refer to them as computing units.

[0142] The constraints include processor performance constraints, memory limit constraints, power consumption limit constraints, delay requirement constraints, and precision loss limit constraints.

[0143] The processor performance constraint is that each inference model layer, when executed on the computing unit, does not exceed the difference between the maximum processing capacity of the board and the preset redundant processing capacity;

[0144] The memory limit constraint is that the memory requirement required for each inference model layer does not exceed the difference between the total board memory and the preset memory redundancy requirement;

[0145] The power consumption limit constraint is that the power consumption required for each inference model layer does not exceed the difference between the total power consumption budget of the board and the preset power consumption redundancy requirement;

[0146] The delay requirement constraint is that the total delay of the inference model does not exceed a preset maximum acceptable delay;

[0147] The precision loss limit constraint is that the precision loss of the model after quantization or pruning does not exceed a preset precision loss threshold.

[0148] In this embodiment, by dynamically adjusting the allocation of inference models and computing units, hardware resources are effectively utilized and resource utilization is improved. Secondly, through optimization such as quantization and pruning, the memory requirements and power consumption of the model are reduced, and the energy efficiency of the system is improved. We automatically find the optimal model conversion strategy through genetic algorithms, realize automatic conversion strategy generation, save human resources and time costs, and can perform optimization while ensuring accuracy, ensuring the accuracy of the inference results.

[0149] The specific steps of using the genetic algorithm to implement the above technical solution are as follows:

[0150] Initialize the basic elements of the genetic algorithm:

[0151] Individual coding: Each individual represents the hierarchical allocation scheme, quantization and pruning operations of an inference model. It can be represented by binary coding or integer coding. The computing unit allocation of each inference model layer can be represented by an integer. For example, assuming there are 3 computing units, the computing unit allocation of each layer can be represented by numbers such as 0, 1, and 2. In addition, the quantization (number of bits) and pruning ratio of each layer can also be represented by integer coding. For example, the number of quantization bits can be represented by an integer (such as 8 for 8-bit quantization), and the pruning ratio can be represented by a real number (such as 0.2 for 20% pruning).

[0152] Genome length: Assuming there is an N-layer model, each layer has two parameters: computing unit allocation and optimization operation (quantization / pruning). Then the gene length of each individual is 2*N (two parameters per layer).

[0153] Initialize the population: Randomly generate a population, assuming that the population size is pop_size, and each individual (solution) represents a set of inference model level allocation and optimization strategies.

[0154] In the initial population, each individual is randomly assigned the computing unit of each model layer, and the number of quantization bits and pruning ratio of each layer are randomly set.

[0155] Calculation of fitness function for each individual: Calculate the fitness value of each individual based on the objective function;

[0156] Quantization and pruning optimization: Based on the calculated fitness value of each individual, it is compared with the set fitness threshold (objective function threshold). If the fitness value of the individual is higher than the fitness threshold and meets the constraints defined above, the calculation amount after quantization and pruning of the individual is calculated based on the above formula;

[0157] Precision evaluation: After each optimization, the precision loss of each layer of the inference model is evaluated in real time to ensure that the precision loss does not exceed the preset threshold;

[0158] Selection: Use the roulette wheel selection strategy to select individuals with higher fitness from the current population as parents (the first K individuals with the highest fitness values ​​can be used as parents, where K is a preset value);

[0159] Crossover: Use a single-point crossover strategy to crossover the selected parent individuals to generate new offspring individuals; for example, perform a crossover operation on the allocation of computing units or the selection of quantization / pruning;

[0160] For example, each individual (i.e., the hierarchical allocation scheme of each inference model) consists of multiple genes, each of which represents a specific decision, such as:

[0161] Compute unit allocation: For each inference model layer, select which compute unit to execute the layer.

[0162] Quantization Bits: Select the number of quantization bits for each inference model layer, such as 8 bits or 4 bits.

[0163] Pruning ratio: Select the pruning ratio for each inference model layer, for example, pruning 20% ​​of the connections.

[0164] By selecting a crossover point as the split point, gene crossover is performed. If the two parent individuals represent the following hardware sub-unit allocation and quantization bit number respectively:

[0165] Parent individual A: [0,1,2,0] (indicating that layer 0 is assigned to computing unit 0, layer 1 is assigned to computing unit 1, layer 2 is assigned to computing unit 2, and layer 3 is assigned to computing unit 0) and [8,8,8,8] (indicating that all layers use 8-bit quantization) Parent individual B: [2,0,1,2] (indicating that layer 0 is assigned to computing unit 2, layer 1 is assigned to computing unit 0, layer 2 is assigned to computing unit 1, and layer 3 is assigned to computing unit 2) and [4,4,4,4] (indicating that all layers use 4-bit quantization);

[0166] If the crossover point is chosen after the second gene, the crossover operation will produce the following two offspring individuals:

[0167] Offspring individual C: [0,1,2,0] (copied from parent A) and [4,4,4,4] (copied from parent B) Offspring individual D: [2,0,1,2] (copied from parent B) and [8,8,8,8] (copied from parent A)

[0168] In this way, the offspring individuals C and D combine the computing unit allocation and quantization bit characteristics of the parent individuals A and B. The choice of crossover point will affect the feature combination of the offspring individuals, thereby affecting the search process of the genetic algorithm and the quality of the final solution.

[0169] Mutation: Perform mutation operations on the offspring individuals generated by crossover, randomly change the computing unit allocation, quantization bit number or pruning ratio of certain inference model layers, and the mutation probability can be set to a lower value to avoid over-expansion of the search space;

[0170] Elite retention: In order to ensure that the excellent individuals of each generation are not lost, a predetermined proportion of excellent individuals are retained in each generation and directly enter the next generation;

[0171] The above constraints need to be considered in the selection, crossover and mutation processes. If the constraints are violated, a lower fitness value is given. Otherwise, the objective function value is calculated normally.

[0172] Based on fitness function calculation, quantization and pruning, accuracy evaluation, selection, crossover, mutation and elite retention, continuous iteration is performed until the preset number of iterations is reached, and finally the individual with the lowest fitness function value is obtained, which is used as the optimal conversion strategy.

[0173] As a further implementation of this embodiment, in the selection of the crossover point of the genetic algorithm, we introduced a performance-aware crossover point selection strategy based on the application scenario, and dynamically adjusted the crossover point by considering the performance of the computing unit and the computing requirements of the inference model layer, so that the crossover operation is more suitable for the application scenario and the search efficiency is improved. This strategy can dynamically adjust the gene exchange method during the crossover process according to the performance indicators of the computing unit and the computing requirements of the inference model layer, so as to find a better inference model conversion strategy; the details are as follows:

[0174] corss point =α*match degree *(1-performance surplus )+β*computation urgency ;

[0175] Where: corss pointIndicates the position of the intersection, and its value range is [0, gene length], where the gene length is equal to the number of inference model layers multiplied by the number of parameters in each layer (for example: computing unit allocation and quantization bit number); α and β are weight coefficients used to balance the impact of matching degree, performance margin and computing demand urgency; match degree Indicates the matching degree between performance indicators and computing requirements. It is the dot product between the performance indicator vector and the computing requirement vector divided by the product of their Euclidean distances. The value range is [0,1]. surplus Indicates the performance margin of the computing unit, which is the ratio of the difference between the available resources of the computing unit and its maximum resources to the maximum resources, and the value range is [0,1]; urgency Indicates the urgency of the computational requirements of the inference model layer. It is the ratio of the computational requirements of the inference model layer to the sum of the computational requirements of all layers, and its value range is [0,1].

[0176] Based on the above method, the intersection point of computing units and inference model layers with higher matching degree will be further back to retain more genetic information. The intersection point of computing units with larger performance margin will also be further back to make full use of their performance. The intersection point of inference model layers with higher computing demand urgency will be further forward to ensure that their performance requirements are met.

[0177] Step 4: Convert the original reasoning model based on the determined conversion strategy to obtain a converted reasoning model, and deploy the converted reasoning model to the AIOT board;

[0178] Step 5: Use the deep learning model in combination with the actual load of the device to dynamically adjust the inference accuracy, batch size, and execution path.

[0179] See also Figure 2 As shown, as a possible implementation of this embodiment, the deep learning model includes an input layer, a shared hidden layer, and a task-specific output layer;

[0180] The input layer is used to input the load data of the device and the characteristics of the reasoning task; the load data includes CPU load, GPU load, memory occupancy, bandwidth usage and power consumption; the characteristics of the reasoning task include the complexity of the model (measured by the computational amount of the model), the size of the data set and the target accuracy level.

[0181] The shared hidden layer receives the features of the input layer and uses a deep neural network to map the input features to a higher-dimensional hidden layer representation, where the hidden layer structure is as follows:

[0182] H1=ReLU(W1·X1+b1);

[0183] H2=ReLU(W2·X2+b2);

[0184] H3=ReLU(W3·X3+b3);

[0185] Where: W i and b i represents the weight and bias of the i-th layer, i is an integer from 1 to 3; X represents the input feature; X1 represents the input feature of the first hidden layer; X2 represents the input feature of the second hidden layer; X3 represents the input feature of the third hidden layer; ReLU represents the activation function; H1 represents the output of the first hidden layer; H2 represents the output of the second hidden layer; H3 represents the output of the third hidden layer;

[0186] The task-specific output layer is used to receive the output features of the shared hidden layer, wherein the task-specific output layer includes an accuracy prediction unit, a batch size prediction unit, and an execution path prediction unit;

[0187] The accuracy prediction unit is a neuron that predicts the accuracy level of the inference model based on the output features of the shared hidden layer:

[0188] Precision = Softmax(W prec H3+b prec );

[0189] Where: W prec Represents the weight matrix of accuracy prediction; b prec It represents the bias term of precision prediction; Precision represents the result of precision prediction;

[0190] The batch size prediction unit is a single neuron that predicts the batch size of the inference model based on the output features of the shared hidden layer:

[0191] Batch Size = Round (W batch H3+b batch );

[0192] Where: Batch Size represents the prediction result of the batch size of the inference model; Round represents the integer function; W batch b represents the weight matrix for batch size prediction; batch Represents the bias term for batch size prediction, with dimension 1;

[0193] The execution path prediction unit is a multi-classification neuron that predicts the appropriate hardware execution path of the inference model based on the output features of the shared hidden layer:

[0194] Path=Softmax(W path H3+b path );

[0195] Where: Path represents the execution path prediction result of the inference model; W path Represents the weight matrix of path prediction; b path Represents the bias term for path prediction.

[0196] The loss function of the deep learning model is as follows:

[0197] L total =λ prec L prec +λ batch L batch +λ path L path ;

[0198] Where: L total represents the total loss function; λ prec , batch and path Represents the weight coefficient; L prec Represents the accuracy prediction loss, using the cross entropy loss function; L batch represents the loss of batch size prediction, using the mean square error loss function; L path Represents the loss of execution path prediction, using the cross entropy loss function.

[0199] As a further implementation of this embodiment, an adaptive weighting mechanism is introduced into the deep learning model to weight the output H3 of the shared hidden layer:

[0200] H task,i =A task,i H3;

[0201] Where: H task,i A represents the weighted output for the i-th prediction task, which is the input of the i-th prediction unit; task,i Represents the weight coefficient of the i-th prediction task, which is learned through training.

[0202] The training of the above-mentioned deep learning model will not be described in detail in this embodiment. Given a specific model structure, how to train the model belongs to conventional technical means in the field, so it will not be described in detail in this embodiment.

[0203] The present invention uses a hardware performance analysis tool to comprehensively evaluate the computing power of the AIOT device, ensuring that the hardware performance indicators of the device, including processing power, memory bandwidth, latency, storage space and power consumption, etc., can be accurately understood, thereby realizing the basis for matching the model with the hardware; then, the original reasoning model is layered, and the analysis tool is used to evaluate the computing amount, memory requirement and complexity of each layer, providing detailed model performance data for subsequent optimization; then, the conversion strategy of the model is optimized based on the hardware performance indicators and the model performance data, ensuring that the model can take into account the hardware limitations and model performance requirements during the conversion process; based on the determined conversion strategy, the original reasoning model is converted to obtain a reasoning model that is both hardware-adaptive and performance-maintaining, and is deployed on the AIOT board, thereby achieving the purpose of seamless migration; secondly, a deep learning model and real-time load monitoring are used to dynamically adjust the reasoning accuracy, batch size and execution path, so that the reasoning model can be adaptively optimized according to the actual operation conditions, further improving the efficiency and performance of the model under limited resources; therefore, the present invention ensures the efficient operation of the reasoning model on the AIOT hardware, and solves the model migration and optimization problems under resource-constrained conditions.

[0204] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A seamless migration and adaptive optimization method of reasoning models for AIOT boards, characterized in that: The following steps are involved: Step 1: Evaluate the computing power of the target AIOT device based on the hardware performance analysis tool, clarify the processing power, memory bandwidth, latency, storage space, and power consumption of the computing unit of the AIOT device, and obtain the performance indicators of each computing unit; Step 2: Layer the original inference model and use analytical tools to evaluate the computational workload, memory requirements, and complexity of each layer of the original inference model. Step 3: Optimize the model conversion strategy based on the performance indicators of each computing unit and the computational workload, memory requirements, and complexity of each layer of the original inference model to obtain the conversion strategy of the original inference model; Step 4: Convert the original reasoning model based on the determined conversion strategy to obtain a converted reasoning model, and deploy the converted reasoning model to the AIOT board; Step 5: Use the deep learning model in combination with the actual load of the device to dynamically adjust the inference accuracy, batch size, and execution path.

2. The method for seamless migration and adaptive optimization of reasoning models for AIOT boards according to claim 1 is characterized in that: In step 2, the original inference model is initially layered into an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer, wherein the computational effort and memory requirements of each layer are obtained according to the following steps: Input layer: The computational complexity of the input layer is O(H×W×C), where H represents the height of the input feature, W represents the width, and C represents the number of channels. The memory requirement of the input layer is H×W×C×B, where B represents the batch size. Convolutional Layer: Where: and Respectively represent the height and width of the convolution kernel; Represents the depth of the output feature map of the convolutional layer, that is, the number of convolution kernels; Represents the height of the output feature map of the convolutional layer; Represents the width of the output feature map of the convolutional layer; Indicates the number of channels of the input image of the convolutional layer; Represents the computational effort of the convolutional layer; The memory requirement of the convolution layer is the memory requirement of each convolution kernel. Plus the memory requirements of the output feature map Pooling layer: Where: Represents the height of the feature map output by the pooling layer; Indicates the width of the feature map output by the pooling layer; Indicates the number of channels of the input image of the pooling layer; Indicates the window size of the pooling layer; Represents the computational effort of the pooling layer; The memory requirement of the pooling layer is Fully connected layer: Where: Represents the input vector dimension of the fully connected layer; Represents the output vector dimension of the fully connected layer; Represents the computational effort of the fully connected layer; The memory requirement of the full connection is the sum of the size of the weight matrix and the size of the output vector Output layer: Where: Represents the output dimension of the fully connected layer in the output layer; Indicates the number of categories; Represents the computational complexity of softmax, which is the exponential calculation for each category; The memory requirement of the output layer is the sum of the memory requirement of the fully connected layer in the output layer and the memory requirement of the softmax, that is, 3. The method for seamless migration and adaptive optimization of reasoning models for AIOT boards according to claim 2 is characterized in that: The complexity of each layer of the original inference model is calculated based on the number of floating-point operations and memory bandwidth of each layer: Where: C i represents the complexity of the i-th layer; F ops represents the floating point operation number of the i-th layer; T model Indicates the processing power of the inference model executed on each computing unit, the time required for each floating-point operation; M i represents the memory requirement of the i-th layer; B model Indicates the amount of data that can be transmitted per second.

4. The method for seamless migration and adaptive optimization of reasoning models for AIOT boards according to claim 1 is characterized in that: The step 3 comprises the following steps: Initialization: Initialize the computing unit of each AIOT board, randomly assign each inference model layer to the computing unit, and obtain the initial deployment plan; Layer scheduling: For each inference model layer, the latency and memory bottleneck of each inference model layer on the current computing unit are calculated, and the layer allocation scheme is dynamically adjusted according to the remaining load of each computing unit; Optimization target calculation: Optimization calculation target function for each level: Total_Cost i =w delay ·Delay i +w memory ·Memory_Load i +w power ·Power_Load i ; Where: Delay i Represents the inference model layer L i In the calculation unit H j Delay on Memory_Load i Represents the inference model layer L i In the calculation unit H j Memory bottleneck on Power_Load i Represents the computing unit H j Run the L i The power consumption of the layer model inference layer; w delay 、w memory and w power Represents the weight coefficient; Total_Cost i Represents the total cost, and the optimization goal is to minimize the total cost; Quantization and pruning optimization judgment: judge whether quantization optimization or pruning optimization is needed based on the set objective function threshold and constraint conditions; The amount of calculation after quantization optimization becomes: Where: b quant Indicates the number of quantization bits; F i represents the computational amount of the i-th layer; Indicates the amount of quantized computation; The memory requirements are: Where: M i Indicates the memory requirement of the i-th layer; Indicates the quantized memory requirements; The pruning optimization analyzes the computational graph structure of each layer and uses a structured pruning method to remove redundant connections and neurons. The amount of computation after pruning is as follows: Where: pruning_ratio represents the pruning ratio; Accuracy evaluation: During the quantization and pruning process, the accuracy change of each layer is evaluated in real time to ensure that the accuracy loss does not exceed the preset threshold; A genetic algorithm is used to dynamically adjust the hierarchical allocation and optimization strategy of each computing unit through selection operations, crossover operations, and mutation operations to minimize the overall objective function and output the final conversion strategy, where the conversion strategy includes the inference model layer allocated to each processing unit on the AIOT board and the pruning and / or quantization operations required for each inference model layer.

5. The method for seamless migration and adaptive optimization of reasoning models for AIOT boards according to claim 4 is characterized in that: The constraints include processor performance constraints, memory limit constraints, power consumption limit constraints, delay requirement constraints, and precision loss limit constraints. The processor performance constraint is that each inference model layer, when executed on the computing unit, does not exceed the difference between the maximum processing capacity of the board and the preset redundant processing capacity; The memory limit constraint is that the memory requirement required for each inference model layer does not exceed the difference between the total board memory and the preset memory redundancy requirement; The power consumption limit constraint is that the power consumption required for each inference model layer does not exceed the difference between the total power consumption budget of the board and the preset power consumption redundancy requirement; The delay requirement constraint is that the total delay of the inference model does not exceed a preset maximum acceptable delay; The precision loss limit constraint is that the precision loss of the model after quantization or pruning does not exceed a preset precision loss threshold.

6. The method for seamless migration and adaptive optimization of reasoning models for AIOT boards according to claim 4 is characterized in that: In the crossover operation, a performance-aware crossover point selection strategy is introduced to dynamically adjust the crossover point by considering the performance of the computing unit and the computing requirements of the inference model layer: corss point =α*match degree *(1-performance surplus )+β*computation urgency ; Where: corss point Indicates the position of the intersection, and its value range is [0, gene length], where the gene length is equal to the number of inference model layers multiplied by the number of parameters in each layer; α and β are weight coefficients used to balance the impact of matching degree, performance margin, and computing demand urgency; match degree Indicates the matching degree between performance indicators and computing requirements. It is the dot product between the performance indicator vector and the computing requirement vector divided by the product of their Euclidean distances. The value range is [0, 1]. surplus The performance margin of the computing unit is the ratio of the difference between the available resources of the computing unit and its maximum resources to the maximum resources, and the value range is [0, 1]; computation urgency Indicates the urgency of the computational requirements of the inference model layer. It is the ratio of the computational requirements of the inference model layer to the sum of the computational requirements of all layers, and its value range is [0, 1].

7. The method for seamless migration and adaptive optimization of reasoning models for AIOT boards according to claim 1 is characterized in that: The deep learning model includes an input layer, a shared hidden layer, and a task-specific output layer; The input layer is used to input the load data of the device and the characteristics of the reasoning task; The shared hidden layer receives the features of the input layer and uses a deep neural network to map the input features to a higher-dimensional hidden layer representation, where the hidden layer structure is as follows: H1=ReLU(W1·X1+b1); H2=ReLU(W2·X2+b2); H3=ReLU(W3·X3+b3); Where: W i and b i represents the weight and bias of the i-th layer, i is an integer from 1 to 3; X represents the input feature; X1 represents the input feature of the first hidden layer; X2 represents the input feature of the second hidden layer; X3 represents the input feature of the third hidden layer; ReLU represents the activation function; H1 represents the output of the first hidden layer; H2 represents the output of the second hidden layer; H3 represents the output of the third hidden layer; The task-specific output layer is used to receive the output features of the shared hidden layer, wherein the task-specific output layer includes an accuracy prediction unit, a batch size prediction unit, and an execution path prediction unit; The accuracy prediction unit is a neuron that predicts the accuracy level of the inference model based on the output features of the shared hidden layer: Precision=Softmax(W prec ·H3+b prec ); Where: W prec Represents the weight matrix of accuracy prediction; b prec It represents the bias term of precision prediction; Precision represents the result of precision prediction; The batch size prediction unit is a single neuron that predicts the batch size of the inference model based on the output features of the shared hidden layer: Batch Size=Round(W batch ·H3+b batch ); Where: BatchSize represents the prediction result of the batch size of the inference model; Round represents the integer function; W batch b represents the weight matrix for batch size prediction; batch Represents the bias term for batch size prediction, with dimension 1; The execution path prediction unit is a multi-classification neuron that predicts the appropriate hardware execution path of the inference model based on the output features of the shared hidden layer: Path=Softmax(W path ·H3+b path ); Where: Path represents the execution path prediction result of the inference model; W path Represents the weight matrix of path prediction; b path Represents the bias term for path prediction.

8. The method for seamless migration and adaptive optimization of reasoning models for AIOT boards according to claim 7 is characterized in that: The loss function of the deep learning model is as follows: L total =λ prec L prec +λ batch L batch +λ path L path ; Where: L total represents the total loss function; λ prec , batch and path Represents the weight coefficient; L prec Represents the accuracy prediction loss, using the cross entropy loss function; L batch represents the loss of batch size prediction, using the mean square error loss function; L path Represents the loss of execution path prediction, using the cross entropy loss function.

9. The method for seamless migration and adaptive optimization of reasoning models for AIOT boards according to claim 7 is characterized in that: An adaptive weighting mechanism is introduced into the deep learning model to weight the output H3 of the shared hidden layer: H task,i =A task,i ·H3; Where: H task,i A represents the weighted output for the i-th prediction task, which is the input of the i-th prediction unit; task,i Represents the weight coefficient of the i-th prediction task, which is learned through training.

Citation Information

Cited By

  • Multi-terminal large model compression strategy selection method based on swarm intelligence optimization

    CN121683898A