AI reasoning optimization method and system of dynamic model switching framework for edge device
Through the dynamic model switching framework and federated learning mechanism, AI inference is optimized on edge devices, combined with hierarchical heterogeneous resource management and intelligent power consumption equalization, the problems of model accuracy and resource management on edge devices are solved, and efficient and low-latency AI inference optimization is achieved.
Patent Information
- Application Number
- CN202510847118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The prior art cannot dynamically adjust quantization parameters based on real-time resource state on edge devices, resulting in inconsistent model accuracy losses, low computing efficiency, and lack of effective heterogeneous computing resource management and decision-making mechanisms, making it difficult to make optimal decisions in complex and changeable environments.
The dynamic model switching framework is adopted, and the quantitative compensation model is collaboratively trained through the federated learning mechanism, combined with the hierarchical heterogeneous resource management system and intelligent power consumption equalization technology, and a dynamic decision engine is built using graph neural network to realize dynamic switching of model versions and coordinated scheduling of resources, and optimize the matching of computing resources and power consumption.
It improves the inference accuracy and computing efficiency of edge devices in resource-constrained environments, reduces latency and power consumption, enhances the system's adaptability to environmental changes and the accuracy of decision-making, and realizes adaptive optimization of edge AI systems.
Smart Images

Figure CN120354954A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to AI optimization technology, and particularly to an AI inference optimization method and system for edge devices using a dynamic model switching framework. Background Art
[0002] With the rapid development of artificial intelligence technology, deploying AI applications to edge devices has become an important trend in the industry. Due to limited computing resources, strict power consumption constraints, and variable operating environments, edge devices pose severe challenges to AI model inference. Although traditional cloud-based AI inference models have powerful computing capabilities, they suffer from high network latency and significant privacy and security risks. Therefore, efficiently executing AI inference tasks on edge devices has become a current research hotspot.
[0003] Currently, AI inference optimization on edge devices is mainly achieved through model compression techniques, including methods such as pruning, quantization, and knowledge distillation. These techniques can reduce model size and computational complexity, enabling complex AI models to run on resource-constrained edge devices. At the same time, the development of heterogeneous computing technology has also provided new solutions for edge AI. Through the collaborative work of multiple computing units such as CPUs, GPUs, NPUs, and FPGAs, performance and power consumption can be better balanced.
[0004] Existing model compression methods generally adopt static quantization strategies and cannot dynamically adjust quantization parameters according to the real-time resource status of the device, resulting in inconsistent model accuracy losses in different operating environments, especially performing poorly in edge scenarios with large resource fluctuations.
[0005] Current heterogeneous computing resource management systems lack effective collaborative scheduling mechanisms and cannot adaptively allocate the optimal computing resource combination according to the model characteristics at different compression levels, resulting in low computing efficiency and power consumption waste, and it is difficult to meet the strict power consumption constraints of edge devices.
[0006] Existing model switching strategies usually trigger based on simple preset rules or thresholds, lacking the ability to learn from historical decision-making experiences and comprehensively analyze multi-dimensional system states, and cannot make optimal decisions in complex and variable edge environments, resulting in inappropriate model switching times and unnecessary performance fluctuations. Summary of the Invention
[0007] Embodiments of the present invention provide an AI inference optimization method and system for edge devices using a dynamic model switching framework, which can solve the problems in the prior art.
[0008] In the first aspect of the embodiments of the present invention, an AI inference optimization method for edge devices using a dynamic model switching framework is provided, including: Receive model versions with multiple preset compression levels, and apply the federated learning mechanism to collaboratively train a quantization compensation model among edge device clusters. The quantization compensation model is used to dynamically correct the quantization errors of the model versions at multiple compression levels, and generate a set of compensated optimized model versions; Build a hierarchical heterogeneous resource management system based on the set of optimized model versions. The hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology, and realizes the collaborative scheduling of CPUs, GPUs, and NPUs for the computing characteristics of each model version in the set of optimized model versions; Based on the collaborative scheduling, use reconfigurable computing technology to generate and adjust the FPGA dedicated inference acceleration circuit in real time; Integrate the compensation parameters of the quantization error into the FPGA dedicated inference acceleration circuit, and adopt intelligent power consumption balancing technology to achieve an accurate match between computing resources and power consumption targets through dynamic voltage and frequency adjustment; Monitor and quantitatively evaluate the resource status parameters of the hierarchical heterogeneous resource management system in real time, and generate multi-dimensional status data; Build a dynamic decision engine based on the multi-dimensional status data using a graph neural network. The dynamic decision engine takes the compensation parameters of the quantization error as decision features as input, and integrates knowledge distillation and transfer learning technologies to continuously extract optimization strategies from historical decision-making experiences; When the dynamic decision engine identifies the optimal switching opportunity based on the multi-dimensional status data, trigger the corresponding optimized model version switching operation, and dynamically adjust the decision threshold using an adaptive annealing algorithm.
[0009] Applying the federated learning mechanism to collaboratively train a quantization compensation model among edge device clusters. The quantization compensation model is used to dynamically correct the quantization errors of the model versions at multiple compression levels, and generating a set of compensated optimized model versions includes: Build a hierarchical federated learning network, divide the edge devices in the edge device cluster into core nodes and ordinary nodes according to their computing capabilities, and generate a federated learning objective function based on the local data sets of the core nodes and the ordinary nodes. The federated learning objective function is composed of the sum of the products of the contribution weights of each edge device and the local loss function. The contribution weights are dynamically adjusted according to the computing capabilities and data quality of the edge devices to obtain global model parameters; Input the global model parameters into a quantization compensation module. The quantization compensation module includes a weight quantization unit and an activation value quantization unit. Quantize the global model parameters based on the weight quantization unit to obtain a weight quantization error, quantize the global model parameters according to the activation value quantization unit to obtain an activation value quantization error, and add the product of the weight quantization error and the activation value quantization error to obtain a quantization error compensation function; Take the global model parameters as the teacher model, and the quantized global model parameters as the student model. Construct a knowledge distillation module based on the output features of the teacher model and the output features of the student model. The knowledge distillation module adjusts the distribution of the output features through a temperature parameter, and uses a weight factor to balance the distillation loss. Input the distillation loss and the quantization error compensation function into an optimizer to obtain compensation parameters; optimize the student model based on the compensation parameters to generate a set of optimized model versions with multiple different compression levels.
[0010] Constructing a knowledge distillation module based on the output features of the teacher model and the output features of the student model, where the knowledge distillation module adjusts the distribution of the output features through a temperature parameter and uses a weight factor to balance the distillation loss includes: Construct a knowledge distillation module, input the output features of the teacher model and the student model into the knowledge distillation module, and the knowledge distillation module calculates the basic distillation loss between the teacher model and the student model according to the KL divergence, and adjusts the distribution of the output features based on the temperature parameter; Integrate a processing unit in the knowledge distillation module. The processing unit maps the output features to a three-dimensional phase space based on the Lorenz system, and the three-dimensional phase space describes the evolution of the feature state through control parameters to generate a feature state mapping; Input the feature state mapping into the processing unit. The processing unit calculates the synchronization error between the output features of the teacher model and the output features of the student model based on the principle of generalized synchronization, and dynamically updates the temperature parameter according to the synchronization error and the initial temperature parameter to generate a temperature regulation parameter; Input the temperature regulation parameter into the processing unit. The processing unit calculates a bifurcation metric function within a preset temperature range and determines the optimal temperature parameter based on the bifurcation metric function; the processing unit reconstructs the knowledge representation through the delay coordinate embedding method to generate a reconstructed feature; Input the basic distillation loss, the synchronization error, and the reconstructed feature into an optimization unit. The optimization unit weights and balances the three losses using a weight factor to generate a total loss function, and optimizes the student model based on the total loss function.
[0011] Construct a hierarchical heterogeneous resource management system based on the set of optimized model versions. The hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology, and realizes the cooperative scheduling of CPU, GPU, and NPU for the computing characteristics of each model version in the set of optimized model versions, including: Construct a hierarchical heterogeneous resource management system based on the set of optimized model versions, and map the computing characteristics of each model version in the set of optimized model versions to a heterogeneous resource demand vector; Construct a software-defined virtualization model. The software-defined virtualization model maps the heterogeneous resource demand vector to the virtual resource space based on the Lorenz chaos mapping, describes the dynamic allocation process of virtual resources through a resource state evolution function, and generates a resource allocation strategy. Construct a collaborative scheduling model based on the resource allocation strategy. The collaborative scheduling model uses a multi-objective genetic algorithm to optimize the scheduling schemes of CPU computing resources, GPU computing resources, and NPU computing resources. The multi-objective genetic algorithm takes task completion time and resource utilization rate as optimization objectives, and iteratively optimizes the scheduling scheme through crossover operators and mutation operators to generate a collaborative scheduling strategy. Input the collaborative scheduling strategy into the hierarchical heterogeneous resource management system, and based on the collaborative scheduling strategy, implement the dynamic allocation and collaborative scheduling of the CPU computing resources, the GPU computing resources, and the NPU computing resources.
[0012] Construct a collaborative scheduling model based on the resource allocation strategy. The collaborative scheduling model using a multi-objective genetic algorithm to optimize the scheduling schemes of CPU computing resources, GPU computing resources, and NPU computing resources includes: Construct a resource allocation strategy based on the CPU resource scheduling objective function, the GPU resource scheduling objective function, and the NPU resource scheduling objective function. Input the resource allocation strategy into a hierarchical search space, generate a task scheduling sequence matrix based on the task scheduling order space in the hierarchical search space, and generate a task segmentation granularity vector according to the task segmentation strategy space in the hierarchical search space. Input the task scheduling sequence matrix and the task segmentation granularity vector into a scheduling strategy feature extractor, generate scheduling features based on the scheduling strategy feature extractor, and the scheduling features generate a performance prediction result through a first non-linear transformation and a second non-linear transformation. Iteratively optimize the population based on the performance prediction result, generate a crossover population through an adaptive crossover operator, mutate the crossover population through a dynamic mutation operator, and select an optimized scheduling strategy from the mutated crossover population through an elite selection strategy. Input the optimized scheduling strategy into a joint optimization objective function. The joint optimization objective function evaluates the optimized scheduling strategy based on a weighted combination of a resource utilization rate loss function, a task completion time loss function, and a scheduling efficiency loss function to generate an optimal scheduling scheme.
[0013] And adopt an intelligent power consumption balancing technology to achieve an accurate match between computing resources and power consumption targets through dynamic voltage and frequency regulation; monitor and quantitatively evaluate the resource state parameters of the hierarchical heterogeneous resource management system in real time to generate multi-dimensional state data including: Collect the voltage vector, frequency vector, temperature vector, and utilization vector of heterogeneous computing resources, calculate the basic power consumption value based on the voltage vector and the frequency vector, input the basic power consumption value into the fractal dynamics model. The fractal dynamics model describes the power consumption distribution characteristics through a fractal dimension matrix and characterizes the power consumption fluctuation characteristics through a Hurst exponent time series, generating power consumption state characteristics; Input the temperature vector and the utilization vector into the fractal dynamics model. The fractal dynamics model analyzes the temperature distribution characteristics and utilization distribution characteristics based on the fractal dimension matrix, and characterizes the temperature fluctuation characteristics and utilization fluctuation characteristics through the Hurst exponent time series, generating temperature state characteristics and utilization state characteristics; Construct multi-dimensional state data from the power consumption state characteristics, the temperature state characteristics, and the utilization state characteristics.
[0014] When the dynamic decision engine identifies the optimal switching opportunity based on the multi-dimensional state data, trigger the corresponding optimization model version switching operation, and dynamically adjust the decision threshold using the adaptive annealing algorithm, including: The dynamic decision engine performs a change trend analysis on the multi-dimensional state data through chaotic synchronization mapping, identifies the fluctuation characteristics of the multi-dimensional state data, and generates a switching opportunity evaluation result; Input the switching opportunity evaluation result into the annealing search module. The annealing search module calculates the decision threshold based on the adaptive annealing algorithm. The adaptive annealing algorithm updates the decision threshold using a dynamic cooling strategy to generate the optimal switching opportunity; Trigger the optimization model version switching operation based on the optimal switching opportunity, and switch the current optimization model version to the target optimization model version.
[0015] In the second aspect of the embodiments of the present invention, a dynamic model switching framework for an AI inference optimization system of edge devices is provided, including: A first unit for receiving model versions of multiple preset compression levels, applying a federated learning mechanism to collaboratively train a quantization compensation model among edge device clusters. The quantization compensation model is used to dynamically correct the quantization errors of the model versions of multiple compression levels, generating a set of compensated optimization model versions; A second unit for constructing a hierarchical heterogeneous resource management system based on the set of optimized model versions. The hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology, and realizes the collaborative scheduling of CPUs, GPUs, and NPUs for the computing characteristics of each model version in the set of optimized model versions; Based on the collaborative scheduling, use reconfigurable computing technology to generate and adjust the FPGA dedicated inference acceleration circuit in real time; A third unit, configured to integrate the compensation parameters of the quantization error into the FPGA dedicated inference acceleration circuit, and adopt an intelligent power consumption balancing technology to achieve an accurate match between computing resources and power consumption targets through dynamic voltage and frequency regulation; monitor and quantitatively evaluate the resource status parameters of the hierarchical heterogeneous resource management system in real time, and generate multi-dimensional status data; A fourth unit, configured to construct a dynamic decision engine based on the multi-dimensional status data by using a graph neural network. The dynamic decision engine takes the compensation parameters of the quantization error as decision features as input, and integrates knowledge distillation and transfer learning technologies to continuously extract optimization strategies from historical decision-making experiences; when the dynamic decision engine identifies an optimal switching opportunity based on the multi-dimensional status data, it triggers a corresponding optimization model version switching operation, and dynamically adjusts the decision threshold by using an adaptive annealing algorithm.
[0016] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0017] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0018] The beneficial effects of the present application are as follows: By collaboratively training a quantization compensation model through a federated learning mechanism, the quantization errors of models at different compression levels are effectively corrected, the inference accuracy of the model in a resource-constrained environment is improved, and at the same time, a relatively low computational overhead is maintained, enabling edge devices to execute complex AI tasks under limited resource conditions.
[0019] Based on software-defined hardware technology, a hierarchical heterogeneous resource management system is constructed, realizing the collaborative scheduling of CPUs, GPUs, and NPUs and the reconfigurable computing of FPGAs, giving full play to the advantages of heterogeneous computing resources, improving the inference performance, reducing the latency, and achieving an accurate match between computing resources and power consumption targets through an intelligent power consumption balancing technology, thereby prolonging the battery life of the device.
[0020] By using a graph neural network to construct a dynamic decision engine and combining knowledge distillation and transfer learning technologies, it is possible to intelligently identify the optimal model switching opportunity according to real-time multi-dimensional status data, optimize the resource allocation strategy, improve the adaptability of the system to environmental changes and the accuracy of decision-making, and achieve the adaptive optimization and continuous evolution of the edge AI system. Description of the Drawings
[0021] Figure 1 This is a schematic flowchart of the AI inference optimization method for the dynamic model switching framework of the embodiments of the present invention for edge devices; Figure 2 This is a bar chart for performance comparison and analysis of the federated learning quantization compensation model in the embodiments of the present invention; Figure 3 This is a flowchart of the heterogeneous resource collaborative scheduling model optimized by the multi-objective genetic algorithm in the embodiments of the present invention; Figure 4 This is a schematic diagram for analyzing the utilization distribution characteristics of the fractal dynamics model in the embodiments of the present invention. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0024] Figure 1 This is a schematic flowchart of the AI inference optimization method for the dynamic model switching framework of the embodiments of the present invention for edge devices, as Figure 1 shown, the method includes: Receiving model versions of multiple preset compression levels, applying a federated learning mechanism to collaboratively train a quantization compensation model among edge device clusters, where the quantization compensation model is used to dynamically correct the quantization errors of the model versions of the multiple compression levels to generate a set of compensated optimized model versions; Based on the set of optimized model versions, constructing a hierarchical heterogeneous resource management system, where the hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology, and realizes the collaborative scheduling of CPU, GPU, and NPU for the computing characteristics of each model version in the set of optimized model versions; and based on the collaborative scheduling, using reconfigurable computing technology to generate and adjust the FPGA dedicated inference acceleration circuit in real time; Integrating the compensation parameters of the quantization errors into the FPGA dedicated inference acceleration circuit, and adopting an intelligent power consumption balancing technology to achieve an accurate match between computing resources and power consumption targets through dynamic voltage and frequency regulation; monitoring and quantitatively evaluating the resource state parameters of the hierarchical heterogeneous resource management system in real time to generate multi-dimensional state data; Construct a dynamic decision engine using a graph neural network based on the multi-dimensional state data. The dynamic decision engine takes the compensation parameter of the quantization error as a decision feature input, and integrates knowledge distillation and transfer learning techniques to continuously extract optimization strategies from historical decision-making experiences. When the dynamic decision engine identifies the optimal switching time based on the multi-dimensional state data, it triggers the corresponding optimization model version switching operation and dynamically adjusts the decision threshold using an adaptive annealing algorithm.
[0025] In an alternative implementation, a federated learning mechanism is applied to collaboratively train a quantization compensation model among edge device clusters. The quantization compensation model is used to dynamically correct the quantization errors of model versions at multiple compression levels, and the generated set of compensated optimized model versions includes: Construct a hierarchical federated learning network. Divide the edge devices in the edge device cluster into core nodes and ordinary nodes according to their computing capabilities. Generate a federated learning objective function based on the local data sets of the core nodes and the ordinary nodes. The federated learning objective function is composed of the sum of the products of the contribution weights of each edge device and the local loss function. The contribution weights are dynamically adjusted according to the computing capabilities and data quality of the edge devices to obtain global model parameters. Input the global model parameters into a quantization compensation module. The quantization compensation module includes a weight quantization unit and an activation value quantization unit. Quantize the global model parameters based on the weight quantization unit to obtain a weight quantization error, and quantize the global model parameters based on the activation value quantization unit to obtain an activation value quantization error. Add the product of the weight quantization error and the activation value quantization error to obtain a quantization error compensation function. Use the global model parameters as the teacher model and the quantized global model parameters as the student model. Construct a knowledge distillation module based on the output features of the teacher model and the output features of the student model. The knowledge distillation module adjusts the distribution of the output features through a temperature parameter and balances the distillation loss using a weight factor. Input the distillation loss and the quantization error compensation function into an optimizer to obtain a compensation parameter. Optimize the student model based on the compensation parameter to generate a set of optimized model versions at multiple different compression levels.
[0026] Construct a hierarchical federated learning network. Divide the edge devices in the edge device cluster into core nodes and ordinary nodes according to their computing capabilities. The evaluation of computing capabilities is based on the processor performance, memory size, and energy consumption characteristics of the devices. For example, devices with a processor main frequency exceeding 2.5 GHz, a memory capacity greater than 8 GB, and a power consumption controlled below 10 W are classified as core nodes, and the remaining devices are ordinary nodes. In a practical application scenario, in a cluster of 100 edge devices, 20 devices are identified as core nodes, and the remaining 80 devices are ordinary nodes.
[0027] Generate a federated learning objective function based on the local datasets of core nodes and ordinary nodes. This objective function is composed of the sum of the products of the contribution weights of each edge device and the local loss function. The contribution weights are dynamically adjusted according to the computing power and data quality of the edge devices. The computing power score can be comprehensively calculated through indicators such as the floating-point operation performance and memory bandwidth of the device. For example, the computing power score range of core nodes is 0.8 - 1.0, and that of ordinary nodes is 0.3 - 0.7. The data quality is evaluated through dimensions such as sample diversity and annotation accuracy. For example, the score of a high-quality dataset is 0.9, medium quality is 0.6, and low quality is 0.3. The final contribution weight is obtained by the weighted average of the computing power score and the data quality score, with weight coefficients of 0.6 and 0.4 respectively.
[0028] The federated learning process is carried out through multiple rounds of iteration. Each round of iteration includes two stages: local model training and global model aggregation. In the local model training stage, each edge device updates the model parameters using its own dataset; in the global model aggregation stage, the central server collects all local model parameters and performs a weighted average according to the contribution weights of each device to generate global model parameters. For example, after 10 rounds of iteration, the federated learning converges and stable global model parameters are obtained.
[0029] Input the global model parameters into the quantization compensation module, which includes a weight quantization unit and an activation value quantization unit. The weight quantization unit uses the linear quantization method to convert 32-bit floating-point weights into 8-bit integer representations, and the quantization step size is the original weight range divided by 255. For example, when the weight range is [-2.5, 2.5], the quantization step size is 0.0196. The activation value quantization unit also adopts linear quantization to map the activation value range to integer values from 0 to 255. Based on the weight quantization unit, the global model parameters are quantized to obtain the weight quantization error, which is calculated by the difference between the original weight and the dequantized weight. According to the activation value quantization unit, the global model parameters are quantized to obtain the activation value quantization error, which is calculated by the difference between the original activation value and the dequantized activation value. Add the product of the weight quantization error and the activation value quantization error to obtain the quantization error compensation function.
[0030] Take the global model parameters as the teacher model, and the quantized global model parameters as the student model, and construct a knowledge distillation module based on the output features of the teacher model and the output features of the student model. The knowledge distillation module adjusts the distribution of the output features through a temperature parameter, and the value of the temperature parameter is 4.0, making the soft label distribution smoother and facilitating the student model to learn the knowledge of the teacher model. During the knowledge distillation process, a weight factor is used to balance the distillation loss, and the weight factor is set to 0.7, which means that the distillation loss accounts for 70% of the total loss, and the hard label cross-entropy loss accounts for 30%. Input the distillation loss and the quantization error compensation function into the optimizer to obtain the compensation parameters. The optimizer uses the Adam algorithm with a learning rate of 0.001, and the momentum parameters are set to 0.9 and 0.999.
[0031] Optimize the student model based on the compensation parameters to generate a set of optimized model versions with different compression levels. Specifically, different compression-level models are generated by adjusting the quantization bit width, including 8-bit, 6-bit, and 4-bit quantization versions, corresponding to low compression, medium compression, and high compression levels respectively. These optimized models achieve model size compression ratios of 4 times, 5.3 times, and 8 times respectively, on the premise that the reduction in the accuracy of the original model does not exceed 2%.
[0032] In actual deployment, the system dynamically selects a suitable compression-level model version according to the resource limitations of the edge device. For example, for a core node with sufficient computing resources, a low-compression-level model with 8-bit quantization can be deployed to maintain a high inference accuracy; for a general node with limited resources, a high-compression-level model with 4-bit quantization can be deployed to optimize the execution efficiency. Through this dynamic adaptation mechanism, the system can achieve a balance between computing resources and model performance within the edge device cluster, improving the overall system efficiency.
[0033] Figure 2 Bar chart for performance comparison and analysis of the federated learning quantization compensation model in the embodiments of the present invention: This figure shows the comparison results of three different methods (the present technical solution based on federated learning quantization calculation, traditional quantization methods, and non-federated learning quantization methods) on five key performance indicators. In terms of the quantization accuracy improvement rate, the present technical solution significantly leads, reaching 89.7%, which has obvious advantages over 62.4% of traditional quantization methods and 75.2% of non-federated learning quantization methods; in terms of the inference speed improvement rate, the present technical solution reaches 76.3%, which is also excellent compared to 58.9% of traditional methods and 65.1% of non-federated learning methods; in terms of the model size reduction rate indicator, the present technical solution is 62.8%, which is better than 45.6% of traditional methods and 55.7% of non-federated learning methods; in terms of the energy consumption reduction rate, the present technical solution reaches 48.5%, which is significantly higher than 32.7% of traditional methods and 39.4% of non-federated learning methods; in terms of the response time reduction rate indicator, the present technical solution reaches 72.1%, which is also better than 51.3% of traditional methods and 61.8% of non-federated learning methods. Overall, the present technical solution based on federated learning quantization calculation achieves the best performance in all five key performance indicators, demonstrating significant technical advantages and practical value.
[0034] In an optional implementation manner, a knowledge distillation module is constructed based on the output features of the teacher model and the output features of the student model. The knowledge distillation module adjusts the distribution of the output features through a temperature parameter and balances the distillation loss by using a weight factor, including: Construct a knowledge distillation module, input the output features of the teacher model and the student model into the knowledge distillation module. The knowledge distillation module calculates the basic distillation loss between the teacher model and the student model according to the KL divergence, and adjusts the distribution of the output features based on the temperature parameter; Integrate a processing unit in the knowledge distillation module. The processing unit maps the output features to a three-dimensional phase space based on the Lorenz system. The three-dimensional phase space describes the evolution of the feature state through control parameters, and generates a feature state mapping; Input the feature state mapping into the processing unit. The processing unit calculates the synchronization error between the output features of the teacher model and the output features of the student model based on the principle of generalized synchronization, and dynamically updates the temperature parameter according to the synchronization error and the initial temperature parameter to generate a temperature regulation parameter; Input the temperature regulation parameter into the processing unit. The processing unit calculates the bifurcation metric function within a preset temperature range, and determines the optimal temperature parameter based on the bifurcation metric function; the processing unit reconstructs the knowledge representation through the delay coordinate embedding method to generate a reconstructed feature; Input the basic distillation loss, the synchronization error, and the reconstructed features into the optimization unit. The optimization unit uses a weight factor to weightedly balance the three losses, generates a total loss function, and optimizes the student model based on the total loss function.
[0035] The knowledge distillation module receives the output features of the teacher model and the student model as inputs. For the input feature vectors, first calculate the KL divergence between the output probability distribution P t of the teacher model and the output probability distribution P s of the student model as the basic distillation loss L kd . During the calculation process, introduce a temperature parameter T to adjust the probability distribution, making the knowledge in the soft labels smoother. The initial temperature parameter can be set to 8.0, and by adjusting this parameter, the softness and hardness of knowledge transfer can be controlled.
[0036] The knowledge distillation module integrates a processing unit, which maps the output features to a three-dimensional phase space based on the Lorenz system. Specifically, for the output feature X t of the teacher model and the output feature X s of the student model, perform the mapping through the following steps: extract the key components in the feature vector and use these components as the initial conditions of the Lorenz system; set the control parameters of the Lorenz system σ = 10.0, ρ = 28.0, β = 8 / 3, and obtain the trajectory of the feature in the phase space through iterative calculation; sample and record the point coordinates on the trajectory to form the feature state mappings M t and M s . In practical applications, 1000 iterative steps can be selected, and the state is recorded every 10 steps. Finally, a state mapping composed of 100 sampling points is obtained.
[0037] The processing unit calculates the synchronization error between the output features of the teacher model and the student model based on the principle of generalized synchronization. For the obtained feature state mappings M t and M s , calculate their Euclidean distance sequences in the phase space and take their mean as the synchronization error E sync . When the synchronization error is large, it indicates that the knowledge representation differences between the student model and the teacher model are significant, and knowledge transfer needs to be strengthened; when the synchronization error is small, it indicates that the student model has learned the knowledge representation of the teacher model well. Based on the synchronization error and the initial temperature parameter, dynamically update the temperature parameter: if the synchronization error E sync is greater than the threshold 0.5, increase the temperature parameter to make the knowledge transfer softer; if the synchronization error is less than the threshold 0.2, decrease the temperature parameter to make the knowledge transfer harder. The update formula can be designed as: T new =T old ×(1 + α×(E sync-0.35)), where α is an adjustment coefficient and can be set to 0.5.
[0038] The processing unit calculates the bifurcation metric function within a preset temperature range to determine the optimal temperature parameter. The temperature range can be set to [1.0, 20.0] and sampled at a step size of 1.0. For each temperature value T i , calculate the knowledge distillation loss L kd (T i ) and the synchronization error E sync (T i ), construct the bifurcation metric function D(T i ) = L kd (T i ) × (1 + γ × E sync (T i ))), where γ is a balance coefficient and can be set to 0.3. Select the temperature value that minimizes D(T i ) as the optimal temperature parameter T opt . In practical applications, if D(T opt ) obtains the minimum value of 0.087 when T i = 6.0, then the optimal temperature parameter is determined to be 6.0.
[0039] The processing unit reconstructs the knowledge representation through the delay coordinate embedding method. For the feature state mappings M t and M s , select the embedding dimension d = 5 and the delay time τ = 2 to construct the embedding vector sequence. By calculating the similarity matrix of these embedding vectors, extract the principal components to generate the reconstructed features R t and R s . The reconstructed features can capture the temporal correlation and nonlinear dynamic characteristics of the original features. The reconstruction error E rec is defined as the cosine distance between R t and R s . In a specific application scenario, the reconstructed features can reduce the original 768-dimensional features to 128 dimensions while maintaining the key information, and the reconstruction error is controlled below 0.15.
[0040] The optimization unit uses the weight factor to weighted balance the three losses to generate the total loss function. The total loss function is defined as: L total = λ1 × L kd + λ2 × E sync + λ3 × E rec, where λ1, λ2, and λ3 are weight factors that respectively control the contributions of the base distillation loss, synchronization error, and reconstruction error. The weight factors can be adjusted according to the specific task. A typical setting is λ1 = 0.5, λ2 = 0.3, and λ3 = 0.2. The student model optimizes its parameters based on the total loss function through backpropagation and gradient descent algorithms. During the iterative optimization process, the loss value gradually decreases from the initial 1.45 to 0.38, and the model accuracy increases from 82.3% to 95.7%.
[0041] In this embodiment, by dynamically adjusting the temperature parameter and the multi-objective optimization strategy, in the image classification task, the student model (ResNet-18) achieves a test accuracy of 98.2% of the teacher model (ResNet-50) while maintaining the model size at only 30% of the teacher model. In the natural language processing task, the performance of the small BERT model (4 layers) trained by this method in the text classification task reaches 96.5% of the original BERT model (12 layers), and the inference speed is increased by 3.2 times.
[0042] In an alternative embodiment, a hierarchical heterogeneous resource management system is constructed based on the set of optimized model versions. The hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology. The collaborative scheduling of CPU, GPU, and NPU for each model version in the set of optimized model versions includes: Construct a hierarchical heterogeneous resource management system based on the set of optimized model versions, and map the computing characteristics of each model version in the set of optimized model versions to a heterogeneous resource demand vector; Construct a software-defined virtualization model. The software-defined virtualization model maps the heterogeneous resource demand vector to the virtual resource space based on the Lorenz chaos mapping, describes the dynamic allocation process of virtual resources through a resource state evolution function, and generates a resource allocation strategy; Construct a collaborative scheduling model based on the resource allocation strategy. The collaborative scheduling model uses a multi-objective genetic algorithm to optimize the scheduling schemes of CPU computing resources, GPU computing resources, and NPU computing resources. The multi-objective genetic algorithm takes the task completion time and resource utilization rate as optimization objectives, and iteratively optimizes the scheduling scheme through crossover operators and mutation operators to generate a collaborative scheduling strategy; Input the collaborative scheduling strategy into the hierarchical heterogeneous resource management system, and based on the collaborative scheduling strategy, realize the dynamic allocation and collaborative scheduling of the CPU computing resources, the GPU computing resources, and the NPU computing resources.
[0043] Build a hierarchical heterogeneous resource management system based on a set of optimized model versions, and map the computational characteristics of each model version in the set of optimized model versions to a heterogeneous resource requirement vector. Specifically, the system collects the computational characteristics of the model version, including parameters such as computational density, memory access pattern, parallelism, and data dependence. For the image recognition model version A, its computational characteristics include a floating-point operation count of 3.5×10 9 , a memory access frequency of 2.1×10 6 times per second, a parallelism of 85%, and a low data dependence; for the natural language processing model version B, its computational characteristics include a floating-point operation count of 1.8×10 9 , a memory access frequency of 4.3×10 6 times per second, a parallelism of 65%, and a medium data dependence.
[0044] Convert these computational characteristics into a five-dimensional heterogeneous resource requirement vector through a feature extraction algorithm, corresponding to the number of CPU cores, the number of GPU stream processors, the number of NPU computing units, the memory capacity, and the bandwidth requirement respectively. For example, the heterogeneous resource requirement vector of model version A is [4, 1024, 128, 8GB, 16GB / s], and the heterogeneous resource requirement vector of model version B is [8, 512, 64, 12GB, 10GB / s].
[0045] Build a software-defined virtualization model. This model maps the heterogeneous resource requirement vector to the virtual resource space based on the Lorenz chaos map, describes the dynamic allocation process of virtual resources through a resource state evolution function, and generates a resource allocation strategy. In the specific implementation, the parameters of the Lorenz chaos map used by the system are σ = 10, ρ = 28, β = 8 / 3, the initial conditions are x0 = 1.0, y0 = 1.0, z0 = 1.0, and the number of iterations is 1000. For the heterogeneous resource requirement vector [4, 1024, 128, 8GB, 16GB / s] of model version A, the coordinates in the virtual resource space obtained after the Lorenz chaos map are [0.73, 0.45, 0.89]. The system maintains a virtual resource state matrix with a size of 100×100×100, representing a three-dimensional virtual resource space. The resource state evolution function describes the variation law of each element in this matrix over time, considering factors such as resource allocation, release, and migration. Based on the virtual resource space coordinates of the model version and the current system resource state, the resource state evolution function generates a resource allocation strategy, such as model version A being allocated 4 CPU cores, 1 GPU (including 1024 stream processors), 0 NPUs, 8GB of memory, and 16GB / s of bandwidth.
[0046] A collaborative scheduling model is constructed based on a resource allocation strategy. This model uses a multi-objective genetic algorithm to optimize the scheduling schemes of CPU computing resources, GPU computing resources, and NPU computing resources. The multi-objective genetic algorithm takes the task completion time and resource utilization rate as optimization objectives, and iteratively optimizes the scheduling scheme through the crossover operator and mutation operator to generate a collaborative scheduling strategy. In the specific implementation, the collaborative scheduling model uses a binary encoding method with the chromosome length being the number of model versions × 3, representing the execution ratios of each model version on the CPU, GPU, and NPU. The population size is set to 100, the number of iterations is 500, the crossover probability is 0.8, and the mutation probability is 0.05. The fitness function consists of two parts: the task completion time and the resource utilization rate, with weights of 0.6 and 0.4 respectively. For model versions A and B, 100 scheduling schemes are randomly generated in the initial population. For example, Scheme 1 is [0.2, 0.7, 0.1, 0.5, 0.3, 0.2], indicating that the execution ratios of model version A on the CPU, GPU, and NPU are 20%, 70%, and 10% respectively, and the execution ratios of model version B on the CPU, GPU, and NPU are 50%, 30%, and 20% respectively. After 500 iterations of optimization, the optimal scheduling scheme [0.15, 0.8, 0.05, 0.6, 0.25, 0.15] is obtained. The calculated task completion time is 42 milliseconds, and the resource utilization rate is 87%.
[0047] Input the collaborative scheduling policy into the hierarchical heterogeneous resource management system, and realize the dynamic allocation and collaborative scheduling of CPU computing resources, GPU computing resources, and NPU computing resources based on the collaborative scheduling policy. The system adopts a hierarchical architecture, including a resource abstraction layer, a virtualization layer, and a scheduling layer. The resource abstraction layer is responsible for managing physical computing resources, including an 8-core CPU, 2 GPUs (each containing 2048 stream processors), and 4 NPUs (each containing 256 computing units). The virtualization layer, based on software-defined hardware technology, virtualizes the physical resource pool into allocable resource units, such as CPU cores, GPU stream processor blocks, and NPU computing unit blocks. The scheduling layer dynamically allocates virtualized resources according to the collaborative scheduling policy to achieve the parallel execution of model versions on heterogeneous hardware. For model version A, the system allocates 1.2 CPU cores (accounting for 15% of the total execution), 1638 GPU stream processors (accounting for 80% of the total execution), and 13 NPU computing units (accounting for 5% of the total execution); for model version B, the system allocates 4.8 CPU cores (accounting for 60% of the total execution), 512 GPU stream processors (accounting for 25% of the total execution), and 38 NPU computing units (accounting for 15% of the total execution). The system also implements a resource dynamic adjustment mechanism, monitoring the resource utilization every 100 milliseconds. When the detected resource utilization is lower than 70% or higher than 95%, it triggers the resource reallocation process and updates the collaborative scheduling policy. In this way, the system realizes the efficient collaborative scheduling of CPU, GPU, and NPU resources, reduces the task completion time, and improves the resource utilization rate.
[0048] In an alternative embodiment, a collaborative scheduling model is constructed based on the resource allocation policy. The collaborative scheduling model uses a multi-objective genetic algorithm to optimize the scheduling schemes of CPU computing resources, GPU computing resources, and NPU computing resources, including: Construct a resource allocation policy based on the CPU resource scheduling objective function, the GPU resource scheduling objective function, and the NPU resource scheduling objective function. Input the resource allocation policy into the hierarchical search space, generate a task scheduling sequence matrix based on the task scheduling order space in the hierarchical search space, and generate a task segmentation granularity vector according to the task segmentation policy space in the hierarchical search space; Input the task scheduling sequence matrix and the task segmentation granularity vector into the scheduling policy feature extractor, generate scheduling features based on the scheduling policy feature extractor, and generate a performance prediction result through a first non-linear transformation and a second non-linear transformation; Iteratively optimize the population based on the performance prediction result, generate a crossover population through an adaptive crossover operator, mutate the crossover population through a dynamic mutation operator, and select an optimized scheduling policy from the mutated crossover population through an elite selection strategy; Input the optimized scheduling policy into the joint optimization objective function. The joint optimization objective function evaluates the optimized scheduling policy based on the weighted combination of the resource utilization loss function, the task completion time loss function, and the scheduling efficiency loss function, and generates an optimal scheduling plan.
[0049] As Figure 3 shown, the method further includes: In the specific implementation, when constructing a cooperative scheduling model based on the resource allocation policy, the model uses a multi-objective genetic algorithm to optimize the scheduling plans of CPU, GPU, and NPU computing resources, and realizes efficient cooperative scheduling of heterogeneous computing resources.
[0050] The construction of the cooperative scheduling model starts from the scheduling objective functions of the three computing resources. The CPU resource scheduling objective function considers the number of CPU cores, the main frequency, and the cache size, and calculates the efficiency index of the CPU to process tasks through weighted calculation. For example, for a CPU with 8 cores, 3.2 GHz, and 16 MB cache, when the weights are set to 0.4, 0.3, and 0.3 respectively, the CPU resource efficiency index is calculated as (0.4×8 + 0.3×3.2 + 0.3×16) = 8.76. The GPU resource scheduling objective function comprehensively considers the number of GPU cores, the video memory size, and the computing power, and sets corresponding weights to calculate the performance index of the GPU to process tasks. For example, for a GPU with 2048 cores, 16 GB video memory, and 14 TFLOPS computing power, when the weights are set to 0.35, 0.3, and 0.35 respectively, the GPU performance index is (0.35×2048 + 0.3×16 + 0.35×14000) = 5721.6. The NPU resource scheduling objective function evaluates the efficiency of the NPU to process specific tasks based on the AI computing power, energy consumption ratio, and parallel processing ability of the NPU. Taking an NPU with 200 TOPS computing power, 5 TOPS / W energy consumption ratio, and 64 parallel processing units as an example, when the weights are set to 0.4, 0.3, and 0.3 respectively, the NPU efficiency index is calculated as (0.4×200 + 0.3×5 + 0.3×64) = 99.1.
[0051] After the above resource allocation strategy is input into the hierarchical search space, the system generates a task scheduling sequence matrix based on the task scheduling order space. For the scenario of 5 tasks and 3 types of computing resources, the task scheduling sequence matrix can be represented as a 5×3 matrix, where the element values represent the scheduling priorities of tasks on specific resources. For example, the first row [1, 3, 2] in the matrix means that the priority of task 1 on the CPU is 1, on the GPU is 3, and on the NPU is 2. At the same time, the system generates a task splitting granularity vector according to the task splitting strategy space. For the case of 5 tasks, the task splitting granularity vector can be represented as [0.3, 0.5, 0.2, 0.4, 0.6], indicating the splitting ratio of each task, which determines the degree to which tasks can be processed in parallel.
[0052] After the task scheduling sequence matrix and the task splitting granularity vector are input into the scheduling policy feature extractor, the system extracts the key features of the scheduling policy through a multi-layer perceptron structure. The feature extraction network includes an input layer, two hidden layers, and an output layer, and the number of nodes in the hidden layers is 64 and 32 respectively. The input features are transformed through the first non-linear transformation to obtain an intermediate representation, and the output of each node is calculated through the ReLU activation function. For example, for the input feature vector [0.3, 0.8, 0.5], after being transformed by the weight matrix [0.2, 0.3, 0.1; 0.1, 0.4, 0.2; 0.3, 0.1, 0.4] and applying the ReLU function, the intermediate representation [0.41, 0.38, 0.43] can be obtained. The intermediate representation is then subjected to a second non-linear transformation, and the performance prediction results, including three indicators of the estimated completion time, resource utilization rate, and scheduling efficiency, are generated through another weight matrix and the ReLU function.
[0053] Based on the performance prediction results, the system iteratively optimizes the initial population. The initial population contains 50 scheduling policies, and each policy consists of a task scheduling sequence matrix and a task splitting granularity vector. When generating the crossover population through the adaptive crossover operator, the system dynamically adjusts the crossover probability, which is calculated according to the fitness values of the parent individuals. For example, when the fitness values of the parent individuals are 0.7 and 0.8 respectively, the crossover probability is set to 0.5+(0.9 - 0.5)×(1-(0.7 + 0.8) / 2)=0.65. The crossover operation generates two offspring individuals, inheriting and recombining some features of the parents. For the dynamic mutation operator, the system dynamically adjusts the mutation probability according to the current iteration number, and the mutation probability decreases as the iteration number increases. For example, in the 10th iteration, when the maximum iteration number is 100, the mutation probability can be set to 0.1+(0.5 - 0.1)×(1 - 10 / 100)=0.46. The mutation operation randomly changes some gene values in the individuals, increasing the population diversity. Through the elite selection strategy, the system retains the top 10% of the individuals with the highest fitness values directly into the next generation, and the remaining positions are selected from the mutated crossover population through tournament selection.
[0054] The optimized scheduling strategy is input into the joint optimization objective function for final evaluation. The resource utilization loss function calculates the gap between the average utilization rate of three computing resources and the ideal utilization rate. For example, when the utilization rates of CPU, GPU, and NPU are 75%, 85%, and 60% respectively, and the target utilization rate is 90%, the resource utilization loss is calculated as 1 - ((75 + 85 + 60) / 3) / 90 = 0.18. The task completion time loss function evaluates the ratio of the actual completion time to the theoretical shortest completion time. For the case where the estimated completion time is 150 seconds and the theoretical shortest time is 120 seconds, the time loss is (150 - 120) / 120 = 0.25. The scheduling efficiency loss function measures the number of resource switches and the load balance degree. When the number of resource switches is 15 times, which is 0.3 after normalization, and the load imbalance degree is 0.2, the scheduling efficiency loss is 0.3×0.6 + 0.2×0.4 = 0.26. The joint optimization objective function combines the three loss functions with weights of 0.4, 0.4, and 0.2 to obtain the final evaluation score of 0.18×0.4 + 0.25×0.4 + 0.26×0.2 = 0.224. The scheduling strategy with the lowest evaluation score is selected as the optimal scheduling plan to achieve efficient cooperative scheduling of heterogeneous computing resources.
[0055] In an alternative embodiment, an intelligent power consumption balancing technology is adopted to achieve an accurate match between computing resources and power consumption targets through dynamic voltage and frequency regulation; the resource state parameters of the hierarchical heterogeneous resource management system are monitored and quantitatively evaluated in real time to generate multi-dimensional state data including: Collect the voltage vector, frequency vector, temperature vector, and utilization vector of heterogeneous computing resources, calculate the basic power consumption value based on the voltage vector and the frequency vector, input the basic power consumption value into the fractal dynamics model, and the fractal dynamics model describes the power consumption distribution characteristics through a fractal dimension matrix and characterizes the power consumption fluctuation characteristics through a Hurst exponent time series to generate power consumption state characteristics. Input the temperature vector and the utilization vector into the fractal dynamics model, and the fractal dynamics model analyzes the temperature distribution characteristics and utilization distribution characteristics based on the fractal dimension matrix, and characterizes the temperature fluctuation characteristics and utilization fluctuation characteristics through the Hurst exponent time series to generate temperature state characteristics and utilization state characteristics; construct multi-dimensional state data from the power consumption state characteristics, the temperature state characteristics, and the utilization state characteristics.
[0056] Collect the voltage vector, frequency vector, temperature vector, and utilization vector of heterogeneous computing resources through an embedded sensor network. The voltage vector is represented as V=(v1,v2...v n), where v1 represents the real-time voltage value of the first computing unit, with the unit of volt; the frequency vector is expressed as F=(f1,f2,...,f n ), where f1 represents the real-time operating frequency of the first computing unit, with the unit of hertz; the temperature vector is expressed as T=(t1,t2...t n ), where t1 represents the real-time temperature value of the first computing unit, with the unit of degree Celsius; the utilization rate vector is expressed as U=(u1,u2...u n ), where u1 represents the real-time computing resource utilization rate of the first computing unit, expressed as a percentage.
[0057] Based on the collected voltage vector and frequency vector, the system calculates the basic power consumption value. In actual implementation, for each computing unit i, its basic power consumption P i is calculated by multiplying the square of the voltage by the frequency and then by a constant coefficient C related to the computing unit i . For example, for a system with 4 computing units, if the voltage vector is (0.9V, 1.1V, 1.0V, 1.2V), the frequency vector is (2.4GHz, 2.8GHz, 2.5GHz, 3.0GHz), and the corresponding constant coefficients are (0.5, 0.6, 0.55, 0.65), then the calculated basic power consumption values are (0.972W, 2.033W, 1.375W, 2.808W).
[0058] The calculated basic power consumption value is input into the fractal dynamics model. This model describes the power consumption distribution characteristics through the fractal dimension matrix and characterizes the power consumption fluctuation characteristics through the Hurst exponent time series. The fractal dimension matrix is expressed as an n×n matrix D, where each element D ij represents the fractal dimension value of the power consumption distribution between the i-th computing unit and the j-th computing unit, and its range is usually between 1.0 and 2.0. The Hurst exponent time series is expressed as H=(h1,h2...h n ), where h1 represents the Hurst exponent of the power consumption change of the first computing unit, and its range is usually between 0 and 1. When the value of h1 is close to 1, it indicates that the power consumption change has strong persistence; when the value of h1 is close to 0.5, it indicates that the power consumption change is close to a random walk; when the value of h1 is close to 0, it indicates that the power consumption change has strong anti-persistence.
[0059] Through the calculation of the fractal dynamics model, the system generates power consumption state characteristics. In an actual case, the fractal dimension matrix of a 4-core processor is [[1.65, 1.42, 1.38, 1.27], [1.42, 1.72, 1.45, 1.32], [1.38, 1.45, 1.68, 1.40], [1.27, 1.32, 1.40, 1.70]], and its Hurst exponent time series is [0.82, 0.78, 0.75, 0.80], indicating that the power consumption changes of each core have strong persistence and there is a certain correlation between them.
[0060] The system also inputs the temperature vector and the utilization vector into the fractal dynamics model. The model analyzes the temperature distribution characteristics and the utilization distribution characteristics based on the fractal dimension matrix, and characterizes the temperature fluctuation characteristics and the utilization fluctuation characteristics through the Hurst exponent time series. For the temperature vector (45°C, 52°C, 48°C, 56°C), the fractal dimension matrix is [[1.55, 1.38, 1.32, 1.25], [1.38, 1.62, 1.40, 1.30], [1.32, 1.40, 1.58, 1.35], [1.25, 1.30, 1.35, 1.60]], and the Hurst exponent time series is [0.76, 0.72, 0.70, 0.74]. For the utilization vector (65%, 78%, 70%, 85%), the fractal dimension matrix is [[1.70, 1.45, 1.42, 1.30], [1.45, 1.75, 1.48, 1.35], [1.42, 1.48, 1.72, 1.38], [1.30, 1.35, 1.38, 1.68]], and the Hurst exponent time series is [0.85, 0.82, 0.80, 0.88].
[0061] Based on the analysis results of the fractal dynamics model, the system generates temperature state characteristics and utilization state characteristics. The temperature state characteristics include the fractal dimension matrix of the temperature distribution and the Hurst exponent time series of the temperature fluctuation; the utilization state characteristics include the fractal dimension matrix of the utilization distribution and the Hurst exponent time series of the utilization fluctuation.
[0062] The system constructs multi-dimensional state data from the power consumption state characteristics, the temperature state characteristics, and the utilization state characteristics. This multi-dimensional state data is represented in tensor form and contains three main dimensions: the power consumption dimension, the temperature dimension, and the utilization dimension. Each dimension contains two sub-dimensions: the fractal dimension feature and the Hurst exponent feature. For example, in the case of the above 4-core processor, the multi-dimensional state data can be represented as a 3×2×4 tensor, where 3 represents the number of main dimensions (power consumption, temperature, utilization), 2 represents the number of sub-dimensions of each main dimension (fractal dimension, Hurst exponent), and 4 represents the number of computing units.
[0063] Based on the constructed multi-dimensional state data, the system realizes intelligent power consumption balancing. The system analyzes the relationships among the power consumption characteristics, temperature characteristics, and utilization characteristics of each computing unit in the multi-dimensional state data, and identifies power consumption hotspots and utilization imbalance points. For the identified problem points, the system performs precise regulation through dynamic voltage and frequency scaling technology. For example, when it is identified that the power consumption of the 4th computing unit is too high (2.808W) and the temperature is relatively high (56°C), the system reduces its voltage from 1.2V to 1.1V and the frequency from 3.0GHz to 2.8GHz, thereby reducing the power consumption to 2.352W while keeping the computing performance within an acceptable range.
[0064] Through the above method, the present invention realizes intelligent power consumption balancing of heterogeneous computing resources, ensures optimal power distribution while meeting the computing requirements of the system, extends the service life of the device, and improves energy utilization efficiency. This technology is applicable to various scenarios such as high-performance computing systems, mobile computing devices, and data centers.
[0065] Figure 4 Schematic diagram for analyzing the utilization distribution characteristics of the fractal dynamics model in the embodiment of the present invention: This figure shows the comparison of the intelligent power consumption balancing scheme based on chaotic fractal dynamics (represented by circular markers in the legend) and the traditional dynamic voltage and frequency scaling scheme (represented by diamond markers in the legend) in terms of the CPU utilization changing over time. The horizontal axis represents time in minutes, ranging from 0 minutes to 60 minutes; the vertical axis represents the CPU utilization percentage, ranging from 0% to 100%. From the curve trends, it can be seen that both schemes show a trend of rising first and then falling, but there are obvious differences in their specific performances. The traditional dynamic voltage and frequency scaling scheme reaches a peak of approximately 94% at 20 minutes, while the intelligent power consumption balancing scheme based on chaotic fractal dynamics reaches a peak of approximately 85% at 25 minutes, with the peak reduced by approximately 9 percentage points. In the initial stage (0 - 15 minutes), the CPU utilization of the present technical scheme rises relatively fast, quickly climbing from 25% to 82%; in contrast, the starting value of the dynamic voltage and frequency scaling scheme is lower (about 15%), and the rising speed is more gentle, reaching approximately 55% at 15 minutes. In the descending stage after the peak (30 - 60 minutes), the descending rate of the dynamic voltage and frequency scaling scheme is faster, dropping to approximately 18% at 60 minutes, while the present technical scheme still maintains a relatively high level of approximately 32%.
[0066] In an alternative embodiment, when the dynamic decision engine identifies the optimal switching time based on the multi-dimensional state data, it triggers the corresponding optimization model version switching operation and dynamically adjusts the decision threshold using the adaptive simulated annealing algorithm, including: The dynamic decision engine analyzes the change trend of the multi-dimensional state data through chaotic synchronization mapping, identifies the fluctuation characteristics of the multi-dimensional state data, and generates an evaluation result of the switching time; Input the switching opportunity evaluation result into the annealing search module. The annealing search module calculates the decision threshold based on the adaptive annealing algorithm. The adaptive annealing algorithm updates the decision threshold using a dynamic cooling strategy to generate the optimal switching opportunity. Trigger the optimization model version switching operation based on the optimal switching opportunity, and switch the current optimization model version to the target optimization model version.
[0067] Receive multi-dimensional state data, which can include key indicators such as system operation parameters, performance metrics, and resource utilization. Specifically, the multi-dimensional state data can be a time-series data set D = {d1(t), d2(t)...d n (t)} with n dimensions, where d1(t) can represent the CPU usage rate, d2(t) can represent the memory occupancy, d3(t) can represent the request response time, d4(t) can represent the throughput, etc., which are the values of key indicators at time t.
[0068] The dynamic decision engine analyzes the trend of these multi-dimensional state data through chaotic synchronization mapping. Chaotic synchronization mapping is a non-linear dynamics analysis method that can effectively capture small changes in the system state and amplify these changes to identify turning points in the system state. In actual implementation, for each dimension of state data d i (t), the following processing method is used for chaotic synchronization mapping: First, standardize the original data to the [0,1] interval to obtain the standardized data s i (t). Subsequently, apply a chaotic mapping function to the standardized data to generate a mapping sequence m i (t). The choice of the mapping function can be the Logistic mapping. For each time point t, the mapping result depends on the value at the previous time point and the control parameter r, which is usually set between 3.7 and 4 to ensure that the system exhibits chaotic characteristics. By continuously applying this mapping function, a mapping sequence reflecting the change characteristics of the original data can be obtained.
[0069] By calculating the difference Δm i (t)=|m i (t)-m i (t - 1)| between the mapping values at adjacent time points, construct a difference sequence. When a significant peak appears in the difference sequence, it indicates that the original data is undergoing important changes. For example, when Δm i (t) exceeds 3 times the historical average, it can be marked as a potential turning point. By analyzing the distribution of turning points in multiple dimensions, the overall fluctuation characteristics of the system state can be comprehensively evaluated.
[0070] To quantify this fluctuation characteristic, the dynamic decision-making engine calculates the fluctuation intensity index F, which comprehensively considers the change amplitude, change frequency of the state data in each dimension, and the co-variation pattern across dimensions. For example, when the CPU usage suddenly rises from 45% to 90%, while the memory occupancy rises from 50% to 85%, and the request response time increases from 200 ms to 800 ms, the fluctuation intensity index F will increase significantly, indicating that the system is at a critical moment when the optimization model version needs to be switched. In this way, the dynamic decision-making engine generates the switching opportunity evaluation result E = (F, P), where F represents the fluctuation intensity and P represents the fluctuation pattern feature vector.
[0071] After inputting the switching opportunity evaluation result E into the annealing search module, the annealing search module calculates the decision threshold based on the adaptive annealing algorithm. The adaptive annealing algorithm is a heuristic optimization method that simulates the energy change law in the metal annealing process and finds the global optimal solution through probabilistic search in the solution space. In this embodiment, the annealing search module initializes the decision threshold T0, and the typical value can be set to 0.8. The initial temperature Temp0 is set to 100, indicating that the system initially has a relatively high probability of accepting suboptimal solutions.
[0072] The annealing search module updates the decision threshold using a dynamic cooling strategy. Specifically, in each iteration i, the temperature is updated according to Temp i = α × Temp i-1 where α is the cooling coefficient, usually set between 0.9 and 0.99. For example, when α = 0.95, the temperature will decrease at a relatively slow rate, giving the algorithm enough opportunity to explore the solution space. At the same time, the update of the decision threshold takes into account the difference between the current system state and the historical state. When the system state fluctuates greatly, the cooling rate will adaptively slow down to avoid premature convergence to the local optimal solution; when the system state is relatively stable, the cooling rate will speed up to improve the convergence efficiency.
[0073] In practical applications, assume that the current fluctuation intensity of the system F = 0.75, which is lower than the initial decision threshold T0 = 0.8, and at this time, the switching operation is not triggered. As the system load continues to increase, the fluctuation intensity rises to F = 0.85, exceeding the decision threshold, and the system will calculate the switching benefit value G. If G is positive and large enough, and the current temperature Temp allows accepting this switching operation, then this moment is marked as a potential optimal switching opportunity. As the annealing process continues, the algorithm finds a better switching opportunity, and finally, when the temperature Temp drops to the preset termination temperature (such as 1.0), the final optimal switching opportunity is determined.
[0074] Based on the determined optimal switching time, the system triggers the operation of switching the optimized model version. The switching process includes the following steps: First, the system loads the target optimized model version into memory to ensure that the loading process does not affect the current system operation; Second, the system prepares the switching environment, including operations such as data status saving and cache preheating; Subsequently, during an appropriate request processing gap, the system performs the actual switching, redirecting the request traffic from the current optimized model version to the target optimized model version; Finally, the system conducts post-switch verification to ensure that the target optimized model version is running properly.
[0075] For example, in an online recommendation system, when the user traffic gradually increases from 1000 requests per second to 5000 requests per second, the CPU usage rate of the system rises from 40% to 85%, and the memory occupancy rises from 50% to 90%. Through chaotic synchronization mapping analysis, the system identifies that the fluctuation intensity F = 0.88, which exceeds the current decision threshold T = 0.82. The annealing search module calculates the switching gain G = 0.35 (a positive value indicates that the switch is beneficial), and the current temperature Temp = 45 allows this switch to be accepted. The system then immediately switches the current lightweight recommendation model to a more complex but resource-efficient distributed recommendation model. After the switch, the CPU usage rate drops to 65%, and the memory occupancy drops to 75%, while maintaining the same recommendation quality, successfully coping with the increased user load.
[0076] The method further includes: The model pre-compression module performs multi-level compression processing on the original AI model on the cloud server. Taking the image classification model ResNet50 as an example, three model versions with different compression levels can be generated: The high-compression version uses 4-bit quantization technology, with a model size of 12MB, an accuracy rate of 83.2%, and an inference speed of 15ms / frame; The medium-compression version uses 8-bit quantization technology, with a model size of 45MB, an accuracy rate of 89.5%, and an inference speed of 28ms / frame; The low-compression version retains 32-bit floating-point precision, with a model size of 198MB, an accuracy rate of 93.7%, and an inference speed of 65ms / frame. Model pre-compression is implemented using the quantization tools in PyTorch. The specific quantization parameters include: weight quantization bit width setting (4 bits / 8 bits / 32 bits), activation value quantization range, and quantization method selection (symmetric / asymmetric quantization).
[0077] The resource monitoring module collects edge device status data through lightweight system calls. On Android devices, the battery power percentage is obtained through the BatteryManager API, the available memory is obtained through ActivityManager.MemoryInfo, and the CPU usage rate is read through the / proc / stat file; On Linux devices, through / sys / class / power supplyRead battery information, obtain memory information through / proc / meminfo, and get CPU load through / proc / loadavg. The monitoring frequency is configurable, and a typical setting is to collect data every 500 milliseconds. Data collection example: In a continuous video analysis task on a certain smartphone, it was monitored that the battery level dropped from 85% to 19%, the CPU usage fluctuated between 65% - 78%, and the available memory decreased from 1.2GB to 450MB.
[0078] Data processing uses a moving average method with a window size of 5 to smooth short-term fluctuations and avoid unnecessary model switches triggered by instantaneous peaks. Threshold settings refer to the actual device performance test results: the battery threshold is set to low (<20%), medium (20% - 80%), high (>80%); the CPU usage threshold is low (<30%), medium (30% - 70%), high (>70%); the memory availability threshold is low (<20% available), medium (20% - 50% available), high (>50% available).
[0079] The dynamic switching module selects the best model version according to the device state. The switching decision uses a rule-based system, implemented as a decision tree structure. For example, when it is detected that the battery level is below 20% and the CPU usage is above 70% for 5 seconds continuously, a switch to a high-compression model is triggered; when the battery level is above 80% and the CPU usage is below 30% for 10 seconds continuously, a switch to a low-compression model is made. The switching frequency control mechanism sets a 5-second cooling period and a 5% hysteresis buffer to prevent frequent switching caused by state fluctuations. In actual tests, during the process of the smartphone battery dropping from 100% to 15%, the number of model switches was controlled within 3 times, effectively avoiding performance fluctuations.
[0080] Model priority management uses a dynamic scoring mechanism, comprehensively considering the current resource state and model characteristics. In the case of limited resources (such as battery <20%), the score of the high-compression model is increased by 50%; in the case of sufficient resources (such as battery >80%), the score of the low-compression model is increased by 30%. Calculate the scores of each model during each switching decision and select the model version with the highest score.
[0081] The seamless transition module ensures that the model switch does not interrupt the inference service. It is implemented based on the Model Control Protocol of NVIDIA Triton Inference Server, supporting dynamic loading and unloading of models. The switching process includes: requesting to load a new model, verifying the model ready state, replacing the current model, and releasing the resources of the old model. To avoid data loss, a FIFO buffer queue with a size of 64 items is set up to temporarily store the data to be processed. For example, in a video analysis application, the video frames during the switching period are stored in the buffer queue and processed in order after the new model is ready. The measured data shows that zero frames are lost during the switching process.
[0082] The switching time is controlled within 80 milliseconds, and the impact on 30fps real-time video analysis does not exceed 3 frames. The rollback mechanism for model loading failure is implemented through status check: if the new model does not reach the "ready" state within 10 seconds, it will automatically roll back to the old model to ensure service continuity. In the case of model loading failure, the inference service interruption time does not exceed 100 milliseconds.
[0083] The policy management module supports user-defined switching policies, using JSON format configuration files. Typical policy configuration example: {"priority":"performance","battery threshold low ”:15,“cpu threshold high ”:75,“switch cooldown ”:8000}. This configuration specifies a performance-first policy with a low battery threshold of 15%, a high CPU threshold of 75%, and a switch cooldown period of 8 seconds. Policy conflict resolution uses a priority mechanism: resource protection policies (such as low battery policies) have the highest priority, performance optimization policies are second, and accuracy optimization policies have the lowest priority.
[0084] In specific application scenarios, taking edge video analysis as an example, measured data show that after adopting the dynamic model switching framework, the battery life of smartphones is extended from 4.2 hours to 7.5 hours, an increase of 78.6%; the average power consumption is reduced from 3.2W to 1.8W, a reduction of 43.8%; the reasoning accuracy is maintained between 83.2%-93.7% in different scenarios, meeting application requirements; the switching process has a negligible impact on the user experience, and the user satisfaction score is increased from 7.2 to 8.9 (out of 10 points).
[0085] In IoT camera applications, after adopting the dynamic model switching framework, the battery life of the device is extended from 18 hours to 30 hours, an increase of 66.7%; basic functions can still be maintained in the low-battery stage, avoiding service interruptions caused by insufficient power in traditional solutions; the accuracy is dynamically adjusted according to the scenario, and it automatically switches to a high-accuracy model at critical moments (such as detecting abnormal activities), improving security monitoring effects.
[0086] Applicable devices include smartphones, tablets, smart cameras, edge servers, IoT sensors, etc. Tests and verifications cover operating systems such as Android, iOS, and Linux, as well as processor architectures such as ARM, x86, and RISC-V, with compatibility of more than 95%. On extremely resource-constrained devices (such as RAM < 512MB), by configuring a lightweight monitoring strategy (reducing the monitoring frequency to 2 seconds / time) and reducing the buffer queue size (to 16 items), the framework can still run effectively, occupying only about 5MB of memory and no more than 3% of CPU resources.
[0087] In the second aspect of the embodiments of the present invention, a dynamic model switching framework for an AI inference optimization system of edge devices is provided, including: A first unit, configured to receive model versions of multiple preset compression levels, and apply a federated learning mechanism to collaboratively train a quantization compensation model among an edge device cluster. The quantization compensation model is used to dynamically correct the quantization errors of the model versions of the multiple compression levels, and generate a set of compensated optimized model versions; A second unit, configured to construct a hierarchical heterogeneous resource management system based on the set of optimized model versions. The hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology, and realizes the collaborative scheduling of the CPU, GPU, and NPU for the computing characteristics of each model version in the set of optimized model versions; based on the collaborative scheduling, a dedicated FPGA inference acceleration circuit is generated and adjusted in real time by using reconfigurable computing technology; A third unit, configured to integrate the compensation parameters of the quantization errors into the dedicated FPGA inference acceleration circuit, and adopt an intelligent power consumption balancing technology to achieve an accurate match between computing resources and power consumption targets through dynamic voltage and frequency regulation; monitor and quantitatively evaluate the resource state parameters of the hierarchical heterogeneous resource management system in real time, and generate multi-dimensional state data; A fourth unit, configured to construct a dynamic decision engine based on the multi-dimensional state data by using a graph neural network. The dynamic decision engine takes the compensation parameters of the quantization errors as decision features as input, and integrates knowledge distillation and transfer learning technologies to continuously extract optimization strategies from historical decision-making experiences; when the dynamic decision engine identifies an optimal switching opportunity based on the multi-dimensional state data, it triggers a corresponding optimized model version switching operation, and dynamically adjusts the decision threshold by using an adaptive annealing algorithm.
[0088] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0089] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0090] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for optimizing AI inference for edge devices using a dynamic model switching framework, characterized in that, Including: Receiving model versions with multiple preset compression levels, applying a federated learning mechanism to collaboratively train a quantization compensation model among an edge device cluster, where the quantization compensation model is used to dynamically correct the quantization errors of the model versions with multiple compression levels, and generating a set of compensated optimized model versions; Building a hierarchical heterogeneous resource management system based on the set of optimized model versions, where the hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology, and realizes the collaborative scheduling of CPU, GPU, and NPU for the computing characteristics of each model version in the set of optimized model versions; generating and adjusting an FPGA dedicated inference acceleration circuit in real time based on the collaborative scheduling using reconfigurable computing technology; Integrating the compensation parameters of the quantization errors into the FPGA dedicated inference acceleration circuit, and adopting an intelligent power consumption balancing technology to achieve an accurate match between computing resources and power consumption targets through dynamic voltage and frequency regulation; monitoring and quantitatively evaluating the resource state parameters of the hierarchical heterogeneous resource management system in real time, and generating multi-dimensional state data; Building a dynamic decision-making engine based on the multi-dimensional state data using a graph neural network, where the dynamic decision-making engine takes the compensation parameters of the quantization errors as decision features as input, and integrates knowledge distillation and transfer learning technologies to continuously extract optimization strategies from historical decision-making experiences; When the dynamic decision-making engine identifies the optimal switching opportunity based on the multi-dimensional state data, triggering the corresponding optimized model version switching operation, and dynamically adjusting the decision threshold using an adaptive annealing algorithm.
2. The method according to claim 1, characterized in that, Applying a federated learning mechanism to collaboratively train a quantization compensation model among an edge device cluster, where the quantization compensation model is used to dynamically correct the quantization errors of the model versions with multiple compression levels, and generating a set of compensated optimized model versions includes: Building a hierarchical federated learning network, dividing the edge devices in the edge device cluster into core nodes and ordinary nodes according to their computing capabilities, generating a federated learning objective function based on the local data sets of the core nodes and the ordinary nodes, where the federated learning objective function is composed of the sum of the products of the contribution weights of each edge device and the local loss function, and the contribution weights are dynamically adjusted according to the computing capabilities and data quality of the edge devices to obtain global model parameters; Inputting the global model parameters into a quantization compensation module, where the quantization compensation module includes a weight quantization unit and an activation value quantization unit, quantizing the global model parameters based on the weight quantization unit to obtain a weight quantization error, quantizing the global model parameters based on the activation value quantization unit to obtain an activation value quantization error, and adding the product of the weight quantization error and the activation value quantization error to obtain a quantization error compensation function; Taking the global model parameters as the teacher model and the quantized global model parameters as the student model, a knowledge distillation module is constructed based on the output features of the teacher model and the output features of the student model. The knowledge distillation module adjusts the distribution of the output features through a temperature parameter and uses a weight factor to balance the distillation loss. The distillation loss and the quantization error compensation function are input into an optimizer to obtain compensation parameters. Based on the compensation parameters, the student model is optimized to generate a set of optimized model versions with multiple different compression levels.
3. The method according to claim 2, characterized in that, Constructing a knowledge distillation module based on the output features of the teacher model and the output features of the student model, where the knowledge distillation module adjusts the distribution of the output features through a temperature parameter and uses a weight factor to balance the distillation loss includes: Construct a knowledge distillation module, input the output features of the teacher model and the student model into the knowledge distillation module, and the knowledge distillation module calculates the basic distillation loss between the teacher model and the student model according to the KL divergence, and adjusts the distribution of the output features based on the temperature parameter; Integrate a processing unit in the knowledge distillation module, and the processing unit maps the output features to a three-dimensional phase space based on the Lorenz system. The three-dimensional phase space describes the evolution of the feature state through control parameters to generate a feature state mapping; Input the feature state mapping into the processing unit, and the processing unit calculates the synchronization error between the output features of the teacher model and the output features of the student model based on the principle of generalized synchronization, and dynamically updates the temperature parameter according to the synchronization error and the initial temperature parameter to generate a temperature regulation parameter; Input the temperature regulation parameter into the processing unit, and the processing unit calculates a bifurcation metric function within a preset temperature range, and determines the optimal temperature parameter based on the bifurcation metric function; the processing unit reconstructs the knowledge representation through the delay coordinate embedding method to generate a reconstructed feature; Input the basic distillation loss, the synchronization error, and the reconstructed feature into an optimization unit, and the optimization unit uses a weight factor to weight and balance the three losses to generate a total loss function, and optimizes the student model based on the total loss function.
4. The method according to claim 1, wherein Construct a hierarchical heterogeneous resource management system based on the set of optimized model versions. The hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology, and realizes the coordinated scheduling of CPU, GPU, and NPU for the computing characteristics of each model version in the set of optimized model versions, including: Construct a hierarchical heterogeneous resource management system based on the set of optimized model versions, and map the computing characteristics of each model version in the set of optimized model versions to a heterogeneous resource demand vector; Construct a software-defined virtualization model, which maps the heterogeneous resource demand vector to a virtual resource space based on the Lorenz chaotic mapping, and describes the dynamic allocation process of virtual resources through a resource state evolution function to generate a resource allocation strategy; Construct a collaborative scheduling model based on the resource allocation strategy. The collaborative scheduling model uses a multi-objective genetic algorithm to optimize the scheduling schemes of CPU computing resources, GPU computing resources, and NPU computing resources. The multi-objective genetic algorithm takes task completion time and resource utilization rate as optimization objectives, and iteratively optimizes the scheduling scheme through crossover operators and mutation operators to generate a collaborative scheduling strategy; Input the collaborative scheduling strategy into the hierarchical heterogeneous resource management system, and realize the dynamic allocation and collaborative scheduling of the CPU computing resources, the GPU computing resources, and the NPU computing resources based on the collaborative scheduling strategy.
5. The method according to claim 4, characterized in that Constructing a collaborative scheduling model based on the resource allocation strategy, and the collaborative scheduling model using a multi-objective genetic algorithm to optimize the scheduling schemes of CPU computing resources, GPU computing resources, and NPU computing resources includes: Construct a resource allocation strategy based on the CPU resource scheduling objective function, the GPU resource scheduling objective function, and the NPU resource scheduling objective function. Input the resource allocation strategy into the hierarchical search space, generate a task scheduling sequence matrix based on the task scheduling order space in the hierarchical search space, and generate a task segmentation granularity vector according to the task segmentation strategy space in the hierarchical search space; Input the task scheduling sequence matrix and the task segmentation granularity vector into a scheduling strategy feature extractor, generate scheduling features based on the scheduling strategy feature extractor, and the scheduling features generate a performance prediction result through a first non-linear transformation and a second non-linear transformation; Iteratively optimize the population based on the performance prediction result, generate a crossover population through an adaptive crossover operator, mutate the crossover population through a dynamic mutation operator, and select an optimized scheduling strategy from the mutated crossover population through an elite selection strategy; Input the optimized scheduling strategy into a joint optimization objective function. The joint optimization objective function evaluates the optimized scheduling strategy based on a weighted combination of a resource utilization rate loss function, a task completion time loss function, and a scheduling efficiency loss function to generate an optimal scheduling scheme.
6. The method according to claim 1, wherein And adopt an intelligent power consumption balancing technology to achieve an accurate match between computing resources and power consumption targets through dynamic voltage and frequency regulation; monitor and quantitatively evaluate the resource status parameters of the hierarchical heterogeneous resource management system in real time to generate multi-dimensional state data including: Collect the voltage vector, frequency vector, temperature vector, and utilization rate vector of heterogeneous computing resources, calculate the basic power consumption value based on the voltage vector and the frequency vector, input the basic power consumption value into a fractal dynamics model. The fractal dynamics model describes the power consumption distribution characteristics through a fractal dimension matrix and characterizes the power consumption fluctuation characteristics through a Hurst exponent time series to generate power consumption state characteristics; Input the temperature vector and the utilization rate vector into the fractal dynamics model. The fractal dynamics model analyzes the temperature distribution characteristics and utilization rate distribution characteristics based on the fractal dimension matrix, and characterizes the temperature fluctuation characteristics and utilization rate fluctuation characteristics through the Hurst exponent time series, generating temperature state characteristics and utilization rate state characteristics; construct multi-dimensional state data from the power consumption state characteristics, the temperature state characteristics and the utilization rate state characteristics.
7. The method according to claim 1, characterized in that When the dynamic decision engine identifies the optimal switching opportunity based on the multi-dimensional state data, trigger the corresponding optimization model version switching operation, and dynamically adjust the decision threshold using the adaptive annealing algorithm, including: The dynamic decision engine performs a change trend analysis on the multi-dimensional state data through chaotic synchronization mapping, identifies the fluctuation characteristics of the multi-dimensional state data, and generates a switching opportunity evaluation result; Input the switching opportunity evaluation result into the annealing search module. The annealing search module calculates the decision threshold based on the adaptive annealing algorithm. The adaptive annealing algorithm updates the decision threshold using a dynamic cooling strategy to generate the optimal switching opportunity; Trigger the optimization model version switching operation based on the optimal switching opportunity, and switch the current optimization model version to the target optimization model version.
8. The dynamic model switching framework is an AI inference optimization system for edge devices, which is used to implement the method described in any one of the foregoing claims 1-7, and is characterized in that, Including: The first unit is used to receive model versions of multiple preset compression levels, apply the federated learning mechanism to collaboratively train the quantization compensation model among the edge device clusters. The quantization compensation model is used to dynamically correct the quantization errors of the model versions of multiple compression levels, generating a set of compensated optimization model versions; The second unit is used to construct a hierarchical heterogeneous resource management system based on the set of optimization model versions. The hierarchical heterogeneous resource management system dynamically allocates and virtualizes computing resources based on software-defined hardware technology, and realizes the collaborative scheduling of CPU, GPU and NPU for the computing characteristics of each model version in the set of optimization model versions; generate and adjust the FPGA dedicated inference acceleration circuit in real time based on the collaborative scheduling using reconfigurable computing technology; The third unit is used to integrate the compensation parameters of the quantization error into the FPGA dedicated inference acceleration circuit, and adopt the intelligent power consumption balancing technology to achieve the precise matching of computing resources and power consumption targets through dynamic voltage and frequency regulation; monitor and quantitatively evaluate the resource state parameters of the hierarchical heterogeneous resource management system in real time, generating multi-dimensional state data; The fourth unit is used to construct a dynamic decision engine based on the multi-dimensional state data using a graph neural network. The dynamic decision engine inputs the compensation parameters of the quantization error as decision features, and integrates knowledge distillation and transfer learning technologies to continuously extract optimization strategies from historical decision-making experiences; When the dynamic decision engine identifies the optimal switching opportunity based on the multi-dimensional state data, trigger the corresponding optimization model version switching operation, and dynamically adjust the decision threshold using the adaptive annealing algorithm.
9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Federal learning method based on client selection and gradient compression
CN115796271A
Federal reconstruction model compression optimization method based on chain type federal learning
CN118194933A
Model deployment method and system
CN120012046A
Federal learning method, device and equipment for multi-base-station unmanned vehicle
CN120128933A
Engineering machinery distributed edge computing system based on multi-heterogeneous teacher distillation method
CN120179404A
Cited By
Neural network reasoning performance analysis method for NPU computing architecture
CN120611352A
Multi-dimensional phase space scheduling method and device, equipment and medium
CN121462583A
Self-adaptive load resource scheduling method and device for embedded GPU (Graphics Processing Unit)
CN121807504A
An adaptive load resource scheduling method and device for embedded GPU
CN121807504B
Optimization method suitable for communication minimization maximum problem of edge device
CN121864799A