AI model collaborative acceleration and energy efficiency optimization method and device
By dividing the AI model into multiple subgraphs and combining deepen reinforcement learning and global energy efficiency models, the problems of AI model collaborative acceleration and energy efficiency optimization caused by device heterogeneity in edge computing environments are solved, and efficient computing and energy efficiency optimization are achieved.
Patent Information
- Application Number
- CN202510209421.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-08
AI Technical Summary
In an edge computing environment, due to device heterogeneity, there are challenges in collaborative acceleration and energy efficiency optimization of AI models, and the existing technology is difficult to effectively solve.
The AI model is constructed into directed acyclic graphs, divided into multiple subgraphs through graph segmentation algorithms, and tasks are allocated and optimized using deepen reinforcement learning algorithms and global energy efficiency models. Combined with compression perception technology, data transmission is optimized, and task allocation strategies are dynamically adjusted to adapt to changes in equipment state.
Improve computing efficiency, ensure calculation accuracy and data consistency, reduce overall energy consumption, extend equipment battery life, and optimize energy efficiency and delay balance.
Smart Images

Figure CN120276836A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an AI model collaborative acceleration and energy efficiency optimization method and device. Background Art
[0002] With the rapid development of applications such as the Internet of Things, smart cities, and autonomous driving, edge computing has gradually become an important way of data processing. Compared with the traditional cloud computing mode, edge computing can process data near the data source, reduce data transmission latency, and improve real-time performance and privacy protection. However, devices in the edge computing environment usually have heterogeneity, that is, there are significant differences in the computing capabilities, storage resources, network bandwidth, and energy consumption characteristics of different devices. This heterogeneity poses a great challenge to the efficient inference of artificial intelligence (AI) models on edge devices.
[0003] Therefore, the problem of how to better perform AI model collaborative acceleration has become an urgent problem to be solved in the industry. Summary of the Invention
[0004] The present invention provides an AI model collaborative acceleration and energy efficiency optimization method and device to solve the problem of how to better perform AI model collaborative acceleration in the prior art.
[0005] The present invention provides an AI model collaborative acceleration and energy efficiency optimization method, including the following steps: Construct the AI model data to be processed into a directed acyclic graph, and based on the graph segmentation algorithm, according to the device information of each collaborative device, divide the directed acyclic graph into multiple model sub-graphs, and then obtain the first model sub-graph tasks corresponding to each collaborative device; wherein, the nodes of the directed acyclic graph represent computing layers, and the edges of the directed acyclic graph represent data dependencies; Based on the deep reinforcement learning algorithm, adjust the task allocation of the first model sub-graph tasks on each collaborative device to obtain the second model sub-graph tasks corresponding to each collaborative device; Based on the global energy efficiency model corresponding to each collaborative device and the multi-objective integer programming algorithm, optimize the task energy efficiency of the second model sub-graph tasks corresponding to each collaborative device to obtain the third model sub-graph tasks corresponding to each collaborative device, and run the corresponding third model sub-graph tasks on each collaborative device.
[0006] According to the AI model collaborative acceleration and energy efficiency optimization method provided by the present invention, after the step of optimizing the task energy efficiency of the second model sub-graph tasks corresponding to each collaborative device based on the global energy efficiency model corresponding to each collaborative device and the multi-objective integer programming algorithm to obtain the third model sub-graph tasks corresponding to each collaborative device, and running the corresponding third model sub-graph tasks on each collaborative device, it further includes: Obtain the device information fed back by each of the collaborative devices during the execution of the corresponding third model sub-graph task; wherein, the device information includes at least one of: device computing power consumption information, device resource utilization information, and device heat dissipation information; Dynamically adjust the model sub-graph tasks corresponding to each of the collaborative devices according to the device information fed back by each collaborative device during the execution of the third model sub-graph task.
[0007] According to an AI model collaborative acceleration and energy efficiency optimization method provided by the present invention, after the step of constructing the to-be-processed AI model data into a directed acyclic graph and dividing the directed acyclic graph into multiple model sub-graphs based on a graph segmentation algorithm according to the device information of each collaborative device to obtain the first model sub-graph tasks corresponding to each collaborative device, the method further includes: Perform sparsity analysis and compression on each of the first model sub-graph tasks through a compressive sensing technique to obtain the compressed first model sub-graph tasks; Determine the encoding strategy for each collaborative device according to the network bandwidth and latency information of each collaborative device, and transmit the compressed first model sub-graph tasks to the corresponding collaborative devices according to the encoding strategy.
[0008] According to an AI model collaborative acceleration and energy efficiency optimization method provided by the present invention, the step of performing task assignment adjustment on the first model sub-graph tasks on each collaborative device based on a deep reinforcement learning algorithm to obtain the second model sub-graph tasks corresponding to each collaborative device includes: Input the real-time state data of each collaborative device into the deep reinforcement learning model in the deep reinforcement learning algorithm, and output the task assignment adjustment strategy corresponding to each collaborative device; wherein, the real-time state data includes: real-time device resource utilization data, real-time device power consumption data, and real-time device task latency data; Perform task assignment adjustment on the first model sub-graph tasks corresponding to each of the collaborative devices according to the task assignment adjustment strategy corresponding to each collaborative device to obtain the second model sub-graph tasks corresponding to each collaborative device.
[0009] According to an AI model collaborative acceleration and energy efficiency optimization method provided by the present invention, before the step of inputting the real-time state data of each collaborative device into the deep reinforcement learning model in the deep reinforcement learning algorithm and outputting the task assignment adjustment strategy corresponding to each collaborative device, the method further includes: A preset deep Q-network model continuously learns the relationship between the real-time state data of each collaborative device and the task assignment adjustment strategy through a reward function to optimize the deep reinforcement learning model; Stop training when the preset training conditions are met to obtain the deep reinforcement learning model.
[0010] According to an AI model collaborative acceleration and energy efficiency optimization method provided by the present invention, the reward function is specifically as follows: Wherein, Rt is the reward value, Ut is the real-time data of the device resource utilization rate, Pt is the real-time data of the device power consumption, Lt is the real-time data of the device task delay, λ1 , λ2 , λ3 are the corresponding weight coefficients respectively.
[0011] The present invention also provides an AI model collaborative acceleration and energy efficiency optimization device, including the following modules: A partitioning module, configured to construct the to-be-processed AI model data into a directed acyclic graph, and based on a graph segmentation algorithm, according to the device information of each collaborative device, partition the directed acyclic graph into multiple model sub-graphs, and obtain the first model sub-graph tasks corresponding to each collaborative device; wherein, the nodes of the directed acyclic graph represent computing layers, and the edges of the directed acyclic graph represent data dependencies; An adjustment module, configured to perform task assignment adjustment on the first model sub-graph tasks on each collaborative device based on a deep reinforcement learning algorithm, and obtain the second model sub-graph tasks corresponding to each collaborative device; An optimization model, configured to perform task energy efficiency optimization on the second model sub-graph tasks corresponding to each collaborative device based on the global energy efficiency model and multi-objective integer programming algorithm corresponding to each collaborative device, obtain the third model sub-graph tasks corresponding to each collaborative device, and run the corresponding third model sub-graph tasks on each collaborative device.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the AI model collaborative acceleration and energy efficiency optimization method as described in any one of the above is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the AI model collaborative acceleration and energy efficiency optimization method as described in any one of the above is implemented.
[0014] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the AI model collaborative acceleration and energy efficiency optimization method as described in any one of the above is implemented.
[0015] The AI model collaborative acceleration and energy efficiency optimization method and device provided by the present invention can make full use of the computing resources of each collaborative device and improve the computing efficiency by dividing the AI model into multiple subgraphs and performing task allocation according to device information. At the same time, this division method takes into account the data dependence relationship, ensuring the correctness of the calculation and the consistency of the data. Through the dynamic adjustment of the DRL model, the task allocation can adapt to the changes in the system state, ensuring the efficient use of computing resources. The design of the reward function enables the model to balance multiple objectives while optimizing the computing performance, taking into account both energy consumption and latency. Through the optimization of the global energy efficiency model and the multi-objective integer programming algorithm, the task allocation not only considers the computing efficiency but also takes into account the energy consumption and energy efficiency. This optimization method can significantly reduce the overall energy consumption while ensuring the system performance, extending the battery life of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the embodiments or the prior art description. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 Schematic flowchart of the AI model collaborative acceleration and energy efficiency optimization method provided by the present invention; Figure 2 System architecture diagram of the overall technical solution provided by the present invention; Figure 3 Schematic flowchart of the dynamic task scheduling and load balancing provided by the present invention; Figure 4 Training and inference flowchart of the deep reinforcement learning model provided by the present invention; Figure 5 Schematic flowchart of the data compression and transmission optimization provided by the present invention; Figure 6 Schematic diagram of the global energy efficiency optimization described in the present invention provided by the present invention; Figure 7 Multi-objective optimization flowchart of the task scheduling and energy efficiency optimization provided by the present invention; Figure 8 Schematic flowchart of the real-time monitoring and feedback control system provided by the present invention; Figure 9 Schematic diagram of the task fault tolerance and recovery mechanism provided by the present invention; Figure 10 Schematic diagram of the cross-device collaborative optimization provided by the present invention; Figure 11 Balanced optimization diagram of energy efficiency and performance provided by the present invention; Figure 12 Schematic structural diagram of the AI model collaborative acceleration and energy efficiency optimization device provided by the present invention; Figure 13 Schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0018] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0019] Figure 1 Schematic flow diagram of the AI model collaborative acceleration and energy efficiency optimization method provided by the present invention, as Figure 1 shown, including: Step 110, constructing the to-be-processed AI model data into a directed acyclic graph, and based on a graph partitioning algorithm, dividing the directed acyclic graph into multiple model subgraphs according to the device information of each collaborative device, so as to obtain first model subgraph tasks corresponding to each collaborative device; wherein, nodes of the directed acyclic graph represent computing layers, and edges of the directed acyclic graph represent data dependencies; In the present invention, in a multi-device collaborative computing scenario, collaborative devices refer to multiple devices jointly participating in a computing task. These devices can be heterogeneous, which means they may vary in hardware configuration, performance and function. For example, collaborative devices can include servers equipped with powerful GPUs, personal computers with integrated graphics processors, and mobile devices with relatively weak computing power but low power consumption, etc. These devices are connected through a network to form a collaborative computing environment to jointly complete complex computing tasks. When executing the AI model collaborative acceleration and energy efficiency optimization method, the collaborative devices dynamically adjust task allocation according to their respective device information and task requirements to achieve efficient utilization of computing resources and minimization of energy consumption.
[0020] In the present invention, first, the to-be-processed AI model data needs to be converted into a directed acyclic graph (DAG). In this graph, nodes represent computing layers of the model, such as convolutional layers, pooling layers or fully connected layers, etc. Each node encapsulates a specific computing operation, and its input and output are specific data tensors. Edges represent the data dependency relationships between these computing layers, reflecting how data is passed from one layer to another in the model. For example, if the output of one layer is used as the input of another layer, there will be a directed edge connecting these two layers.
[0021] Finally, the DAG is divided into multiple model subgraphs, and each subgraph corresponds to a first model subgraph task on a collaborative device. The size, computational load, and data dependency relationship of each subgraph are optimized to ensure that in subsequent collaborative computing, each device can efficiently execute its own tasks, and the performance and energy efficiency of the entire system are maximized.
[0022] Step 120: Based on the deep reinforcement learning algorithm, adjust the task allocation of the first model subgraph tasks on each collaborative device to obtain the second model subgraph tasks corresponding to each collaborative device. Using the deep reinforcement learning (DRL) algorithm, a deep Q-network (DQN) model is established. This model automatically adjusts the task allocation strategy of the first model subgraph tasks on each collaborative device by continuously learning the relationship between the device state and the scheduling result.
[0023] The DRL model selects the optimal action according to the current system state and dynamically adjusts the task allocation. For example, if the resource utilization rate of a certain device is too high, the model may migrate some tasks to other devices to achieve load balancing.
[0024] Step 130: Based on the global energy efficiency model corresponding to each collaborative device and the multi-objective integer programming algorithm, optimize the task energy efficiency of the second model subgraph tasks corresponding to each collaborative device to obtain the third model subgraph tasks corresponding to each collaborative device, and run the corresponding third model subgraph tasks on each collaborative device.
[0025] In the present invention, a global energy efficiency model is constructed, which comprehensively considers factors such as the computational power consumption, resource utilization rate, and heat dissipation characteristics of each collaborative device. By evaluating the energy efficiency performance of each device in real time, the model can provide a basis for energy efficiency optimization of task scheduling.
[0026] Combined with the multi-objective integer programming algorithm, further optimize the second model subgraph tasks corresponding to each collaborative device. The optimization objectives include minimizing the overall system energy consumption, maximizing the computational performance, and meeting the real-time requirements of the tasks.
[0027] On the premise of meeting the system performance requirements, minimize the overall energy consumption by adjusting the task allocation strategy. For example, allocate high-power consumption tasks to devices with higher energy efficiency, or reduce the power consumption of devices when they are idle.
[0028] In the present invention, by dividing the AI model into multiple subgraphs and performing task allocation according to device information, the computing resources of each collaborative device can be fully utilized, improving the computing efficiency. At the same time, this division method takes into account data dependencies, ensuring the correctness of the calculation and the consistency of the data. Through the dynamic adjustment of the DRL model, the task allocation can adapt to changes in the system state, ensuring the efficient utilization of computing resources. The design of the reward function enables the model to balance multiple objectives while optimizing the computing performance, taking into account both energy consumption and latency. Through the optimization of the global energy efficiency model and the multi-objective integer programming algorithm, the task allocation not only considers computing efficiency but also takes into account energy consumption and energy efficiency. This optimization method can significantly reduce the overall energy consumption while ensuring the system performance, extending the battery life of the device.
[0029] Optionally, after the step of performing task energy efficiency optimization on the second model subgraph tasks corresponding to each collaborative device based on the global energy efficiency model and the multi-objective integer programming algorithm corresponding to each collaborative device, obtaining the third model subgraph tasks corresponding to each collaborative device, and running the corresponding third model subgraph tasks on each collaborative device, the method further includes: Obtaining device information fed back by each collaborative device during the running of the corresponding third model subgraph task; wherein, the device information includes at least one of: device computing power consumption information, device resource utilization information, and device heat dissipation information; Dynamically adjusting the model subgraph tasks corresponding to each collaborative device according to the device information fed back by each collaborative device during the running of the third model subgraph task.
[0030] In the present invention, during the task running process, the system continuously collects the running state information of each collaborative device, including the computing power consumption, resource utilization, and heat dissipation of the device. Through this feedback information, the system can understand the energy efficiency performance and running state of each device in real time.
[0031] The system dynamically adjusts the task allocation of each collaborative device according to the collected device information in combination with the global energy efficiency model. For example, if the computing power consumption of a certain device is too high or the resource utilization is close to saturation, the system can migrate some tasks to other devices for running to optimize the overall energy efficiency and system performance. This dynamic adjustment mechanism ensures that the system always maintains high energy efficiency performance and computing performance during the running process.
[0032] For example, by constructing an energy efficiency function and combining the computing power consumption, resource utilization, and heat dissipation characteristics of the device, the system can evaluate and optimize the energy efficiency performance of each device.
[0033] On the other hand, the system can also use the MIP algorithm to perform task scheduling among multiple devices to minimize the overall energy consumption while meeting the latency requirements of the tasks.
[0034] In the embodiment of the present invention, during the operation of the system, the real-time monitoring module continuously collects the resource utilization, task execution status and data transmission status of each device, and dynamically adjusts the parameters of each module through the feedback mechanism. The system continuously optimizes the task scheduling strategy and data transmission path based on the real-time data to ensure that the best performance is always maintained in a changing environment.
[0035] Optionally, after the step of constructing the AI model data to be processed as a directed acyclic graph, and dividing the directed acyclic graph into multiple model subgraphs based on the graph segmentation algorithm according to the device information of each collaborative device, and obtaining the first model subgraph task corresponding to each collaborative device, the step further includes: Performing sparsity analysis and compression on each of the first model subgraph tasks by using a compressed sensing technology to obtain a compressed first model subgraph task; According to the network bandwidth and delay information of each of the collaborative devices, the encoding strategy of each of the collaborative devices is determined, so that the compressed first model subgraph task is transmitted to the corresponding collaborative device according to the encoding strategy.
[0036] In the present invention, the data in each of the first model subgraph tasks is analyzed and compressed by compressed sensing technology. This technology can reduce the amount of data storage and transmission by utilizing the sparse characteristics of data while ensuring data integrity. During the compression process, the system performs sparsity analysis on the intermediate calculation results, finds the sparse representation in the data, and then reconstructs the original data through a small number of measurement values, thereby achieving efficient data compression.
[0037] The encoding strategy of each device is determined based on the network bandwidth and delay information of each of the collaborative devices. The selection of the encoding strategy will take into account the current network conditions, such as bandwidth size, delay level, etc. For example, when the network bandwidth is sufficient, a high-precision encoding method is used to ensure data quality; when the network bandwidth is tight, a low-precision encoding method is used to reduce the amount of data. At the same time, in order to further reduce the transmission delay, the encoding parameters and framework will be adjusted according to the delay information during adaptive encoding, so that the encoded data can reach the target device as soon as possible during the transmission process.
[0038] According to the encoding strategy, the compressed first model subgraph task is transmitted to the corresponding collaborative device. After the data is compressed and encoded, the transmission volume is greatly reduced, and the transmission can be completed faster under limited network bandwidth. In addition, through the adaptive encoding strategy, it is possible to maintain good transmission efficiency and data quality in different network environments, thereby improving the data transmission efficiency and computing performance of the entire system.
[0039] Optionally, the task assignment adjustment of the first model sub-graph tasks on each collaborative device based on the deep reinforcement learning algorithm to obtain the second model sub-graph tasks corresponding to each collaborative device includes: Input the real-time status data of each collaborative device into the deep reinforcement learning model in the deep reinforcement learning algorithm, and output the task assignment adjustment strategy corresponding to each collaborative device; wherein, the real-time status data includes: real-time data of device resource utilization rate, real-time data of device power consumption, and real-time data of device task latency; According to the task assignment adjustment strategy corresponding to each collaborative device, adjust the task assignment of the first model sub-graph tasks corresponding to each collaborative device to obtain the second model sub-graph tasks corresponding to each collaborative device.
[0040] In the present invention, the real-time status data of each collaborative device is input into the deep reinforcement learning model in the deep reinforcement learning algorithm, and the task assignment adjustment strategy corresponding to each collaborative device is output. Among them, the real-time status data includes real-time data of device resource utilization rate, real-time data of device power consumption, and real-time data of device task latency. By constructing a state space, the deep reinforcement learning model can learn and optimize the task scheduling strategy, balance the computing load between devices, and improve resource utilization. The model generates an optimal task assignment strategy according to the current state of the device, such as information on CPU / GPU utilization rate, power consumption, and task queue length, to adapt to the dynamic environmental requirements.
[0041] According to the task assignment adjustment strategy corresponding to each collaborative device, adjust the task assignment of the first model sub-graph tasks corresponding to each collaborative device to obtain the second model sub-graph tasks corresponding to each collaborative device. Under the guidance of the multi-objective optimization strategy, the system comprehensively considers multiple factors such as computing performance, energy consumption, and latency. By designing a multi-objective reward function, it realizes the adaptive scheduling of tasks and the efficient utilization of resources. For example, when the resource utilization rate of a certain device is too high, the system will migrate some tasks to other devices to achieve load balancing; when the power consumption of a certain device is too high or the task latency is too large, the system will re-adjust the task assignment and preferentially assign tasks to devices with high energy efficiency or small latency, thereby optimizing the overall performance of the system.
[0042] Optionally, before the step of inputting the real-time status data of each collaborative device into the deep reinforcement learning model in the deep reinforcement learning algorithm and outputting the task assignment adjustment strategy corresponding to each collaborative device, it further includes: The preset deep Q-network model continuously learns the relationship between the real-time status data of each collaborative device and the task assignment adjustment strategy through the reward function to optimize the deep reinforcement learning model; When the preset training conditions are met, stop training to obtain the deep reinforcement learning model.
[0043] In the present invention, the preset deep Q - network model continuously learns the relationship between the real - time state data of each collaborative device and the task allocation adjustment strategy through the reward function to optimize the deep reinforcement learning model. During the model training process, the system collects the real - time state data of the device (such as resource utilization rate, power consumption, task latency, etc.) and the effect feedback of the corresponding task allocation adjustment strategy through interaction with the environment. The reward function evaluates and scores each task allocation adjustment strategy according to indicators such as task execution efficiency, energy consumption, and latency. The deep Q - network model continuously adjusts its own parameters based on these reward values, learns a better task allocation strategy, and thus optimizes the performance of the model.
[0044] When the preset training conditions are met, the training is stopped to obtain the deep reinforcement learning model. The preset training conditions can be the convergence conditions of the model. For example, the value of the reward function changes less than a set threshold in several consecutive iterations; it can also be that the number of training rounds reaches the set maximum value, or the performance on the validation set reaches the expected goal, etc. When one of these conditions is met, it is considered that the model has been sufficiently trained and the training process can be stopped. The deep reinforcement learning model obtained at this time has good generalization ability and adaptability, and can output the optimal task allocation adjustment strategy according to the real - time state data of the collaborative device, providing a decision basis for subsequent task scheduling.
[0045] The reward function is specifically: where Rt is the reward value, Ut is the real - time data of the device resource utilization rate, Pt is the real - time data of the device power consumption, Lt is the real - time data of the device task latency, λ1 , λ2 , λ3 are the corresponding weight coefficients respectively.
[0046] In an optional embodiment, the embodiment of the present invention further describes the graph segmentation algorithm. The AI model usually consists of multiple computing layers, such as convolutional layers, fully - connected layers, and pooling layers, etc. The traditional model layering method is usually based on static rules and is difficult to adapt to dynamic changes in device resources and network conditions. Therefore, the present invention proposes an improved graph segmentation algorithm for intelligent layering and task mapping of the model in a heterogeneous device environment.
[0047] First, represent the AI model as a directed acyclic graph (DAG), where each node represents a computing layer and the edge represents the data dependency relationship between layers.
[0048] Using an improved spectral clustering algorithm or the Kernighan-Lin algorithm, the DAG is divided into multiple subgraphs (i.e., computing tasks). When dividing, factors such as the computational complexity of nodes (e.g., FLOPs), the computing power of devices, the energy consumption model, and communication latency are considered to ensure that tasks can be mapped to the most suitable devices for operation.
[0049] The goal of hierarchical computing is to maximize computing efficiency and resource utilization while ensuring model accuracy. By allocating computing tasks to different edge devices, the total computing time and energy consumption can be reduced, achieving collaborative acceleration across devices.
[0050] The basic process of the improved graph segmentation algorithm is as follows: Graph construction: The AI model is constructed as a directed acyclic graph (DAG), where nodes represent computing layers and edges represent data dependencies.
[0051] Node weight calculation: Calculate weights for each node, and the weight values reflect the complexity of the computing layer (e.g., FLOPs) and resource requirements.
[0052] Edge weight calculation: Calculate weights for each edge, and the weight values reflect the data transfer volume and dependency relationship between layers.
[0053] Graph segmentation: Based on the weights of nodes and edges, the graph is segmented into several subgraphs through an improved spectral clustering algorithm or the Kernighan-Lin algorithm. Each subgraph represents a computing task, and the data dependency between tasks is minimized, thereby reducing the communication overhead across devices.
[0054] More specifically, in the embodiments of the present invention, the model will be automatically represented as a directed acyclic graph (DAG), and the model will be intelligently layered through an improved graph segmentation algorithm. The system intelligently divides the model into multiple computing tasks according to the computing power, energy consumption characteristics, and network conditions of devices, and allocates these tasks to the most suitable edge devices.
[0055] For example, in the monitoring system of a smart city, multiple cameras collect video data in real time and perform preliminary processing through edge computing devices, such as object detection and motion tracking. Through the hierarchical computing and dynamic task scheduling mechanism of the present invention, the system can allocate simple tasks (such as image preprocessing) to edge devices with lower computing power, while allocating complex tasks (such as behavior analysis and anomaly detection) to higher-performance devices or the cloud for processing, thereby optimizing the response speed and energy efficiency of the overall system.
[0056] In urban traffic management, edge devices are responsible for real-time monitoring of road condition data and performing preliminary analysis, such as traffic flow statistics and accident detection. Through the energy efficiency optimization algorithm of the present invention, the system can ensure real-time performance while saving energy, providing effective decision support for traffic management.
[0057] In an industrial manufacturing environment, edge devices monitor the operating status of production lines in real time for fault detection and quality control. Through the hierarchical computing and task scheduling mechanism of the present invention, the system can process a large amount of monitoring data in real time and complete preliminary analysis and fault warning locally. Complex diagnosis and optimization tasks can be completed by more powerful computing devices or cloud servers, thereby improving production efficiency and product quality.
[0058] The system predicts potential faults and performs preventive maintenance by collecting and analyzing the operating data of devices in real time. The system utilizes the intelligent data transmission optimization strategy of the present invention to reduce the occupancy of network bandwidth while ensuring data transmission efficiency, providing guarantee for the continuous and efficient operation of industrial devices.
[0059] In autonomous driving, in-vehicle edge devices need to process sensor data and make decisions in real time. Through the cross-device hierarchical computing mechanism of the present invention, simple perception tasks such as obstacle detection can be performed on in-vehicle devices, while complex path planning and environment modeling tasks can be completed through cooperation with roadside units or cloud servers to ensure the real-time and accuracy of driving decisions.
[0060] In a vehicle networking environment, vehicles need to continuously exchange information to maintain coordination. The intelligent data transmission optimization strategy of the present invention can reduce data transmission delay, improve the cooperation efficiency between vehicles, and further enhance the safety and smoothness of the overall traffic system.
[0061] Through the intelligent hierarchical and dynamic scheduling mechanism, this system can make full use of the computing resources of heterogeneous devices and significantly improve the inference efficiency of AI models.
[0062] Based on the optimization strategy of the global energy efficiency model, the system significantly reduces the overall energy consumption while performing high-performance tasks, and extends the battery life of the device.
[0063] Through intelligent data compression and adaptive transmission optimization, the system reduces the data transmission delay between devices and improves the real-time performance and response speed.
[0064] Optionally, Figure 2 The system architecture diagram of the overall technical solution provided by the present invention is as Figure 2As shown, it includes AI model layering and task mapping, data compression and transmission optimization, global energy efficiency optimization and collaborative computing, dynamic task scheduling and load balancing, and real-time monitoring and feedback. The AI model layering and task mapping module divides the AI model into multiple sub-graph tasks and collaborates with the data compression and transmission optimization module to compress and encode the tasks. The dynamic task scheduling and load balancing module dynamically adjusts task allocation according to the real-time status of the devices through deep reinforcement learning technology, and closely collaborates with the AI model layering, data compression, global energy efficiency optimization, and real-time monitoring modules. The global energy efficiency optimization and collaborative computing module evaluates the device performance based on the energy efficiency model, dynamically adjusts task allocation, and interacts with the dynamic task scheduling and real-time monitoring modules. The real-time monitoring and feedback module continuously monitors the system operation status, and dynamically adjusts the task scheduling and data transmission strategies through the feedback mechanism to ensure the efficient operation of the system and energy efficiency optimization Figure 3 The flowchart of dynamic task scheduling and load balancing provided by the present invention is shown, presenting a task scheduling flowchart based on the Deep Q-Network (DQN). The process starts from the "Start" node and sequentially goes through the following steps: Monitor the device resource status: The system first monitors the resource status of each device to obtain the current resource usage.
[0065] Determine the state space and action space: According to the monitored device resource status, determine the state space and action space in the deep Q-network. The state space represents all possible states that the system may be in, and the action space represents all possible actions that can be executed in each state.
[0066] Calculate the task requirements: The system calculates the current task requirements to be processed to determine the computing resources to be allocated.
[0067] Execute the Deep Q-Network (DQN) model: Utilize the deep Q-network model to select the optimal action (i.e., the task allocation strategy) according to the current state and task requirements.
[0068] Update the task allocation strategy: According to the output of the DQN model, update the task allocation strategy to optimize resource utilization and task execution efficiency.
[0069] Complete the task: Execute the updated task allocation strategy to complete the current task scheduling.
[0070] End: The process ends.
[0071] Through continuous learning and optimization of the deep Q-network model, the entire process dynamically adjusts the task allocation strategy to improve the resource utilization rate and task execution efficiency of the system.
[0072] Figure 4 The flowchart of the training and inference of the deep reinforcement learning model provided by the present invention is asFigure 4 As shown below: Collect device status data: Collect the status data of the device, which is used to determine the current status; Define the state space and action space: Define the state space and action space. The state space represents all possible states that the system may be in, and the action space represents all actions that can be executed in each state.
[0073] Set the reward function: Set the reward function, which is used to evaluate the quality of each action and guide the model to learn the optimal policy; Use the collected state data and the set reward function to train the deep Q-network model. The model learns the relationship between states and actions through interaction with the environment to maximize the cumulative reward.
[0074] Task inference and execution: Use the trained DQN model for task inference and execution. Select the optimal action (task allocation strategy) according to the current state and execute the corresponding task.
[0075] Feedback optimization: During the task execution process, collect feedback information for optimizing the task allocation strategy. The feedback information may include the results of task execution, system performance metrics, etc.
[0076] Optimize the task allocation strategy: Adjust and optimize the task allocation strategy according to the feedback information to improve the overall performance of the system.
[0077] Figure 5 For the data compression and transmission optimization flowchart provided by the present invention, first analyze the sparsity of the data to determine which parts of the data are sparse and which are dense. This step helps with subsequent compression operations.
[0078] According to the sparsity of the data, apply compressive sensing technology to compress the data. Compressive sensing technology can significantly reduce the amount of data while maintaining the main features of the data.
[0079] Select a suitable coding strategy to further optimize the compressed data. The coding strategy will be adaptively adjusted according to the characteristics of the data and transmission requirements.
[0080] Transmit the compressed and encoded data. This step ensures that the bandwidth and storage space occupied by the data during transmission are minimized.
[0081] At the receiving end, decode and reconstruct the transmitted compressed data to recover the original data or a form close to the original data.
[0082] Figure 6 For the global energy efficiency optimization schematic diagram described in the present invention of the present invention, start: The process starts.
[0083] Collect energy efficiency data of devices: Collect the energy efficiency data of each device, which is used to evaluate the energy efficiency performance of the device.
[0084] Build a global energy efficiency model: Build a global energy efficiency model based on the collected energy efficiency data. This model is used to evaluate and optimize the energy efficiency of the entire system.
[0085] Evaluate the energy efficiency of devices: Use the global energy efficiency model to evaluate the energy efficiency performance of each device and determine the energy efficiency of the device under the current task allocation.
[0086] Optimize the task allocation strategy: Optimize the task allocation strategy according to the energy efficiency evaluation results of the devices. The goal of optimization is to improve the energy efficiency of the overall system and reduce energy consumption.
[0087] Adjust the task allocation in real time: Adjust the task allocation of each device in real time according to the optimized task allocation strategy. Ensure that the task allocation remains optimal in a dynamic environment.
[0088] End: The process ends.
[0089] Figure 7 This is the multi-objective optimization flowchart for task scheduling and energy efficiency optimization provided by the present invention. Start: The process starts from the "Start" node.
[0090] Calculate the task requirements: First, calculate the current task requirements, including the amount of computation, data volume, etc.
[0091] Evaluate the device resources and energy efficiency: Evaluate the resource status (such as CPU, memory, network bandwidth, etc.) and energy efficiency performance (such as power consumption, heat dissipation, etc.) of each device.
[0092] Select a task scheduling strategy: Select a suitable task scheduling strategy according to the task requirements and the evaluation results of the device resources and energy efficiency.
[0093] Optimize the computing time: Reduce the computing time of the task and improve the computing efficiency by optimizing the task scheduling strategy.
[0094] Optimize the task real-time performance: Ensure that the task can be completed in real time and meet the real-time requirements.
[0095] Optimize the system energy consumption: Reduce the overall energy consumption of the system and improve the energy efficiency by optimizing the task scheduling strategy.
[0096] Determine the optimal scheduling strategy: Comprehensively consider the computing time, task real-time performance, and system energy consumption to determine the optimal task scheduling strategy.
[0097] End: The process ends.
[0098] Figure 8Flowchart of the real-time monitoring and feedback control system provided by the present invention. Real-time monitoring of device status: The system monitors the status of each device in real time to obtain the latest device operation information.
[0099] Collect device resource data: Collect the resource data of the device, including CPU usage rate, memory occupancy, network bandwidth, etc.
[0100] Analyze data and evaluate: Analyze and evaluate the collected data to determine the performance of the current system and the task execution situation.
[0101] Feedback and adjust the scheduling strategy: According to the analysis results, feedback and adjust the task scheduling strategy to optimize the task execution efficiency.
[0102] Optimize task execution: According to the adjusted scheduling strategy, optimize the task execution to improve the task completion speed and quality.
[0103] Continuous monitoring and feedback: Continuously monitor the task execution situation and further optimize the task scheduling strategy according to the feedback information.
[0104] Figure 9 Schematic diagram of the task fault tolerance and recovery mechanism provided by the present invention. Task allocation: Allocate tasks to the corresponding execution units or devices.
[0105] Execute tasks: Execute the allocated tasks.
[0106] Monitor task execution status: During the task execution process, monitor the task execution status in real time.
[0107] Fault detected: If a fault occurs in the task execution during the monitoring process, enter the fault handling process.
[0108] Trigger the fault tolerance mechanism: Once a fault is detected, trigger the fault tolerance mechanism.
[0109] Re-allocate tasks: Re-allocate the faulty tasks to other available execution units or devices.
[0110] Resume task execution: Resume the task execution on the new execution unit or device.
[0111] Task completed: The task execution is completed.
[0112] Figure 10 Schematic diagram of cross-device collaborative optimization provided by the present invention. As Figure 10 shown, select the computing task: Select a suitable computing task according to the task requirements and device status.
[0113] Device 1: The task is allocated to Device 1 for computing.
[0114] Optimized Task Execution: Device 1 performs optimizations during task execution to improve computational efficiency and performance.
[0115] Data Transmission and Collaboration: Device 1 conducts data transmission and collaborative work with other devices (Device 2 and Device 3) to ensure the smooth execution of tasks.
[0116] Task Completion: The task is completed on Device 1.
[0117] Figure 11 The balanced optimization diagram of energy efficiency and performance provided by the present invention is as Figure 11 described. Evaluate device energy efficiency: Evaluate the energy efficiency performance of the device to determine the current energy efficiency status of the device.
[0118] Calculate performance requirements: Based on the evaluation results, calculate the performance requirements of the current task.
[0119] Select optimization strategy: According to the energy efficiency and performance requirements, select a suitable optimization strategy.
[0120] Energy efficiency optimization: If the selected strategy is energy efficiency optimization, perform energy efficiency optimization measures.
[0121] Performance optimization: If the selected strategy is performance optimization, perform performance optimization measures.
[0122] Balance energy efficiency and performance: Find a balance between energy efficiency and performance to ensure that the system can meet the performance requirements while maximizing energy efficiency.
[0123] Strategy implementation: Implement the selected optimization strategy into the system.
[0124] The AI model collaborative acceleration and energy efficiency optimization device provided by the present invention will be described below. The AI model collaborative acceleration and energy efficiency optimization device described below can be mutually corresponding and referenced to the AI model collaborative acceleration and energy efficiency optimization method described above.
[0125] Figure 12 The structural schematic diagram of the AI model collaborative acceleration and energy efficiency optimization device provided by the present invention is as Figure 12 shown, including: The partitioning module 1210 is used to construct a directed acyclic graph for the AI model data to be processed, and based on the graph segmentation algorithm, according to the device information of each collaborative device, divide the directed acyclic graph into multiple model subgraphs, and then obtain the first model subgraph tasks corresponding to each collaborative device; wherein, the nodes of the directed acyclic graph represent computational layers, and the edges of the directed acyclic graph represent data dependencies; The adjustment module 1220 is used to perform task assignment adjustment on the first model subgraph tasks on each collaborative device based on the deep reinforcement learning algorithm to obtain the second model subgraph tasks corresponding to each collaborative device; The optimization model 1230 is used to optimize the task energy efficiency of the second model subgraph tasks corresponding to each of the collaborative devices based on the global energy efficiency models and multi-objective integer programming algorithms corresponding to the respective collaborative devices, obtain the third model subgraph tasks for each of the collaborative devices, and run the corresponding third model subgraph tasks on each of the collaborative devices.
[0126] In the present invention, by dividing the AI model into multiple subgraphs and performing task allocation according to device information, the computing resources of each collaborative device can be fully utilized, improving the computing efficiency. At the same time, this division method takes into account the data dependency relationship, ensuring the correctness of the calculation and the consistency of the data. Through the dynamic adjustment of the DRL model, the task allocation can adapt to changes in the system state, ensuring the efficient utilization of computing resources. The design of the reward function enables the model to balance multiple objectives while optimizing the computing performance, taking into account both energy consumption and latency. Through the optimization of the global energy efficiency model and multi-objective integer programming algorithm, the task allocation not only considers the computing efficiency but also takes into account energy consumption and energy efficiency. This optimization method can significantly reduce the overall energy consumption while ensuring the system performance, extending the battery life of the device.
[0127] Figure 13 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 13 shown, the electronic device may include: a processor 1310, a communication interface 1320, a memory 1330, and a communication bus 1340. Among them, the processor 1310, the communication interface 1320, and the memory 1330 communicate with each other through the communication bus 1340. The processor 1310 can call the logical instructions in the memory 1330 to execute the AI model collaborative acceleration and energy efficiency optimization method, which includes: constructing the to-be-processed AI model data into a directed acyclic graph, and based on the graph segmentation algorithm, according to the device information of each collaborative device, dividing the directed acyclic graph into multiple model subgraphs to obtain the first model subgraph tasks corresponding to each collaborative device; wherein, the nodes of the directed acyclic graph represent computing layers, and the edges of the directed acyclic graph represent data dependencies; Based on the deep reinforcement learning algorithm, adjust the task allocation of the first model subgraph tasks on each of the collaborative devices to obtain the second model subgraph tasks corresponding to each of the collaborative devices; Based on the global energy efficiency models and multi-objective integer programming algorithms corresponding to the respective collaborative devices, optimize the task energy efficiency of the second model subgraph tasks corresponding to the respective collaborative devices, obtain the third model subgraph tasks for each of the collaborative devices, and run the corresponding third model subgraph tasks on each of the collaborative devices.
[0128] In addition, when the logical instructions in the above-mentioned memory 1330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0129] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the AI model collaborative acceleration and energy efficiency optimization method provided by the above-mentioned various methods. The method includes: constructing the AI model data to be processed into a directed acyclic graph, and based on a graph partitioning algorithm, according to the device information of each collaborative device, dividing the directed acyclic graph into multiple model subgraphs to obtain the first model subgraph tasks corresponding to each collaborative device; wherein, the nodes of the directed acyclic graph represent computing layers, and the edges of the directed acyclic graph represent data dependencies. Based on a deep reinforcement learning algorithm, task assignment adjustment is performed on the first model subgraph tasks on each collaborative device to obtain the second model subgraph tasks corresponding to each collaborative device. Based on the global energy efficiency models corresponding to each of the collaborative devices and a multi-objective integer programming algorithm, task energy efficiency optimization is performed on the second model subgraph tasks corresponding to each of the collaborative devices to obtain the third model subgraph tasks corresponding to each of the collaborative devices, and the corresponding third model subgraph tasks are run on each of the collaborative devices.
[0130] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the AI model collaborative acceleration and energy efficiency optimization method provided by the above-mentioned various methods. The method includes: constructing the AI model data to be processed into a directed acyclic graph, and based on a graph partitioning algorithm, according to the device information of each collaborative device, dividing the directed acyclic graph into multiple model subgraphs to obtain the first model subgraph tasks corresponding to each collaborative device; wherein, the nodes of the directed acyclic graph represent computing layers, and the edges of the directed acyclic graph represent data dependencies. Based on the deep reinforcement learning algorithm, the first model sub-graph tasks on each collaborative device are adjusted for task allocation to obtain the second model sub-graph tasks corresponding to each collaborative device; Based on the global energy efficiency models corresponding to each of the collaborative devices and the multi-objective integer programming algorithm, the task energy efficiency of the second model sub-graph tasks corresponding to each of the collaborative devices is optimized to obtain the third model sub-graph tasks corresponding to each of the collaborative devices, and the corresponding third model sub-graph tasks are run on each of the collaborative devices.
[0131] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0132] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An AI model collaborative acceleration and energy efficiency optimization method, characterized in that, Including: Construct the AI model data to be processed into a directed acyclic graph, and based on the graph segmentation algorithm, according to the device information of each collaborative device, divide the directed acyclic graph into multiple model sub-graphs, and then obtain the first model sub-graph task corresponding to each collaborative device; wherein, the nodes of the directed acyclic graph represent computing layers, and the edges of the directed acyclic graph represent data dependencies; Based on the deep reinforcement learning algorithm, perform task allocation adjustment on the first model sub-graph tasks on each collaborative device to obtain the second model sub-graph tasks corresponding to each collaborative device; Based on the global energy efficiency model and multi-objective integer programming algorithm corresponding to each collaborative device, perform task energy efficiency optimization on the second model sub-graph tasks corresponding to each collaborative device to obtain the third model sub-graph tasks corresponding to each collaborative device, and run the corresponding third model sub-graph tasks on each collaborative device.
2. The AI model collaborative acceleration and energy efficiency optimization method according to claim 1, wherein After the step of performing task energy efficiency optimization on the second model sub-graph tasks corresponding to each collaborative device based on the global energy efficiency model and multi-objective integer programming algorithm corresponding to each collaborative device to obtain the third model sub-graph tasks corresponding to each collaborative device, and running the corresponding third model sub-graph tasks on each collaborative device, it further includes: Obtain the device information fed back by each collaborative device during the running of the corresponding third model sub-graph task; wherein, the device information includes at least one of: device computing power consumption information, device resource utilization information, and device heat dissipation information; Dynamically adjust the model sub-graph tasks corresponding to each collaborative device according to the device information fed back by each collaborative device during the running of the third model sub-graph task.
3. The AI model collaborative acceleration and energy efficiency optimization method according to claim 1, characterized in that, After the step of constructing the AI model data to be processed into a directed acyclic graph, and based on the graph segmentation algorithm, according to the device information of each collaborative device, dividing the directed acyclic graph into multiple model sub-graphs to obtain the first model sub-graph task corresponding to each collaborative device, it further includes: Perform sparsity analysis and compression on each of the first model sub-graph tasks through compressive sensing technology to obtain the compressed first model sub-graph tasks; Determine the coding strategy for each collaborative device according to the network bandwidth and latency information of each collaborative device, and transmit the compressed first model sub-graph tasks to the corresponding collaborative device according to the coding strategy.
4. The AI model collaborative acceleration and energy efficiency optimization method according to claim 1, characterized in that, The step of performing task allocation adjustment on the first model sub-graph tasks on each collaborative device based on the deep reinforcement learning algorithm to obtain the second model sub-graph tasks corresponding to each collaborative device includes: Input the real-time state data of each collaborative device into the deep reinforcement learning model in the deep reinforcement learning algorithm, and output the task allocation adjustment strategy corresponding to each collaborative device; wherein, the real-time state data includes: real-time device resource utilization data, real-time device power consumption data, real-time device task latency data; Perform task allocation adjustment on the first model sub-graph tasks corresponding to each collaborative device according to the task allocation adjustment strategy corresponding to each collaborative device to obtain the second model sub-graph tasks corresponding to each collaborative device.
5. The AI model collaborative acceleration and energy efficiency optimization method according to claim 4, wherein Before the step of inputting the real-time status data of each collaborative device into the deep reinforcement learning model of the deep reinforcement learning algorithm and outputting the task allocation adjustment strategy corresponding to each collaborative device, the method further includes: The preset deep Q-network model continuously learns the relationship between the real-time status data of each collaborative device and the task allocation adjustment strategy through a reward function to optimize the deep reinforcement learning model; When the preset training conditions are met, stop training to obtain the deep reinforcement learning model.
6. The AI model collaborative acceleration and energy efficiency optimization method according to claim 5, wherein The reward function is specifically: Among them, Rt is the reward value, Ut is the real-time data of the device resource utilization rate, Pt is the real-time data of the device power consumption, Lt is the real-time data of the device task delay, λ1 , λ2 , λ3 are the corresponding weight coefficients respectively.
7. An AI model collaborative acceleration and energy efficiency optimization device, characterized in that, including: A partitioning module for constructing the data of the AI model to be processed into a directed acyclic graph, and based on a graph segmentation algorithm, dividing the directed acyclic graph into multiple model subgraphs according to the device information of each collaborative device, so as to obtain the first model subgraph tasks corresponding to each collaborative device; wherein, the nodes of the directed acyclic graph represent computing layers, and the edges of the directed acyclic graph represent data dependencies; An adjustment module for adjusting the task allocation of the first model subgraph tasks on each collaborative device based on the deep reinforcement learning algorithm to obtain the second model subgraph tasks corresponding to each collaborative device; An optimization model for optimizing the task energy efficiency of the second model subgraph tasks corresponding to each collaborative device based on the global energy efficiency model and the multi-objective integer programming algorithm corresponding to each collaborative device, to obtain the third model subgraph tasks corresponding to each collaborative device, and running the corresponding third model subgraph tasks on each collaborative device.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the AI model collaborative acceleration and energy efficiency optimization method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the AI model collaborative acceleration and energy efficiency optimization method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the AI model collaborative acceleration and energy efficiency optimization method according to any one of claims 1 to 6.