Multi-device cooperative task execution method and device and readable storage medium

By sharding and deploying large models on multiple edge computing devices and using Monte Carlo tree search and load-aware strategies to optimize task allocation, the problems of large model processing accuracy and efficiency on edge computing devices are solved, and efficient parallel data processing is achieved.

CN120745816AActive Publication Date: 2025-10-03JIANGNAN UNIV

Patent Information

Application Number
CN202510843051.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

When large models are deployed on edge computing devices in existing technologies, it is impossible to balance data processing accuracy and efficiency, resulting in the inability to execute tasks efficiently and accurately.

Method used

Divide large models into multiple model shards and execute them collaboratively on multiple edge computing devices. Optimize task allocation through the Monte Carlo tree search algorithm and load-aware upper confidence bound strategy to achieve efficient collaborative processing of model shards.

Benefits of technology

Through multi-device collaborative processing, the accuracy and efficiency of data detection and recognition tasks are improved, the problem of limited resources of a single edge computing device is solved, and efficient parallel data processing is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745816A_ABST
    Figure CN120745816A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of edge reasoning, and relates to a multi-device cooperative task execution method and device and a readable storage medium. Dividing a large model used for realizing a to-be-executed task into M model fragments which are executed in sequence; loading and calculating each model fragment on mutually communicated edge calculation equipment, and obtaining loading time, calculation time and communication time delay to construct a simulation model of the task to be executed; using a Monte Carlo tree search algorithm, a load awareness upper confidence limit strategy and the to-be-executed task simulation model, based on the execution sequence of the model fragments, performing iterative search on the M model fragments in sequence to allocate edge computing devices until a preset iterative convergence condition is reached, and obtaining a target task allocation strategy; based on the target task distribution strategy, the M model fragments are deployed on the corresponding edge computing devices respectively, so that the to-be-executed task is cooperatively and efficiently achieved through the model fragments on the multiple edge computing devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of edge reasoning technology, and in particular to a method, apparatus, and computer-readable storage medium for executing collaborative tasks across multiple devices. Background Art

[0002] With the rapid development of IoT devices and deep learning technology, deep learning models, leveraging the powerful representational capabilities of multi-layer neural networks, are widely used in scenarios such as computer vision, natural language processing, and intelligent control. Examples include accurate lesion identification in medical images, real-time analysis of complex road conditions by self-driving cars, intelligent recommendations on e-commerce platforms, and semantic understanding of user commands by smart home devices. Among various deep learning models, large deep learning models (Transformer) architectures based on self-attention mechanisms significantly improve their ability to learn from input data when processing large amounts of data. For example, large models such as the GPT series, BERT, and T5, with their superior contextual understanding and generation capabilities, can produce results that are more consistent with the actual data in applications such as object recognition, image processing, and intelligent question-answering. Therefore, in high-precision, time-sensitive applications such as production equipment status and data monitoring and intelligent vehicle control within the Industrial Internet of Things, large models are often deployed on edge computing devices. By inputting the text or image data to the edge computing devices, the data processing capabilities of the large models are combined with the efficient data processing capabilities of edge computing to achieve low-latency data processing and closed-loop decision-making.

[0003] However, edge computing devices are limited by bottlenecks in computing, storage, and memory. Taking memory as an example, a typical edge computing device such as the Raspberry Pi 5 has a memory capacity of only 8GB. In contrast, the parameter scale of large models has exploded in recent years. For example, the GPT-3 model contains 175 billion trainable parameters, and the parameter scale of the GPT-4 model is estimated to have reached the trillion level, almost dozens of times that of GPT-3. In particular, when large models are applied to scenarios with multi-source heterogeneous data characteristics such as industrial Internet of Things data monitoring, intelligent traffic control, and power grid energy scheduling, the data processing volume of the large models is greatly increased. This means that the memory usage of mainstream large models often reaches tens or even hundreds of GB, far exceeding the memory capacity of general edge computing devices, making loading and running large models on these edge computing devices a daunting task.

[0004] To solve this problem, two methods have been proposed in the existing technology: one is to reduce the computational complexity of large models, reduce storage requirements and model complexity through model optimization technology, which specifically includes model pruning, compression and quantization. By removing some neurons, connections or hierarchical structures in the large model, or encoding or compressing the model parameters into a smaller form, or converting the floating-point parameters in the model into low-precision integer values, the optimized large model can be deployed on edge computing devices. However, the simplification of the model structure will inevitably reduce the model's ability to learn input data, which also leads to the optimized large model being unable to fully extract the feature information of the input data when processing input information such as images and text, thereby outputting low-precision processing results. Another method is to split the large model into multiple model shards, each model shard undertakes a part of the computing task, and then deploy the model shards on the edge computing device. Since the model's processing of input data is divided into serial small-scale data processing processes, the edge computing device only needs to load and process the parameters of a single model shard in a single data processing process, which reduces the single calculation amount of the edge computing model and avoids the resource consumption of the complex network structure in the large model for simultaneous processing of input data. However, this strategy of exchanging time for space will reduce the data processing rate and cannot output data processing results in a timely manner, resulting in an inability to meet the high timeliness requirements of data processing in scenarios such as intelligent traffic control and production data monitoring.

[0005] In summary, when using large models to process input data such as images and text in the existing technology, it is impossible to take into account both data processing accuracy and data processing efficiency, which makes it impossible to perform tasks such as data detection and recognition based on large models efficiently and accurately. Summary of the Invention

[0006] To this end, the technical problem to be solved by the present invention is to overcome the problem in the existing technology that when using large models to process input data such as images and texts, it is impossible to take into account both data processing accuracy and data processing efficiency, thereby making it impossible to perform tasks such as data detection and recognition based on large models efficiently and accurately.

[0007] To solve the above technical problems, the present invention provides a multi-device collaborative task execution method, comprising: S10: Build a large model based on the task to be executed, and divide the large model into M model slices for sequentially processing the task to be executed; S20: Load and calculate each model slice on each interconnected edge computing device, obtain the loading time, calculation time and communication delay of each model slice on each edge computing device, and thus build a simulation model of the task to be executed; S30: Using the Monte Carlo tree search algorithm and the load-aware upper confidence bound strategy, based on the execution order of the model shards and the cumulative reward values ​​of the task allocation strategies of the previous n-1 iterations, iteratively search for edge computing devices for the M model shards in turn to obtain the task allocation strategy of the nth iteration; where n ≥ 1, when n = 1, the reward value of the task allocation strategy of the n-1th iteration is a preset value; S40: Using the simulation model of the task to be executed to simulate the task allocation strategy of the nth iteration, obtain the reward value of the task allocation strategy of the nth iteration, update n=n+1 and return to step S30 until the preset iteration convergence condition is reached to obtain the target task allocation strategy; S50: Based on the target task allocation strategy, the M model shards are respectively deployed on the corresponding edge computing devices, so as to utilize the model shards on multiple edge computing devices to collaboratively implement the tasks to be executed.

[0008] Preferably, when the task to be performed is an intelligent question-answering task, the large model is an intelligent question-answering robot model, the input of the large model is the query text, and the output of the large model is the answer text; When the task to be performed is an image classification task, the large model is an image classification model, the input of the large model is the image to be classified, and the output of the large model is the image category prediction probability.

[0009] Preferably, step S30 includes: S300: Initialize m=1; S301: Using the load-aware upper confidence bound strategy, based on the cumulative reward value of each edge computing device under the task allocation strategy of the previous n-1 iterations, calculate the priority of allocating the mth model shard to each edge computing device, thereby allocating the mth model shard; S302: Update m=m+1 and return to step S301. If, based on the allocation results of the first m-1 model shards, there is an edge computing device that has not been selected for allocation of the m-th model shard in the first n-1 iterations, the m-th model shard is allocated to the edge computing device. S303: Randomly assign edge computing devices to the m+1th to Mth model shards, thereby obtaining the task allocation strategy for the nth iteration based on the allocation results of the 1st to Mth model shards.

[0010] Preferably, step S301 includes: The number of times each edge computing device was selected under the task allocation strategy in the previous n-1 iterations, the cumulative reward value, and the current load of each edge computing device are input into the load-aware upper confidence limit calculation formula to obtain the priority of allocating the m-th model shard to each edge computing device; Assign the mth model shard to the edge computing device with the highest priority.

[0011] Preferably, the load-aware upper confidence limit calculation formula is expressed as: , in, The formula for calculating the load-aware upper confidence limit is: Indicates the task allocation strategy for the first n-1 iterations. The cumulative reward value of edge computing devices; Indicates the task allocation strategy for the first n-1 iterations. The number of times an edge computing device is selected; Indicates that under the task allocation strategy of the first n-1 iterations, when The total number of times an edge computing device is selected as a child node; Indicates that the mth model shard is assigned to the edge computing devices; Indicates that the mth model shard is assigned to the After the edge computing device The load of edge computing devices; 、 represents the preset positive constant hyperparameter.

[0012] Preferably, step S40 includes: The simulation model of the task to be executed is used to simulate the actual loading time, actual computing time, and actual communication delay of each edge computing device under the task allocation strategy of the nth iteration; The actual loading time, actual computing time, and actual communication delay of each edge computing device are converted to obtain the loading reward, computing reward, and communication delay reward of each edge computing device; The loading reward, computing reward, and communication delay reward of each edge computing device are weighted and summed to obtain the reward value of each edge computing device under the task allocation strategy of the nth iteration.

[0013] Preferably, the calculation formula of the reward value of each edge computing device under the task allocation strategy of the nth iteration is expressed as: , in, Indicates the The reward value of an edge computing device; Represents the conversion function from time to reward; Indicates the The actual loading time of each edge computing device; express The weight of Indicates the The actual computing time of each edge computing device; express The weight of Indicates the The actual communication delay of each edge computing device; express The weight of .

[0014] Preferably, before step S30, the following steps are further included: Based on the preset loading time and preset calculation time of each model shard, as well as the loading time and calculation time of each edge computing device for each model shard, a set of available edge computing devices for each model shard is constructed, so as to iteratively search for edge computing devices for M model shards in the set of available edge computing devices for each model shard.

[0015] The present invention also provides a multi-device collaborative task execution device, comprising: The model building and sharding module is used to build a large model based on the tasks to be executed, and divide the large model into M model shards for sequentially processing the tasks to be executed; The modeling and analysis module is used to load and calculate each model slice on each interconnected edge computing device, obtain the loading time, calculation time and communication delay of each model slice on each edge computing device, and thus build a simulation model of the task to be executed; The model allocation module is used to use the Monte Carlo tree search algorithm and the load-aware upper confidence bound strategy to iteratively search edge computing devices for M model shards in sequence based on the execution order of the model shards and the cumulative reward values ​​of the task allocation strategies of the previous n-1 iterations to obtain the task allocation strategy of the nth iteration; where n ≥ 1, when n = 1, the reward value of the task allocation strategy of the n-1th iteration is the preset value; The environment simulation module is used to simulate the task allocation strategy of the nth iteration using the simulation model of the task to be executed, obtain the reward value of the task allocation strategy of the nth iteration, update n=n+1 and return to execute the steps executed by the iterative allocation module until the preset iterative convergence condition is reached, and obtain the target task allocation strategy; The collaborative task execution module is used to deploy M model shards on corresponding edge computing devices based on the target task allocation strategy, so as to utilize the model shards on multiple edge computing devices to collaboratively implement the tasks to be executed.

[0016] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-mentioned multi-device collaborative task execution method are implemented.

[0017] The multi-device collaborative task execution method provided in this application has the following beneficial effects: This application considers deploying multiple model shards of a large model on multiple edge computing devices in a dispersed manner, and completing the data exchange and data processing process in parallel through collaborative communication between multiple edge computing devices, so as to alleviate the problem that a single edge computing device cannot carry a large model due to resource constraints; specifically, firstly, a large model is built based on the task to be executed (i.e., feature extraction and detection of input data such as images and texts), and the large model is divided into multiple model shards using model sharding technology. A model shard may contain one or more basic units, and a pre-evaluation is performed on each edge computing device and all model shards participating in the execution of the collaborative task, and the loading time, calculation time, and communication delay of each model shard are quantified, thereby building a simulation model based on the performance parameters of each edge computing device, and using the simulation model to accurately Accurately simulate the actual loading, computing and communication delays in the execution of collaborative tasks of multiple devices under different task allocation strategies; further, since there is usually a data dependency between model shards, that is, the operation process of each model shard has obvious timing, the model shard allocation problem can be converted into a timing decision problem. At the same time, considering that each model shard can be deployed on multiple edge computing devices, the allocation decision space of the model shard can be abstracted as a multi-branch tree structure. Based on this feature, this application chooses to use the Monte Carlo tree search algorithm to allocate multiple model shards. On this basis, the priority of each edge computing device in the evaluation allocation process is evaluated based on the load-aware upper confidence limit node selection strategy, thereby achieving load balancing between multiple edge computing devices and avoiding the problem of excessive load on some edge computing devices. The solution provided in this application divides the complete data processing process into multiple small-scale data processing tasks. By allocating these small-scale data processing tasks to different edge computing devices, the natural parallelism of distributed deployment is utilized to achieve high parallelization of data processing. This not only solves the problem that a single edge computing device cannot carry a large model due to limited resources, but also retains the structure and parameter processing performance of the large model, thereby improving the output accuracy and efficiency of tasks such as data detection and recognition based on large models. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein: Figure 1 A comparison chart of tasks performed by a single device and a network of devices provided for this application; Figure 2 Flowchart of the multi-device collaborative task execution method provided by this application; Figure 3 Schematic diagram of the multi-device collaborative task execution principle provided by this application; Figure 4 Schematic diagram of the lightweight simulation framework for large model edge deployment provided by this application; Figure 5 This is a diagram of the task allocation principle based on the Monte Carlo tree search algorithm provided in this application. DETAILED DESCRIPTION

[0019] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0020] See also Figure 1 , Figure 1 Shown is a comparison diagram of the execution of tasks by a single device and a networked device (multi-device collaboration) provided by this application. Since the large model used to perform tasks often requires multi-stage processing of the input data, for example, a large model used for face recognition first needs to preprocess the input face image, then perform feature extraction on the preprocessed image, and finally perform identification and classification based on the extracted features to output the results. If the large model used to implement the task to be executed is deployed in a single device, the device needs to load and run the model architecture corresponding to the current step when executing each step of the model. This sequential execution method results in a longer time for a single task and higher performance requirements such as device resources and operating speed.

[0021] The current mainstream large models generally adopt a unitized and hierarchical structure, which is usually composed of a series of basic units stacked in sequence. These units are regarded as relatively independent computing units, responsible for performing specific tasks such as feature extraction and information processing. If the large model is divided into multiple model shards using model sharding technology, a model shard may contain one or more basic units, and there is usually a data dependency relationship between model shards, which makes the model operation process have obvious time sequence. By deploying the model shards on different edge computing devices, Figure 1 Taking the dual-device collaboration scenario shown in as an example, in the initial stage, device A and device B can load the first model shards they are responsible for in parallel (for example, model shard 1 and model shard 2). When device A completes the calculation of model shard 1, it immediately loads model shard 3. At the same time, device B can calculate model shard 2 based on the calculation result of model shard 1 output by device A, and so on. A single task is executed in a pipeline form. Compared with the traditional single-device task execution method, this multi-device collaborative task execution method greatly reduces the task execution time. Moreover, since the models are distributed and loaded on different edge computing devices, it significantly alleviates the problem that a single device is difficult to carry large models due to resource constraints.

[0022] Based on the above principles, this application provides a multi-device collaborative task execution method, such as Figure 2 As shown, the method specifically includes: S10: Build a large model based on the tasks to be executed, and divide the large model into M model slices for sequentially processing the tasks to be executed.

[0023] S20: Load and calculate each model slice on each interconnected edge computing device, obtain the loading time, calculation time and communication delay of each edge computing device for each model slice, and thus build a simulation model of the task to be executed.

[0024] S30: Using the Monte Carlo tree search algorithm and the load-aware upper confidence bound strategy, based on the execution order of the model shards and the cumulative reward values ​​of the task allocation strategies of the first n-1 iterations, iteratively search for edge computing devices for the M model shards in turn to obtain the task allocation strategy of the nth iteration; where n≥1, when n=1, the reward value of the task allocation strategy of the n-1th iteration is a preset value.

[0025] S40: Use the simulation model of the task to be executed to simulate the task allocation strategy of the nth iteration to obtain the reward value of the task allocation strategy of the nth iteration, update n=n+1 and return to execute step S30 until the preset iteration convergence condition is reached to obtain the target task allocation strategy.

[0026] S50: Based on the target task allocation strategy, the M model shards are respectively deployed on the corresponding edge computing devices, so as to utilize the model shards on multiple edge computing devices to collaboratively implement the tasks to be executed.

[0027] Specifically, the large model of the task to be executed is recorded as , and divide it into Model shards executed sequentially: , and at the same time The connected edge computing devices are denoted as Because of the edge computing devices Execute model sharding on It includes two steps: model shard loading and model shard calculation. Therefore, the loading time and calculation time are recorded as and At the same time, considering the communication delay when the edge computing device transmits the calculation result of the current model fragment to the edge computing device where the next model fragment is located, Represents edge devices and The communication delay between Considering the limited resources of edge computing devices, an edge computing device can only perform one task at any one time: loading or calculating a model shard.

[0028] like Figure 3 The figure shows the principle diagram of multi-device collaborative task execution provided by this application. It can be seen that this application coordinates the memory, computing power, communication and other resources of multiple edge computing devices, and builds a "load-wait-compute-communication" workflow at the level of a single edge computing device from the global perspective of "multi-model sharding-multi-device". It optimizes the overall reasoning process, thereby significantly reducing end-to-end reasoning latency and effectively controlling peak memory usage. Specifically, by collecting the key parameters of the current edge environment, a lightweight simulation environment is constructed, and exploration and search are performed in the task allocation stage based on the simulation environment. The allocation scheme of the model shards is quickly found, and the scheme is applied to the actual deployment environment to drive the collaborative task execution of multiple edge computing devices.

[0029] Specifically, in different application scenarios, the tasks to be performed in step S10 are different, and the constructed large models are also different. When the task to be performed is an intelligent question-answering task, the large model is an intelligent question-answering robot model, the input of the large model is the query text, and the output of the large model is the answer text; When the task to be executed is an image classification task, the large model is an image classification model, which can specifically adopt one of Vision Transformer, MSG-Transformer, and Pyramid Vision Transformer. The input of the large model is the image to be classified, and the output of the large model is the image category prediction probability.

[0030] For example, in an autonomous driving control scenario covering a large area, a large number of high-definition images need to be recognized and detected, and then the control instructions of each autonomous driving vehicle are controlled based on the recognition and detection results. At this time, the large model architecture is complex, with many parameters, and the required memory resources are large. Therefore, the computing power of a single edge computing device is limited. By deploying multiple interconnected edge computing devices, the large model shards can be deployed on multiple edge computing devices, and the image to be detected and recognized can be input into the device where the first model shard is deployed. Multiple devices can then collaboratively process the image to be detected and recognized, and finally output the recognition detection results.

[0031] Furthermore, the main purpose of step S20 is to pre-evaluate each edge computing device and all model shards involved in the collaborative task execution, quantify the loading time, calculation time of each model shard, and the communication delay between two edge computing devices, so that a simulation model can be constructed based on the performance parameters of each edge computing device. This simulation model can be used to accurately simulate the actual loading, calculation and communication delays in the execution of multi-device collaborative tasks under different task allocation strategies, thereby realizing effective simulation of diversified and heterogeneous large-model edge deployment scenarios.

[0032] In addition, not all edge computing devices have the ability to successfully load and run each model shard. Therefore, step S20 can also evaluate the feasibility of deploying each model shard on each edge computing device based on the memory usage of each model shard.

[0033] like Figure 4 Shown is a schematic diagram of a lightweight simulation framework for edge-side deployment of large models provided by this application. The model in this framework is the large model of the task to be executed, the shard is the basic computing unit or layer for building the large model, the device represents the edge computing device used to deploy each shard, and the environment represents the global manager of the entire simulation scenario, which drives the time-based multi-device collaborative task execution simulation process according to the task allocation strategy.

[0034] Furthermore, since there is usually a data dependency between model shards, that is, the execution process of each model shard has obvious temporal sequence, the model shard allocation problem can be transformed into a temporal decision problem, such as Figure 5 As shown in Figure 2, given that each model shard can be deployed on multiple potential edge computing devices, the allocation decision space of the model shard can be abstracted as a multi-branch tree structure.

[0035] The Monte Carlo Tree Search (MCTS) algorithm is a widely used search algorithm in decision-making processes. It cleverly combines the concepts of Monte Carlo simulation and tree search. It is particularly suitable for decision-making problems with incomplete information, large state spaces, or high complexity. Its core concept is to iteratively evaluate the potential value of the current decision based on simulating potential future behaviors and outcomes. This method can efficiently explore and evaluate each branch of the decision tree without the need for an exhaustive search. Therefore, in environments with limited computing resources and time, MCTS demonstrates remarkable efficiency and speed. At the same time, given the search space characteristics and complexity of the model shard allocation problem, this application uses the MCTS algorithm to solve the allocation strategy for multiple model shards.

[0036] The Monte Carlo tree search algorithm consists of four main steps: selection, expansion, simulation, and backtracking. The selection step is to start from the root node of the search tree and traverse downward according to a specific strategy until it reaches a node that has not been fully explored; the expansion step means that if the selected node corresponds to a state that has not been fully explored (that is, there are legal actions that have not been evaluated), one or more new child nodes are added under the node; the simulation step means using a random strategy to start from the newly expanded node and perform a complete Monte Carlo simulation to evaluate the potential value of the node; the final backtracking step means that after each simulation is completed, the reward value obtained is backpropagated along all nodes on the path from the root node to the end state of this simulation to update the number of visits to each node on the path And the accumulated reward value .

[0037] By repeatedly iterating the above four steps of selection, expansion, simulation, and backtracking, the search tree is continuously refined and the estimation accuracy of node values ​​is improved. Finally, the optimal path can be obtained based on the node statistics after convergence.

[0038] Building on this foundation, the node traversal strategy during the MCTS algorithm's selection phase is crucial. The core task of this phase is to determine the path to traverse downward in the search tree based on the currently known information, until a node requiring further expansion or simulation is encountered. This directly determines the balance between exploration and exploitation, profoundly impacting the search direction and efficiency of MCTS. Excessive bias toward exploitation can cause the search to converge prematurely to a local optimum, missing the global optimal solution. On the other hand, excessive bias toward exploration can waste excessive computational resources and time on low-value paths, reducing overall search efficiency. Therefore, designing an effective selection strategy that strikes the right balance between exploration and exploitation is key to ensuring that Monte Carlo tree search is both efficient and comprehensive.

[0039] Among the many MCTS node selection strategies, the upper confidence bound (UCB) algorithm is widely used due to its effectiveness in balancing exploration and utilization. However, this application found that when the UCB algorithm is directly applied to the model sharding allocation scenario in this application, it fails to consider the load balancing between edge computing devices. As a result, during the search process, MCTS may waste a large amount of computing and time resources exploring allocation paths with uneven load distribution and poor overall performance. To address this problem, this application proposes a new node selection strategy based on load-aware upper confidence bound (LA-UCB).

[0040] Specifically, when evaluating the selection priority of child nodes in order to allocate edge computing devices to each model shard based on priority, LA-UCB, in addition to inheriting the original mechanism of the UCB algorithm, also introduces an evaluation item specifically for measuring load balancing.

[0041] Specifically, step S30 includes the following steps: S300: Initialize m=1.

[0042] S301: Using the load-aware upper confidence limit strategy, based on the cumulative reward value of each edge computing device under the task allocation strategy of the previous n-1 iterations, calculate the priority of allocating the mth model shard to each edge computing device, and thus allocate the mth model shard.

[0043] S302: Update m=m+1 and return to execute step S301 until, based on the allocation results of the first m-1 model shards, there is an edge computing device that has not been selected to allocate the mth model shard in the first n-1 iterations, then allocate the mth model shard to the edge computing device.

[0044] S303: Randomly assign edge computing devices to the m+1th to Mth model shards, thereby obtaining the task allocation strategy for the nth iteration based on the allocation results of the 1st to Mth model shards.

[0045] Specifically, step S301 corresponds to the selection step in the MCTS algorithm. This application uses a load-aware upper confidence limit strategy to replace the upper confidence limit strategy in the selection step of the traditional MCTS algorithm, so that the load balancing problem of each edge computing device can be considered when selecting edge computing devices for each model shard; step S302 corresponds to the expansion step in the MCTS algorithm, and step S303 corresponds to the simulation step in the MCTS algorithm.

[0046] Furthermore, step S301 includes: The number of times each edge computing device is selected under the task allocation strategy in the first n-1 iterations, the cumulative reward value, and the load of each edge computing device after the m-th model shard is allocated to each edge computing device are input into the load-aware upper confidence limit calculation formula to obtain the priority of allocating the m-th model shard to each edge computing device; Assign the mth model shard to the edge computing device with the highest priority.

[0047] Specifically, the load-aware upper confidence limit calculation formula is expressed as: , in, The formula for calculating the load-aware upper confidence limit is: Indicates the task allocation strategy for the first n-1 iterations. The cumulative reward value of edge computing devices; Indicates the task allocation strategy for the first n-1 iterations. The number of times an edge computing device is selected; Indicates that under the task allocation strategy of the first n-1 iterations, when The total number of times an edge computing device is selected as a child node; Indicates that the mth model shard is assigned to the edge computing devices; Indicates that the mth model shard is assigned to the After the edge computing device The load of edge computing devices; 、 represents the preset positive constant hyperparameter.

[0048] Furthermore, the present application uses a simulation model to simulate the allocation strategy obtained by the simulation step of the MCTS algorithm, thereby obtaining the reward value of the current allocation strategy, and using it as the reward value of the backtracking step in the MCTS algorithm. By backpropagating the reward value along the allocation results of the M model slices, the number of times each edge computing device is selected in the current allocation strategy is updated. And the accumulated reward value .

[0049] Specifically, step S40 includes: The simulation model of tasks to be executed is used to simulate the actual loading time, actual computing time, and actual communication delay of each edge computing device under the task allocation strategy of the nth iteration.

[0050] The actual loading time, actual computing time, and actual communication delay of each edge computing device are converted respectively to obtain the loading reward, computing reward, and communication delay reward of each edge computing device.

[0051] The loading reward, computing reward, and communication delay reward of each edge computing device are weighted and summed to obtain the reward value of each edge computing device under the task allocation strategy of the nth iteration.

[0052] Specifically, the calculation formula of the reward value of each edge computing device under the task allocation strategy of the nth iteration is expressed as: , in, Indicates the The reward value of an edge computing device; Represents the conversion function from time to reward; Indicates the The actual loading time of each edge computing device; express The weight of Indicates the The actual computing time of each edge computing device; express The weight of Indicates the The actual communication delay of each edge computing device; express The weight of .

[0053] The following is a further explanation of the above task allocation process through a specific example:

[0054] Assume that the large model is divided into model shard 1, model shard 2, and model shard 3, which are executed sequentially. The calculation of model shard 2 depends on the calculation result of model shard 1, and the calculation of model shard 3 depends on the calculation result of model shard 2. At the same time, there are edge computing devices ED1, ED2, and ED3 with different properties such as computing power and memory size. It is necessary to obtain the optimal task allocation strategy to deploy the three model shards to the three edge computing devices, so that the output results and execution performance of the collaborative task execution of multiple devices can reach the best.

[0055] The entire task allocation process can be viewed as a search tree. The root node of the tree represents the initial state where no model shard has been assigned to an edge computing device. The first-level nodes of the tree represent the assignment of model shard 1 to an edge computing device, the second-level nodes represent the assignment of model shard 2 to an edge computing device, and the third-level nodes represent the assignment of model shard 3 to an edge computing device.

[0056] Step 1: Select an edge computing device for model shard 1 based on the load-aware upper confidence bound strategy. Specifically, the LA-UCB strategy evaluates the priority of assigning model shard 1 to ED1, ED2, and ED3. It not only considers which edge computing device performed better in the historical iteration process, but also gives a certain chance to those edge computing devices that have been selected less frequently. In addition, it also considers load balancing among various devices.

[0057] Assuming that ED2 is finally selected for model shard 1, the algorithm will start from the first-level node representing "model shard 1 is already on ED2" and continue to use the LA-UCB strategy to select a device (ED1, ED2, or ED3) for model shard 2 and generate the second-level nodes. This process continues until it reaches a "leaf node" or a node that has not been fully explored. For example: when allocating for model shard 2, starting from the first-level node "model shard 1 is already on ED2", past iterations have tried to allocate model shard 2 to ED1 and ED2, but have never tried to allocate it to ED3, that is, the node "model shard 1 is on ED2" currently has only two child nodes, "model shard 2 is on ED1" and "model shard 2 is on ED2", which indicates that it has reached a "leaf node" or a node that has not been fully explored.

[0058] Step 2: Expansion. When reaching a "leaf node," or a node that hasn't been fully explored, the algorithm creates a new child node in the tree. This node represents the decision: "Assign model shard 2 to ED3, given that model shard 1 has already been assigned to ED2." This is the process of adding a new branch to the end of the existing tree. This new child node is the starting point for the next "simulation."

[0059] Step 3: Simulation. Starting with a brand new child node, the algorithm needs to quickly evaluate the long-term value of this allocation decision path. Specifically, the allocations for model shards 1 and 2 have been determined (ED2 and ED3), leaving only model shard 3 unallocated. To quickly reach a result, the algorithm uses a random strategy to complete the allocation for all remaining shards (in this case, only model shard 3).

[0060] Assume that the system randomly selects ED1 for model shard 3, thus obtaining a complete allocation plan: {model shard 1 → ED2, model shard 2 → ED3, model shard 3 → ED1}. That is, it is necessary to simulate the execution of this complete plan in the simulation environment and calculate its reward value R.

[0061] Step 4: Backtracking: Propagate the reward value R obtained in the simulation step backward to update the information of all nodes along the allocation decision path. Specifically, this reward value R is propagated back to each node along the path. In this example, the following three nodes are updated: 1. The root node (initial state); 2. The node "Model Shard 1 → ED2"; and 3. The node "Model Shard 2 → ED3".

[0062] For each node on the current task allocation decision path, its visit count is incremented by 1, and the cumulative reward value is added to the R obtained in this simulation. Effectively, if the path {model shard 1 → ED2, model shard 2 → ED3, ...} ultimately achieves a high reward (i.e., a short total time), the value of the decisions "model shard 1 → ED2" and "model shard 2 → ED3" is increased. In future "selection" steps, the load-aware upper confidence bound strategy will make this promising path more likely to be chosen again.

[0063] Furthermore, given the resource constraints and significant heterogeneity of edge computing devices, not all edge computing devices can meet the operational requirements (i.e., successfully load and complete computations) for each model shard in a given large model. Therefore, ensuring the feasibility of each model shard on each edge computing device is a key constraint that must be strictly considered during the search process for model shard allocation. To address this issue, this application employs a feasibility assurance mechanism based on hierarchical pruning.

[0064] Specifically, before starting the MCTS-based model shard allocation search, the present application first constructs a list that clearly identifies the feasibility relationship between the model shard and the edge computing device based on the evaluation data obtained in step S20. The list records in detail which model shards cannot be successfully deployed and executed on the edge computing device due to the hardware resource limitations of the specific edge computing device. During the MCTS search process, when the algorithm needs to allocate the model shard to the current When evaluating all possible edge computing devices, a pre-generated feasibility list is queried. If an edge computing device Marked as sharding the model If not feasible, shard the model Assigned to edge computing devices This potential decision option will be immediately considered an illegal action and removed from the possible expansion options of the current node. This means that the MCTS child node representing the illegal assignment will not be generated, and its corresponding search branch will be pruned immediately, thus effectively avoiding subsequent exploration and simulation of infeasible paths. That is, only when the edge computing device Confirmed to shard the model When feasible, the allocation decision is considered a legal action and may be included in the search tree expansion and evaluation process.

[0065] Specifically, before step S30, it also includes: based on the preset loading time and preset calculation time of each model shard, and the loading time and calculation time of each edge computing device for each model shard, constructing a set of available edge computing devices for each model shard, so as to iteratively search for edge computing devices for M model shards in turn in the set of available edge computing devices for each model shard.

[0066] To more clearly demonstrate the technical solution provided by this application, several specific examples are provided below to illustrate how to implement the multi-device adaptive collaborative task execution solution (MAC-IP) based on large model edge inference provided by this application in a practical environment. The following examples demonstrate the feasibility and application effects of this method: Example 1 of the present application provides a collaborative task execution method based on two edge computing devices. In this embodiment, edge computing device A and edge computing device B serve as two edge computing devices for collaborative task execution. The computing resources and memory resources of each edge computing device are limited, and it is impossible to load and run the entire large model at the same time.

[0067] First, the large model is split into four shards: model shard 1, model shard 2, model shard 3, and model shard 4. To achieve collaborative task execution, edge computing device A and edge computing device B communicate through the network to exchange calculation results and complete the transfer of data dependencies.

[0068] The model shard allocation strategy obtained in this embodiment is: edge computing device A is responsible for loading and calculating model shard 1 and model shard 3, and edge computing device B is responsible for loading and calculating model shard 2 and model shard 4.

[0069] Based on the above allocation strategy, the final multi-device collaborative task execution process is as follows: edge computing device A and edge device B load model shard 1 and model shard 2 in parallel, and the loading process and the calculation process are performed alternately; edge computing device A starts to calculate model shard 1 after loading model shard 1, and edge computing device B waits for the calculation result of model shard 1 after loading model shard 2; once edge computing device A completes the calculation of model shard 1, it immediately transmits the calculation result of model shard 1 to edge computing device B and starts loading model shard 3 at the same time; after receiving the calculation result of model shard 1, edge computing device B continues to calculate model shard 2, and after the calculation is completed, it transmits the calculation result of model shard 2 to edge computing device A, and loads model shard 4 at the same time; after receiving the calculation result of model shard 2, edge computing device A continues to calculate model shard 3, and after the calculation is completed, it transmits the calculation result of model shard 3 to edge computing device B; after receiving the calculation result of model shard 3, edge computing device B continues to calculate model shard 4, and after the calculation is completed, it transmits the calculation result of model shard 4 to edge computing device A.

[0070] Compared with traditional single-device task execution, using two devices to collaboratively execute tasks significantly improves task execution efficiency and reduces end-to-end latency. At the same time, since tasks are assigned to two devices for processing, the memory usage and computing load of each device are effectively balanced, avoiding performance bottlenecks caused by resource overload on a single device.

[0071] Example 2 of the present application provides a method for executing collaborative tasks on multiple devices. In this embodiment, edge computing devices A, B, C, and D with heterogeneous performance serve as four edge computing devices for executing collaborative tasks, where edge computing devices A and B are high-performance devices, and C and D are low-performance devices.

[0072] First, the large model is split into eight shards: model shards 1 to 8. To achieve collaborative task execution, edge computing devices A, B, C, and D communicate over the network to exchange computing results and complete the transfer of data dependencies.

[0073] The model shard allocation strategy obtained in this embodiment is: edge computing device A (high performance) is responsible for loading and calculating model shard 1, model shard 2, and model shard 3; edge computing device B (high performance) is responsible for loading and calculating model shard 4, model shard 5, and model shard 6; edge computing device C (low performance) is responsible for loading and calculating model shard 7; edge computing device D (low performance) is responsible for loading and calculating model shard 8.

[0074] This allocation connects consecutive shards in sequence: the input passes through model shards 1→2→3 (edge ​​computing device A), then through model shards 4→5→6 (edge ​​computing device B), and then through model shard 7 (edge ​​device C) and model shard 8 (edge ​​device D), forming a pipeline; high-performance devices take on more shards, and low-performance devices take on fewer shards to avoid resource overload.

[0075] Based on the above allocation strategy, the final multi-device collaborative task execution process is as follows: Edge computing device A loads model shards 1-3 into memory in parallel, Edge computing device B loads model shards 4-6 in parallel, Edge computing device C loads model shard 7, and Edge computing device D loads model shard 8. Loading can be carried out in parallel with the previous round of inference calculations to reduce latency.

[0076] When a new input arrives, edge computing device A first performs forward computations in the order of model shard 1, model shard 2, and model shard 3, obtaining the third-level output. Edge computing device A transmits the third-level output to edge computing device B and simultaneously begins loading model shard 1 for the next input. Upon receiving this, edge computing device B computes model shards 4, 5, and 6, obtaining the sixth-level output and transmitting it to edge computing device C. Upon receiving this, edge computing device C computes model shard 7, obtaining the seventh-level output and transmitting it to edge computing device D. Upon receiving this, edge computing device D computes model shard 8, outputs the final result, and returns it to the initiator. If pipeline parallelism is employed, multiple inputs can be processed concurrently across edge computing devices: while edge computing device B is processing the current input, edge computing device A can start model shards 1-3 for the next input, thereby improving throughput.

[0077] It's worth noting that during operation, MAC-IP continuously monitors the CPU / GPU utilization, memory usage, and network latency of each edge computing device. If an edge computing device becomes overloaded or network fluctuations increase, temporary adjustments can be made: some model shards can be reallocated to idle or high-performance edge computing devices, and the number of parallel inputs can be adjusted to alleviate pressure. This online adjustment ensures the continued efficient and stable execution of tasks in heterogeneous and dynamic environments.

[0078] This multi-device collaborative task execution significantly improves task execution efficiency, especially on resource-constrained edge computing devices, where tasks can be rationally allocated to optimize resource utilization. Because each edge computing device allocates tasks based on its resource availability, overall system performance is greatly improved. This allows for efficient model loading and computation on low-performance edge computing devices, avoiding performance bottlenecks.

[0079] Based on the multi-device collaborative task execution method provided in the above embodiment, the embodiment of the present application further provides a multi-device collaborative task execution apparatus, which specifically includes: The model building and sharding module is used to build a large model based on the tasks to be executed, and divide the large model into M model shards for sequentially processing the tasks to be executed.

[0080] The modeling and analysis module is used to load and calculate each model shard on each interconnected edge computing device, obtain the loading time, calculation time and communication delay of each edge computing device for each model shard, and thus build a simulation model of the task to be executed.

[0081] The model allocation module is used to use the Monte Carlo tree search algorithm and the load-aware upper confidence bound strategy to iteratively search edge computing devices for M model shards in turn based on the execution order of the model shards and the cumulative reward values ​​of the task allocation strategies of the first n-1 iterations to obtain the task allocation strategy of the nth iteration; where n≥1, when n=1, the reward value of the task allocation strategy of the n-1th iteration is a preset value.

[0082] The environment simulation module is used to use the simulation model of the task to be executed to simulate the task allocation strategy of the nth iteration, obtain the reward value of the task allocation strategy of the nth iteration, update n=n+1 and return to execute the steps executed by the iterative allocation module until the preset iterative convergence condition is reached, and obtain the target task allocation strategy.

[0083] The collaborative task execution module is used to deploy M model shards on corresponding edge computing devices based on the target task allocation strategy, so as to utilize the model shards on multiple edge computing devices to collaboratively implement the tasks to be executed.

[0084] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned multi-device collaborative task execution method are implemented.

[0085] Furthermore, the Industrial Internet of Things includes numerous production equipment and equipment status monitoring sensors. In order to timely understand and follow up on production status, intelligent question-answering robots are often used to parse input query questions and output parsing results, thereby realizing real-time query of production data. Due to the heterogeneous characteristics of multi-source data and the extremely large amount of data in the Industrial Internet of Things, the number of parameters and scale of intelligent question-answering robots are large. Therefore, in order to ensure the timeliness and accuracy of production data query, the embodiment of the present application provides an application of a multi-device collaborative task execution method in the intelligent question-answering of the Industrial Internet of Things, which specifically includes: Step 1: Build a large model of an intelligent question-answering robot that implements the status changes of various production equipment in the Industrial Internet of Things and the query tasks of various sensor monitoring data; divide the large model of the intelligent question-answering robot into M model shards that sequentially process the input query questions.

[0086] Step 2: Load and calculate each model shard on each interconnected edge computing device in the industrial Internet of Things, obtain the loading time, calculation time and communication delay of each model shard on each edge computing device, and thus build an intelligent question-answering robot simulation model.

[0087] Specifically, each edge computing device is deployed on the edge computing node of the industrial Internet of Things and is directly connected to the production equipment and sensor network in the industrial Internet of Things.

[0088] Step 3: Using the Monte Carlo tree search algorithm and the load-aware upper confidence bound strategy, based on the execution order of the model shards and the cumulative reward values ​​of the task allocation strategies of the first n-1 iterations, iteratively search for edge computing devices for each of the M model shards to obtain the task allocation strategy of the nth iteration; where n ≥ 1, when n = 1, the reward value of the task allocation strategy of the n-1th iteration is the preset value.

[0089] Step 4: Use the intelligent question-answering robot simulation model to simulate the input query question based on the task allocation strategy of the nth iteration, obtain the reward value of the task allocation strategy of the nth iteration, update n=n+1 and return to execute step 3 until the preset iterative convergence condition is reached, and obtain the target task allocation strategy.

[0090] Step 5: Deploy M model shards on corresponding edge computing devices based on the target task allocation strategy, so as to answer the input query questions using the model shards on multiple edge computing devices.

[0091] In a specific example, the task to be executed (i.e., the input query question) is the power consumption of production equipment O in the Industrial Internet of Things. The intelligent question-answering robot model is divided into model shards 1 to 4. Each model shard performs the following tasks: Model shard 1 performs natural language parsing on the input query question. Model shard 2 matches production equipment O from the Industrial Internet of Things based on the parsed question. Model shard 3 obtains parameters such as the usage time, operating power, voltage and current of production equipment O. Model shard 4 calculates the power consumption of production equipment O based on the parameters output by model shard 3. The industrial Internet of Things includes edge computing devices A and B. The execution process of the intelligent question-answering task obtained in this embodiment is as follows:

[0092] Edge computing device A and edge computing device B load model shard 1 and model shard 2 in parallel, and the loading process and the calculation process are performed alternately.

[0093] Edge computing device A starts calculating model shard 1 immediately after loading model shard 1, and edge computing device B waits for edge computing device A to complete the calculation and obtain the calculation result of model shard 1 after loading model shard 2.

[0094] After edge computing device A completes the calculation of model shard 1, it transmits the calculation result to edge computing device B through the network and starts loading model shard 3 at the same time.

[0095] Edge computing device B starts to calculate model shard 2 using the calculation results of model shard 1 transmitted from edge computing device A.

[0096] After edge computing device B completes the calculation of model shard 2, it transmits the calculation result to edge computing device A and loads model shard 4 for calculation at the same time.

[0097] After receiving the calculation result of model shard 2, edge computing device A continues to calculate model shard 3 and prepares to transmit its calculation result to edge computing device B.

[0098] After edge computing device A completes the calculation of model shard 3, it transmits the calculation result to edge computing device B. Edge computing device B continues to process model shard 4 and transmits the final output result to edge computing device A for further processing or feedback to the user.

[0099] By distributing the large intelligent Q&A robot model across multiple edge computing devices, not only is the intelligent Q&A response speed improved, but it can also process larger amounts of data, ensuring that the large intelligent Q&A robot model can access large amounts of data from production equipment and sensors in real time and provide timely feedback. Furthermore, the collaborative reasoning of multiple edge computing devices ensures that the memory and computing resources of each edge computing device are effectively managed through load balancing between them, thus avoiding system crashes or response failures caused by resource overload.

[0100] This application uses distributed reasoning and dynamic task allocation to slice and distribute the computing tasks of large models to multiple edge computing devices, achieving high parallelization of model loading, calculation and communication. It not only greatly improves the reasoning speed, but also reduces the waiting time between devices during the reasoning process through pipeline parallelism, thereby significantly reducing the end-to-end reasoning delay and giving full play to the potential of the edge computing environment. At the same time, the large model is split into multiple subtasks through sharding technology and assigned to different devices for processing. This approach effectively alleviates the resource bottleneck of a single device, allowing edge computing devices to collaboratively complete large model reasoning tasks without exceeding resource limits. In addition, this technical solution can not only achieve efficient deployment of large model reasoning in the existing edge computing environment, but also has high scalability. With the addition of more edge devices, MAC-IP can flexibly adjust the reasoning strategy according to the heterogeneity and load conditions between devices, further improving the computing power and reasoning efficiency of the system, and adapting to application needs of different scales.

[0101] In addition, when allocating multiple model shards, by adopting a load-aware upper confidence bound (LA-UCB) strategy, the computing load between devices can be effectively balanced to avoid excessive device load, thereby improving the overall performance of the system. This ensures that the computing power and load of each device are fully considered when allocating tasks, thereby optimizing resource utilization.

[0102] In summary, the MAC-IP technical solution provided by this application demonstrates significant advantages in improving the efficiency of large-model edge reasoning, reducing resource consumption, and enhancing system stability and scalability. It fills the gap in existing technologies for deploying large models in edge computing environments and demonstrates strong technological advancement. At the same time, through experimental data verification in both simulation and real-world environments, the solution provided by this application can achieve a large-model data processing acceleration ratio of 1.81 to 2.53 times, effectively reducing peak memory usage, and achieving a memory compression ratio of 1.34 to 1.79 times, greatly alleviating the memory pressure on resource-constrained edge computing devices.

[0103] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0104] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0105] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0107] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A multi-device collaborative task execution method, characterized in that: include: S10: Build a large model based on the task to be executed, and divide the large model into M model slices for sequentially processing the task to be executed; S20: Load and calculate each model slice on each interconnected edge computing device, obtain the loading time, calculation time and communication delay of each model slice on each edge computing device, and thus build a simulation model of the task to be executed; S30: Using the Monte Carlo tree search algorithm and the load-aware upper confidence bound strategy, based on the execution order of the model shards and the cumulative reward values ​​of the task allocation strategies of the previous n-1 iterations, iteratively search for edge computing devices for the M model shards in turn to obtain the task allocation strategy of the nth iteration; where n ≥ 1, when n = 1, the reward value of the task allocation strategy of the n-1th iteration is a preset value; S40: Using the simulation model of the task to be executed to simulate the task allocation strategy of the nth iteration, obtain the reward value of the task allocation strategy of the nth iteration, update n=n+1 and return to step S30 until the preset iteration convergence condition is reached to obtain the target task allocation strategy; S50: Based on the target task allocation strategy, the M model shards are respectively deployed on the corresponding edge computing devices, so as to utilize the model shards on multiple edge computing devices to collaboratively implement the tasks to be executed.

2. The multi-device collaborative task execution method according to claim 1, characterized in that: When the task to be executed is an intelligent question-answering task, the large model is an intelligent question-answering robot model, the input of the large model is the query text, and the output of the large model is the answer text; When the task to be performed is an image classification task, the large model is an image classification model, the input of the large model is the image to be classified, and the output of the large model is the image category prediction probability.

3. The multi-device collaborative task execution method according to claim 1, characterized in that: Step S30 includes: S300: Initialize m=1; S301: Using the load-aware upper confidence bound strategy, based on the cumulative reward value of each edge computing device under the task allocation strategy of the previous n-1 iterations, calculate the priority of allocating the mth model shard to each edge computing device, thereby allocating the mth model shard; S302: Update m=m+1 and return to step S301. If, based on the allocation results of the first m-1 model shards, there is an edge computing device that has not been selected for allocation of the m-th model shard in the first n-1 iterations, the m-th model shard is allocated to the edge computing device. S303: Randomly assign edge computing devices to the m+1th to Mth model shards, thereby obtaining the task allocation strategy for the nth iteration based on the allocation results of the 1st to Mth model shards.

4. The multi-device collaborative task execution method according to claim 3, characterized in that: Step S301 includes: The number of times each edge computing device is selected under the task allocation strategy in the first n-1 iterations, the cumulative reward value, and the load of each edge computing device after the m-th model shard is allocated to each edge computing device are input into the load-aware upper confidence limit calculation formula to obtain the priority of allocating the m-th model shard to each edge computing device; Assign the mth model shard to the edge computing device with the highest priority.

5. The multi-device collaborative task execution method according to claim 4, characterized in that: The calculation formula for the load-aware upper confidence limit is expressed as: , in, The formula for calculating the load-aware upper confidence limit is shown in Figure 2. Indicates the task allocation strategy for the first n-1 iterations. The cumulative reward value of edge computing devices; Indicates the task allocation strategy for the first n-1 iterations. The number of times an edge computing device is selected; Indicates that under the task allocation strategy of the first n-1 iterations, when The total number of times an edge computing device is selected as a child node; Indicates that the mth model shard is assigned to the edge computing devices; Indicates that the mth model shard is assigned to the After the edge computing device The load of edge computing devices; 、 represents the preset positive constant hyperparameter.

6. The multi-device collaborative task execution method according to claim 1, characterized in that: Step S40 includes: The simulation model of the task to be executed is used to simulate the actual loading time, actual computing time, and actual communication delay of each edge computing device under the task allocation strategy of the nth iteration; The actual loading time, actual computing time, and actual communication delay of each edge computing device are converted to obtain the loading reward, computing reward, and communication delay reward of each edge computing device; The loading reward, computing reward, and communication delay reward of each edge computing device are weighted and summed to obtain the reward value of each edge computing device under the task allocation strategy of the nth iteration.

7. The multi-device collaborative task execution method according to claim 6, characterized in that: The calculation formula of the reward value of each edge computing device under the task allocation strategy of the nth iteration is expressed as: , in, Indicates the The reward value of an edge computing device; Represents the conversion function from time to reward; Indicates the The actual loading time of each edge computing device; express The weight of Indicates the The actual computing time of each edge computing device; express The weight of Indicates the The actual communication delay of each edge computing device; express The weight of .

8. The multi-device collaborative task execution method according to claim 1, characterized in that: Before step S30, the following steps are also included: Based on the preset loading time and preset calculation time of each model shard, as well as the loading time and calculation time of each edge computing device for each model shard, a set of available edge computing devices for each model shard is constructed, so as to iteratively search for edge computing devices for M model shards in the set of available edge computing devices for each model shard.

9. A multi-device collaborative task execution device, characterized in that: include: The model building and sharding module is used to build a large model based on the tasks to be executed, and divide the large model into M model shards for sequentially processing the tasks to be executed; The modeling and analysis module is used to load and calculate each model slice on each interconnected edge computing device, obtain the loading time, calculation time and communication delay of each model slice on each edge computing device, and thus build a simulation model of the task to be executed; The model allocation module is used to use the Monte Carlo tree search algorithm and the load-aware upper confidence bound strategy to iteratively search edge computing devices for M model shards in sequence based on the execution order of the model shards and the cumulative reward values ​​of the task allocation strategies of the previous n-1 iterations to obtain the task allocation strategy of the nth iteration; where n ≥ 1, when n = 1, the reward value of the task allocation strategy of the n-1th iteration is the preset value; The environment simulation module is used to simulate the task allocation strategy of the nth iteration using the simulation model of the task to be executed, obtain the reward value of the task allocation strategy of the nth iteration, update n=n+1 and return to execute the steps executed by the iterative allocation module until the preset iterative convergence condition is reached, and obtain the target task allocation strategy; The collaborative task execution module is used to deploy M model shards on corresponding edge computing devices based on the target task allocation strategy, so as to utilize the model shards on multiple edge computing devices to collaboratively implement the tasks to be executed.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the multi-device collaborative task execution method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Edge computing task allocation method based on deep Monte Carlo tree search

    CN110427261A

  • Block chain task processing method combined with artificial intelligence and related device thereof

    CN117135159A

  • Intelligent computing power resource scheduling method based on cloud edge collaboration

    CN119003184A

  • Cross-chain fragmentation scheduling method based on DAG

    CN119271380A

  • Digital twin-based edge-end collaborative scheduling method for heterogeneous tasks and resources

    US20250086005A1

Cited By

  • Industrial edge-oriented lightweight AI model dynamic loading system

    CN121560417A