Large model operation method and device based on edge device, equipment and medium

By preprocessing and optimizing the large model, combined with the model collaborative operation framework and data cache management algorithm, the problem of unreasonable resource allocation, low operation efficiency and latency on edge devices of the large model is solved, and efficient resource utilization and application scenario expansion is achieved.

CN120196449AActive Publication Date: 2025-06-24INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510668615.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The application of large models on edge devices faces problems such as unreasonable resource allocation, low operational efficiency, prominent latency problems and insufficient hardware capabilities.

Method used

By preprocessing and optimizing the large model to be run, quantifying, pruning, knowledge distillation and sparse utilization are performed based on the hardware characteristics of edge devices, combining the model collaborative operation framework and data cache management algorithm, task priority and resources are dynamically allocated to achieve data preloading and transmission optimization.

Benefits of technology

It improves the operation efficiency of large models on edge devices, optimizes resource allocation, reduces latency, fully utilizes the hardware capabilities of edge devices, and expands application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196449A_ABST
    Figure CN120196449A_ABST
Patent Text Reader

Abstract

The invention discloses a large model operation method and device based on edge equipment, equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the preprocessing of a to-be-operated large model, and carrying out the optimization of the to-be-operated large model after preprocessing based on the hardware characteristics of the edge equipment; determining the task priority of a to-be-executed task corresponding to the optimized model based on the resource condition of the current edge device; allocating the to-be-executed task to a pre-established model collaborative operation framework according to the task priority, and analyzing the to-be-executed task in the allocation process, so as to deploy the to-be-executed task to a load balancing node of the model collaborative operation framework according to an analysis result; and managing the related parameter data by using a preset data cache management algorithm to pre-load to-be-used data in the related parameter data into the edge device, so that the optimized model executes a corresponding to-be-executed task based on the to-be-used data through a load balancing node of the model collaborative operation framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large model operation method, device, equipment and medium based on edge devices. Background Art

[0002] At present, the application of large models on edge devices faces many challenges. On the one hand, the existing edge computing solutions mostly use static task allocation, which leads to low task processing efficiency when resources are tight, and cannot fully utilize device capabilities when resources are sufficient; on the other hand, the caching strategy of model parameters and intermediate results in the existing technology is relatively simple. During the operation of large models, the need to temporarily load data increases the running time, reduces the running efficiency, and further affects the user experience; on the one hand, when complex tasks are processed, the delay problem is more prominent; on the other hand, the hardware capability development of edge devices is limited, and its hardware capabilities cannot be fully utilized, especially in low-power and high-performance scenarios. The problem is more prominent. In other words, the existing technology limits the widespread application of large models on edge devices, unreasonable resource allocation, low operating efficiency, and limited application scenarios.

[0003] In summary, how to optimize the running process of large models on edge devices to solve application problems caused by the shortcomings of current technology is a technical problem that needs to be solved urgently. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a large model operation method, device, equipment and medium based on edge devices, which can optimize the operation process of large models on edge devices to solve the application problems caused by the shortcomings of current technologies. The specific scheme is as follows: In a first aspect, the present application provides a large model operation method based on an edge device, comprising: Preprocessing the large model to be run, and optimizing the preprocessed large model to be run based on the hardware characteristics of the edge device to obtain an optimized model; Based on the current resource situation of the edge device, determine the task priority of the task to be executed corresponding to the optimized model; Allocate the tasks to be executed to a pre-built model collaborative operation framework according to the task priority, and analyze the tasks to be executed during the allocation process, so as to deploy the tasks to be executed to the load balancing node of the model collaborative operation framework according to the obtained analysis results; the model collaborative operation framework is a framework constructed by the edge device, the edge server and the cloud; Use a preset data cache management algorithm to manage the relevant parameter data required for running the optimized model, so as to preload the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative operation framework; the data to be used is the data predicted by the preset data cache management algorithm and needed to be used in the next step during the running of the optimized model.

[0005] Optionally, the preprocessing of the large model to be run includes: Use a preset quantization technique to compress the weights of the large model to be run, and remove redundant parameters in the large model to be run through a preset pruning technique to obtain a pruned large model; Transfer the knowledge of the pruned large model to a small model through a preset knowledge distillation technique, and redesign the model structure and adjust the parameters of the transferred model to obtain an adjusted model; Use a gradient-based sparse training method to mine the sparse characteristics of the adjusted model to obtain a compressed model with a sparse structure; Correspondingly, the optimization of the large model to be run after preprocessing based on the hardware characteristics of the edge device includes: Optimize the compressed model with a sparse structure based on the hardware characteristics of the edge device.

[0006] Optionally, the optimization of the large model to be run after preprocessing based on the hardware characteristics of the edge device to obtain an optimized model includes: Utilize the characteristics of the field programmable gate array in the edge device to map the target computing module in the large model to be run after preprocessing to the field programmable gate array; Based on the clock speed and pipeline delay rate of the field programmable gate array, calculate the computing frequency of the computing engine corresponding to the target computing module, and design the computing engine according to the computing frequency; If it is detected that a graphics processor is configured in the edge device, utilize the single instruction multiple data architecture of the graphics processor to calculate the sub-batch size based on the total batch size of the batch data and the number of cores of the graphics processor, so as to divide the batch data according to the sub-batch size and allocate the divided data to the cores of the graphics processor for parallel computing to obtain an optimized model.

[0007] Optionally, determining the task priority of the task to be executed corresponding to the optimized model based on the current resource situation of the edge device includes: Determine the corresponding task type weight according to the task type of the to-be-executed task corresponding to the optimized model, and determine the corresponding task priority processing requirement weight based on whether the to-be-executed task meets the user-defined task priority processing condition; Determine the corresponding urgency weight according to the urgency of the to-be-executed task, and determine the corresponding resource dependence weight according to the dependence degree of the to-be-executed task on resources; Calculate the task priority of the to-be-executed task based on the task type weight, the task priority processing requirement weight, the urgency weight, and the resource dependence weight; Correspondingly, the allocation of the to-be-executed task to the pre-built model collaborative operation framework includes: When the task priority is not lower than the first preset priority, allocate the to-be-executed task to the edge device in the pre-built model collaborative operation framework; When the task priority is not lower than the second preset priority and lower than the first preset priority, allocate the to-be-executed task to the edge server in the model collaborative operation framework; When the task priority is lower than the second preset priority, allocate the to-be-executed task to the cloud in the model collaborative operation framework; Among them, the first preset priority is higher than the second preset priority.

[0008] Optionally, before allocating the to-be-executed task to the pre-built model collaborative operation framework according to the task priority, it further includes: After the edge device and the edge server in the model collaborative operation framework are started, register their respective resource information to the cloud in the model collaborative operation framework through the edge device and the edge server respectively, so that the cloud manages the allocation situation of the to-be-executed task according to the resource information and the task priority; the edge device, the edge server, and the cloud in the model collaborative operation framework communicate with each other through a preset communication protocol.

[0009] Optionally, the analysis of the to-be-executed task to deploy the to-be-executed task to the load balancing node of the model collaborative operation framework according to the obtained analysis result includes: Analyze the to-be-executed task to determine the computational complexity, data volume, and real-time requirement of the to-be-executed task, and determine the corresponding relevant performance indicators of several load balancing nodes of the model collaborative operation framework; Calculate the task allocation fitness of the to-be-executed task on the corresponding load balancing node respectively based on the computing complexity, the data volume, the real-time requirement, and the relevant performance indicators, so as to allocate the to-be-executed task to the target load balancing node according to the task allocation fitness; the target load balancing node is the load balancing node corresponding to the target task allocation fitness among the several load balancing nodes, and the target task allocation fitness is the task allocation fitness with the smallest value among the task allocation fitnesses of the to-be-executed tasks corresponding to each load balancing node.

[0010] Optionally, the managing the relevant parameter data required for running the optimized model by using a preset data cache management algorithm to preload the data to be used in the relevant parameter data into the edge device includes: Calculate the data access frequency of the relevant parameter data required for running the optimized model by using the exponential weighted moving average method, and determine the importance level of the relevant parameter data; Cache the relevant parameter data into the target cache location according to the importance level and the data access frequency; the target cache location includes a memory cache location, a flash cache location, and a cloud cache location; Extract and analyze the features of the cached data in the target cache location to predict the next predicted access frequency of the cached data through a preset autoregressive moving average model; Determine the corresponding cache location score according to the relevant information of the cached data stored in the target cache location; Calculate the cache value of the cached data based on the cache location score, the predicted access frequency, and the importance level of the cached data, so as to adjust the cache location of the cached data according to the cache value, and at the same time use the user behavior prediction data output by the preset user behavior model and the current state of the model collaborative operation framework to determine the data to be used in the cached data, so as to preload the data to be used into the memory cache location of the edge device.

[0011] Optionally, during the process of executing the corresponding to-be-executed task based on the data to be used, it further includes: When the data to be used needs to be transmitted in the model collaborative operation framework, perform a preprocessing operation of compressing and encapsulating the data to be used to obtain processed data; Determine the data priority and data type of the processed data, and calculate the target channel bandwidth for transmitting the processed data based on the task priority of the to-be-executed task and the total available bandwidth; Determine the target transmission channel of the processed data according to the data priority, the data type, and the target channel bandwidth, so as to transmit the processed data by using the target transmission channel; Moreover, during the process of transmitting the processed data, calculate the check value of the processed data by using a preset data check algorithm, so as to determine whether an abnormality occurs in the transmission of the processed data according to the check value.

[0012] Optionally, during the process of performing the corresponding to-be-executed tasks based on the to-be-used data, it further includes: Continuously monitor the task queue of the edge server; When it is detected that there is an emergency task in the task queue, pause or adjust other tasks in the task queue except the emergency task; Determine the corresponding resource allocation ratio according to the target task priority of the emergency task, and determine the target computing resource amount of the emergency task and the current available computing resource amount of the edge server; Calculate the resource allocation amount of the emergency task based on the resource allocation ratio, the target computing resource amount, and the current available computing resource amount, so as to allocate the corresponding computing resources from the other tasks according to the obtained resource allocation amount, so as to execute the emergency task by using the computing resources; If after allocating the computing resources, it is detected that the current edge server cannot execute the emergency task by using the computing resources, unload the target computing link of the emergency task to the cloud, so as to use the cloud to assist the edge server in executing the emergency task.

[0013] In a second aspect, the present application provides a large model running device based on an edge device, including: A model optimization module, configured to perform preprocessing on a to-be-run large model, and optimize the preprocessed to-be-run large model based on the hardware characteristics of the edge device, so as to obtain an optimized model; A task priority determination module, configured to determine the task priority of the to-be-executed tasks corresponding to the optimized model based on the resource situation of the current edge device; A task analysis module, configured to allocate the to-be-executed tasks to a pre-built model collaborative running framework according to the task priority, and during the allocation process, analyze the to-be-executed tasks, so as to deploy the to-be-executed tasks to the load balancing nodes of the model collaborative running framework according to the obtained analysis results; the model collaborative running framework is a framework constructed by the edge device, the edge server, and the cloud; A data management module, which is used to manage relevant parameter data required for running the optimized model by using a preset data cache management algorithm, so as to preload the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative operation framework; the data to be used is the data predicted by the preset data cache management algorithm and needed to be used in the next step during the running of the optimized model.

[0014] In a third aspect, the present application provides an electronic device, including: A memory, which is used to store a computer program; A processor, which is used to execute the computer program to implement the foregoing method for running a large model based on an edge device.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, which is used to store a computer program; wherein, when the computer program is executed by a processor, the foregoing method for running a large model based on an edge device is implemented.

[0016] In this embodiment, the large model to be run is preprocessed, and the preprocessed large model to be run is optimized based on the hardware characteristics of the edge device to obtain an optimized model; based on the resource situation of the current edge device, the task priority of the task to be executed corresponding to the optimized model is determined; according to the task priority, the task to be executed is assigned to a pre-built model collaborative operation framework, and during the assignment process, the task to be executed is analyzed to deploy the task to be executed to the load balancing node of the model collaborative operation framework according to the obtained analysis result; the model collaborative operation framework is a framework constructed by the edge device, the edge server, and the cloud; a preset data cache management algorithm is used to manage the relevant parameter data required to run the optimized model, so as to preload the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative operation framework; the data to be used is the data predicted by the preset data cache management algorithm and needed to be used in the next step during the process of running the optimized model. As can be seen from the above, in this application, the large model to be run is first preprocessed, and the preprocessed large model to be run is optimized according to the hardware characteristics of the edge device. At the same time, according to the resource situation of the current edge device, the task priority of the task to be executed of the optimized model is determined, so as to assign the task to be executed to the model collaborative operation framework according to the task priority. During the assignment process, according to the analysis result of the task to be executed, it is deployed to the load balancing node of the model collaborative operation framework. At the same time, a preset data cache management algorithm is used to manage the relevant parameter data, so as to preload the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative operation framework. In this way, through the above process of this application, optimizing the large model to be run based on the hardware characteristics of the edge device can give full play to the hardware capabilities of the edge device to improve the running efficiency of the large model; determining the task priority of the task to be executed according to the resource situation of the edge device and dynamically allocating the task to be executed according to the task priority can fully and reasonably achieve resource allocation and improve system performance; using a preset data cache management algorithm to manage the relevant parameter data and preloading the data to be used, loading the relevant data into the cache of the edge device in advance, reducing the waiting time for model inference, and improving the running efficiency; constructing a model collaborative operation framework composed of an edge device, an edge server, and the cloud can effectively solve the latency problem, and then optimize the running process of the large model on the edge device to solve the application problems caused by the deficiencies of the current technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.

[0018] Figure 1 Flowchart of a large model operation method based on edge devices disclosed in this application; Figure 2 Flowchart of a specific large model operation method based on edge devices disclosed in this application; Figure 3 Schematic structural diagram of a large model operation device based on edge devices disclosed in this application; Figure 4 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0020] Currently, the application of large models on edge devices faces many challenges. On the one hand, existing edge computing solutions mostly adopt static task allocation methods, resulting in low task processing efficiency when resources are scarce, and the device capabilities cannot be fully utilized when resources are sufficient. On the other hand, the caching strategies for model parameters and intermediate results in the existing technology are relatively simple. During the operation of large models, the running time increases due to the need to temporarily load data, reducing the running efficiency and further affecting the user experience. On the one hand, the latency problem is more prominent when dealing with complex tasks. On the other hand, the development of the hardware capabilities of edge devices is limited, and their hardware capabilities cannot be fully utilized, especially in the scenarios of low power consumption and high performance. That is, the existing technology limits the wide application of large models on edge devices, with unreasonable resource allocation, low running efficiency, and limited application scenarios.

[0021] To overcome the above technical problems, this application provides a large model operation method based on edge devices to optimize the operation process of large models on edge devices and solve the application problems caused by the deficiencies of current technologies.

[0022] See Figure 1As shown in the figure, an embodiment of the present invention discloses a method for running a large model based on an edge device, including: Step S11: Preprocess the large model to be run, and optimize the preprocessed large model to be run based on the hardware characteristics of the edge device to obtain an optimized model.

[0023] In this embodiment, when a large model needs to be run on an edge device, in order to reduce the number of parameters and computational complexity of the large model, the large model to be run is first preprocessed to make it more suitable for running on resource-constrained edge devices, and the preprocessed large model to be run is optimized based on the hardware characteristics of the edge device to obtain an optimized model, further improving the inference speed of the model. Among them, the large model to be run includes, but is not limited to, large models in the fields of production manufacturing, quality control, process optimization, etc.; the hardware of the hardware characteristics includes, but is not limited to, a graphics processing unit (i.e., GPU, Graphics Processing Unit), a field programmable gate array (i.e., FPGA, Field Programmable Gate Array), etc.

[0024] It should be noted that the operations of the preprocessing include, but are not limited to, quantization, pruning, knowledge distillation, etc., and its processing flow is as follows: Use a preset quantization technique to compress the weights of the large model to be run, and remove redundant parameters in the large model to be run through a preset pruning technique to obtain a pruned large model; Transfer the knowledge of the pruned large model to a small model through a preset knowledge distillation technique, and redesign and adjust the parameters of the model structure after migration to obtain an adjusted model; Use a gradient-based sparse training method to mine the sparse characteristics of the adjusted model to obtain a compressed model with a sparse structure; Correspondingly, optimizing the preprocessed large model to be run based on the hardware characteristics of the edge device includes: optimizing the compressed model with a sparse structure based on the hardware characteristics of the edge device. That is, use a preset quantization technique to compress the weights of the large model to be run, compress its weights from 32-bit floating-point numbers to 8-bit or lower-precision integers, and at the same time remove redundant parameters in the large model to be run through a preset pruning technique to reduce the model volume and obtain a pruned large model. Subsequently, transfer the knowledge of the pruned large model to a lightweight small model through a preset knowledge distillation technique to improve the performance of the small model, and perform structure reparameterization on the model after migration, redesign and adjust its model structure to make it more suitable for running on edge devices while maintaining high accuracy to obtain an adjusted model. Specifically, the convolutional layer of the convolutional neural network of the model after migration can be reparameterized, and multiple convolutional kernels can be fused into an equivalent convolutional kernel to reduce the number of parameters. The fusion formula can be specifically: ; Among them, is the weight coefficient; is the i-th convolution kernel; represents the equivalent convolution kernel after fusion. After obtaining the adjusted model, perform sparsity utilization processing on the adjusted model, adopt a gradient-based sparse training method to mine the sparse characteristics in the adjusted model, further reduce the computational amount and storage requirements of the model through sparsity calculation and storage optimization, and dynamically adjust the sparsity of the weights during the training process to obtain a compressed model with a naturally formed sparse structure. The specific calculation formula for the sparsity can be: ; Among them, s represents the sparsity; non_zero_params are non-zero parameters, used to represent the number of non-zero weight parameters in the model; total_params are the total parameters, used to represent the total number of weight parameters in the model. Therefore, the preprocessed large model to be run in the step of optimizing the preprocessed large model to be run based on the hardware characteristics of the edge device is the compressed model with a sparse structure formed.

[0025] It should be further noted that after preprocessing the to-be-run large model, this embodiment can customize and optimize the preprocessed to-be-run large model according to the hardware characteristics of the edge device to give full play to the computing power of the hardware. The processing flow is as follows: Utilize the characteristics of the field programmable gate array (FPGA) in the edge device to map the target computing module in the preprocessed to-be-run large model to the FPGA; Based on the clock speed and pipeline delay rate of the FPGA, calculate the computing frequency of the computing engine corresponding to the target computing module to design the computing engine according to the computing frequency; If it is detected that a graphics processing unit (GPU) is configured in the edge device, utilize the single instruction multiple data (SIMD) architecture of the GPU to calculate the sub-batch size based on the total batch size of the batch data and the number of cores of the GPU, so as to split the batch data according to the sub-batch size and distribute the split data to the cores of the GPU for parallel computing to obtain the optimized model. Among them, the target computing module includes, but is not limited to, the convolutional layer, the fully connected layer, etc. That is to say, since an FPGA is usually configured in the edge device, the programmable logic and hardware acceleration capabilities of the FPGA can be utilized to map the target computing module in the preprocessed to-be-run large model to the FPGA to achieve efficient hardware acceleration. Specifically, based on the clock speed and pipeline delay rate of the FPGA, the computing frequency of the computing engine corresponding to the target computing module can be calculated to design the computing engine according to the computing frequency. For example, the convolutional layer in the model can be mapped to the programmable logic unit of the FPGA to design a dedicated convolutional computing engine. Among them, the specific calculation formula of the computing frequency can be: ; Among them, f represents the actual computing frequency of the convolutional computing engine of the FPGA; Clock_Speed is the clock speed of the FPGA; Pipeline_Delay_Rate is the pipeline delay rate, that is, the delay rate in the FPGA pipeline design. This embodiment can optimize the pipeline design through the calculated computing frequency, reduce latency, and improve computing throughput. In addition, if it is detected that a GPU is configured in the edge device, this embodiment can make full use of the SIMD (single instruction multiple data) architecture of the GPU to transform the model for parallel computing optimization to accelerate computing-intensive tasks such as matrix operations. Specifically, the batch data can be split into multiple sub-batches and distributed to different GPU cores for parallel computing. Among them, the sub-batch size of the sub-batch can be calculated based on the total batch size of the batch data and the number of cores of the GPU, so as to split the batch data according to the sub-batch size and distribute the split data to the cores of the GPU for parallel computing to obtain the optimized model. The specific calculation formula of the sub-batch size can be: ; Among them, b is the sub-batch size, which is used to represent the sub-batch size obtained by dividing the total batch of batch data and allocated to each GPU core; Total_Batch_Size is the total batch size, which is used to represent the total batch size during model training or inference; GPU_Core_Count is the number of GPU cores, which is used to represent the number of available cores in the GPU. In this way, when a large model needs to run on an edge device in this embodiment, by preprocessing the large model to be run through technologies such as quantization, pruning, knowledge distillation, structural reparameterization, and sparsity utilization, it is possible to reduce the number of parameters, computational complexity, and storage requirements of the large model while maintaining high precision, making it more suitable for running on resource-constrained edge devices; customizing and optimizing the preprocessed large model to be run based on the hardware characteristics of the edge device to give full play to the computing power of the hardware, achieve efficient hardware acceleration, reduce latency, increase computational throughput, and further improve the inference speed of the model.

[0026] Step S12: Based on the resource situation of the current edge device, determine the task priority of the task corresponding to the optimized model.

[0027] In this embodiment, after obtaining the optimized model, the task priority of the task corresponding to the optimized model can be determined based on the resource situation of the current edge device, so as to perform task allocation on the task to be executed according to the obtained task priority, and intelligently determine that the task is more suitable to be executed at any one of the local, cloud, or edge server, so as to achieve efficient collaborative processing of the task. Among them, the resource situation refers to the real-time resource status of the edge device, including but not limited to the CPU (i.e., Central Processing Unit, central processor) usage rate, memory occupancy rate, network bandwidth, etc.

[0028] It should be noted that the determination of the task priority in this embodiment can be comprehensively based on multiple factors, and its processing flow is as follows: Determine the corresponding task type weight according to the task type of the task to be executed corresponding to the optimized model, and determine the corresponding task priority processing requirement weight based on whether the task to be executed meets the user-defined task priority processing condition; Determine the corresponding urgency weight through the urgency of the task to be executed, and determine the corresponding resource dependence weight according to the dependence degree of the task to be executed on resources; Calculate the task priority of the task to be executed based on the task type weight, the task priority processing requirement weight, the urgency weight, and the resource dependence weight. Among them, the user-defined task priority processing condition refers to whether the task to be executed is a task that the user clearly requires to be processed with priority. That is, determine the corresponding task type weight according to the task type of the task to be executed corresponding to the optimized model. For example, tasks with high real-time requirements such as device status monitoring and fault warning have relatively high task priorities, and their task type weights can be set to 0.6; Batch processing tasks such as data statistical analysis have relatively low task priorities, and their task type weights can be set to 0.3; Determine the corresponding task priority processing requirement weight based on whether the task to be executed meets the user-defined task priority processing condition. If it is a task that the user clearly requires to be processed with priority, its priority will be correspondingly increased, and its task priority processing requirement weight can be set to 0.5, and the task priority processing requirement weight of ordinary tasks can be set to 0.2; Determine the corresponding urgency weight through the urgency of the task to be executed. For example, the urgency of emergency tasks such as handling sudden failures is higher than that of regular tasks, and the urgency weight of emergency tasks can be set to 0.7, and the urgency weight of regular tasks can be set to 0.3; Determine the corresponding resource dependence weight according to the dependence degree of the task to be executed on resources. For tasks with high resource requirements, if the current resources are tight, their priorities will be dynamically adjusted. Among them, the resource dependence weight of high-resource-dependent tasks can be set to 0.4, and the resource dependence weight of low-resource-dependent tasks can be set to 0.1. After determining the task type weight, the task priority processing requirement weight, the urgency weight, and the resource dependence weight, use the above weights to calculate the task priority of the task to be executed. The specific calculation formula of the task priority can be: ; Among them, is the task type weight; is the task priority processing requirement weight; is the urgency weight; is the weight of resource dependence. In this way, based on the resource situation of the current edge device and the relevant information of the task to be executed, this embodiment determines the task priority by comprehensively considering various factors such as task type, task priority processing requirements, task urgency, and task resource dependence, which can optimize resource allocation, meet business requirements, ensure the timely execution of critical tasks, and improve the operation efficiency of the large model.

[0029] Step S13: Allocate the task to be executed to the pre-built model collaborative operation framework according to the task priority, and during the allocation process, analyze the task to be executed to deploy the task to be executed to the load balancing node of the model collaborative operation framework according to the obtained analysis result; the model collaborative operation framework is a framework built by the edge device, the edge server, and the cloud.

[0030] In this embodiment, the task to be executed corresponding to the determined task priority is allocated to the pre-built model collaborative operation framework, and during the allocation, the task to be executed is analyzed to deploy the task to be executed to the load balancing node of the model collaborative operation framework according to the obtained analysis result. Among them, the model collaborative operation framework is a unified edge-cloud collaborative framework built by the edge device, the edge server, and the cloud, which can support seamless data transmission and task collaboration between the edge device, the edge server, and the cloud; the cloud, as the collaborative control center in the model collaborative operation framework, is used to reasonably allocate the task to be executed to each load balancing node according to the task requirements and resource status, and coordinate the task execution order and data interaction between each load balancing node to ensure the efficient completion of the task.

[0031] It is understandable that before allocating the task to be executed to the pre-built model collaborative operation framework, it is necessary to first build the model collaborative operation framework. Specifically, it is necessary to first formulate a unified communication protocol to standardize the data transmission format, communication process, and error handling mechanism between edge devices, edge servers, and the cloud, ensuring seamless communication between nodes. For example, the data packet format can be defined as follows: Header (20 bytes): contains data type, source address, destination address, timestamp, etc. Payload (variable length): actual transmitted data. Footer (4 bytes): contains a checksum. Subsequently, resource discovery and registration are carried out, and the processing flow is as follows: when the edge device and the edge server in the model collaborative operation framework are started, they respectively register their own resource information with the cloud in the model collaborative operation framework, so that the cloud can manage the allocation of the task to be executed according to the resource information and the task priority; the edge device, the edge server, and the cloud in the model collaborative operation framework communicate with each other through a preset communication protocol. Among them, the resource information includes but is not limited to device type, computing power, storage capacity, etc. That is, when the edge device and the edge server in the model collaborative operation framework are started, they respectively register their own resource information with the cloud in the model collaborative operation framework, so that the cloud can maintain a global resource directory according to the resource information and the task priority, and manage the allocation of the task to be executed, including querying and allocating during task scheduling.

[0032] It should be noted that after determining the task priority, the process of allocating the to-be-executed task to the pre-built model collaborative operation framework according to the task priority is as follows: When the task priority is not lower than the first preset priority, the to-be-executed task is allocated to the edge device in the pre-built model collaborative operation framework; When the task priority is not lower than the second preset priority and lower than the first preset priority, the to-be-executed task is allocated to the edge server in the model collaborative operation framework; When the task priority is lower than the second preset priority, the to-be-executed task is allocated to the cloud in the model collaborative operation framework; where the first preset priority is higher than the second preset priority. That is, the allocation of the to-be-executed task is carried out according to the set first preset priority and the second preset priority. For example, the first preset priority can be set to 0.6 and the second preset priority can be set to 0.3. When the task priority, that is, the P value, is not lower than 0.6, the to-be-executed task is allocated to the edge device in the pre-built model collaborative operation framework, and it is preferentially executed locally on the edge device. If the local resources are insufficient, it is allocated to the adjacent edge server. If the local resources of the adjacent edge server are still insufficient, it is allocated to the cloud dedicated resources; When the P value is not lower than 0.3 and lower than 0.6, the to-be-executed task is allocated to the edge server in the model collaborative operation framework, and it is preferentially executed on the edge server. If the resources of the edge server are tight, part of the to-be-executed task is allocated to the cloud; When the P value is lower than 0.3, the to-be-executed task is allocated to the cloud in the model collaborative operation framework, and batch processing is mainly carried out in the cloud. It can be understood that in this embodiment, the resource usage situation of the edge device can be continuously collected, the task priority can be determined according to the obtained resource monitoring data, and tasks can be dynamically allocated according to the task priority. For example, for tasks with tight resources, part of the tasks can be automatically offloaded to the cloud or the edge server; For devices with sufficient resources, more tasks are preferably executed locally. By dynamically adjusting the allocation of tasks between the edge device and the cloud, the resource utilization rate can be optimized and the latency can be reduced.

[0033] It should be further pointed out that in the process of allocating the tasks to be executed, it is also necessary to perform task allocation for the load balancing nodes of the model collaborative operation framework according to the analysis of the tasks to be executed. The processing flow is as follows: Analyze the tasks to be executed to determine the computational complexity, data volume, and real-time requirements of the tasks to be executed, and determine the corresponding relevant performance indicators of each of the several load balancing nodes of the model collaborative operation framework; Calculate the task allocation suitability of the tasks to be executed on the corresponding load balancing nodes respectively based on the computational complexity, the data volume, the real-time requirements, and the relevant performance indicators, so as to allocate the tasks to be executed to the target load balancing node according to the task allocation suitability; The target load balancing node is the load balancing node corresponding to the target task allocation suitability among the several load balancing nodes, and the target task allocation suitability is the task allocation suitability with the smallest value among the task allocation suitabilities of the tasks to be executed corresponding to each load balancing node. That is to say, analyze the tasks to be executed to determine the computational complexity, data volume, and real-time requirements of the tasks to be executed, and at the same time, monitor the resource status of edge devices, edge servers, and the cloud in real time, including but not limited to CPU, memory, network bandwidth, etc., so as to evaluate the processing capabilities of each load balancing node based on the above dynamic data, determine the corresponding relevant performance indicators of each of the several load balancing nodes of the model collaborative operation framework, so as to calculate the task allocation suitability of the tasks to be executed on the corresponding load balancing nodes respectively according to the determined above elements. That is, for each task to be executed, calculate its task allocation suitability on all available nodes. Since the smaller the value of the task allocation suitability, the higher the suitability of the task to be executed on this node, that is, the more suitable the node is to execute this task, the task to be executed can be allocated to the load balancing node corresponding to the target task allocation suitability among the several load balancing nodes, where the target task allocation suitability is the task allocation suitability with the smallest value among the task allocation suitabilities of the tasks to be executed corresponding to each load balancing node. The specific calculation formula of the task allocation suitability can be: ; Among them, C represents the computing complexity; D represents the data volume; T represents the real-time requirement; Node_Capability is the processing capability, that is, the computing capability of the load balancing node, which can be expressed as the number of operations that can be executed per second or the benchmark score; Node_Storage is the storage capability, that is, the storage capacity of the load balancing node, in bytes; Node_Latency is the latency characteristic, that is, the average latency of the load balancing node to process tasks, which can be expressed as a normalized value (between 0 and 1), where 0 represents the lowest latency. In this way, after determining the task priority in this embodiment, the tasks to be executed are dynamically allocated according to the task priority, which can optimize the resource utilization rate and reduce the latency; building a model collaborative operation framework composed of edge devices, edge servers and the cloud can achieve efficient collaboration of data among edge devices, edge servers and the cloud, ensure seamless communication between nodes, and then execute the tasks to be executed more efficiently, improving the processing efficiency; reasonably allocating tasks to edge devices, edge servers and the cloud according to the resource status and task characteristics can avoid single-point overload and resource waste.

[0034] Step S14: Use a preset data cache management algorithm to manage the relevant parameter data required for running the optimized model, so as to pre-load the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative operation framework; the data to be used is the data that needs to be used next predicted by the preset data cache management algorithm during the process of running the optimized model.

[0035] In this embodiment, a preset data cache management algorithm is used to manage the relevant parameter data required for running the optimized model, so as to pre-load the data to be used in the relevant parameter data into the cache of the edge device in advance, reduce the data transmission latency, and improve the task processing efficiency, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative operation framework. Among them, the relevant parameter data includes the model parameters and intermediate results of the optimized model; the data to be used is the data that needs to be used next predicted by the preset data cache management algorithm during the process of running the optimized model.

[0036] It should be noted that the memory cache can store data with high access frequency and high importance, such as the key parameters of the currently executing task, frequently interacted data, etc., to ensure fast access; the flash cache can store data with medium access frequency and relatively high importance, such as the model parameters used recently, relatively important intermediate results, etc., to ensure relatively fast read and write speed and long-term preservation; the cloud cache can store data with low access frequency and relatively low importance, such as historical data, backup data, etc., so as to obtain it on demand through the network and save the storage resources of edge devices. Therefore, the processing flow of managing the relevant parameter data using the preset data cache management algorithm is as follows: calculating the data access frequency of the relevant parameter data required to run the optimized model using the exponential weighted moving average method, and determining the importance of the relevant parameter data; caching the relevant parameter data into the target cache location according to the importance and the data access frequency; the target cache location includes the memory cache location, the flash cache location, and the cloud cache location. That is, the exponential weighted moving average method is used to calculate the data access frequency of the relevant parameter data and determine the importance of the relevant parameter data. Among them, the specific calculation formula of the data access frequency can be: ; Among them, is the access frequency at time t. If there is no historical data for this data, its value is set to 0.5. If there is historical data, the actual data at t - 1 is taken; α is the smoothing coefficient, which determines the weight distribution of the historical AF value and the current access times. The default value is 0.2. After running for a period of time, the absolute value of the slope of the access frequency change curve in the previous 30 minutes is taken as the smoothing coefficient; is the data access times at time t. After determining the data access frequency and the importance, the relevant parameter data is cached into the memory cache location, the flash cache location, or the cloud cache location according to the importance and the data access frequency. For example, when the data access frequency of the current data is not less than 0.7 and the importance is expressed as high, the current data can be stored in the memory cache location; when the data access frequency of the current data is lower than 0.7 and not less than 0.3, and the importance is expressed as high or medium, the current data can be stored in the flash cache location; when the data access frequency of the current data is lower than 0.3 and the importance is expressed as low, the current data can be stored in the cloud cache location.

[0037] It should be further pointed out that after managing the relevant parameter data, this embodiment can analyze user behavior and environmental perception information through machine learning algorithms to predict the user's possible next operations, and pre-load the data to be used in the relevant parameter data into the cache of the edge device in advance. The processing flow is as follows: Extract and analyze the features of the cache data in the target cache location to predict the predicted access frequency of the cache data in the next step through a preset autoregressive moving average model; Determine the corresponding cache location score according to the relevant information of the cache data stored in the target cache location; Calculate the cache value of the cache data based on the cache location score, the predicted access frequency, and the importance of the cache data, so as to adjust the cache location of the cache data according to the cache value, and at the same time use the user behavior prediction data output by the preset user behavior model and the current state of the model collaborative operation framework to determine the data to be used in the cache data, and pre-load the data to be used into the memory cache location of the edge device. Among them, the features include but are not limited to the access history, relevance, timeliness, etc. of the cache data; the preset autoregressive moving average model is used to predict the future access frequency of the data; the relevant information includes storage cost and access speed; the preset user behavior model is used to predict the data that the user may need at a specific time point, and can be a model based on time series analysis, including a Markov chain model, and the transition probability matrix in the Markov chain model is obtained through historical data statistics. That is, extract and analyze the features of the cache data in the target cache location to extract the access time series of the data, and establish a preset autoregressive moving average model (i.e., ARIMA model, Autoregressive Integrated Moving Average Model) to predict the predicted access frequency of the cache data in the next step through the preset autoregressive moving average model. Among them, the formula of the preset autoregressive moving average model can be specifically: ; where B is the lag operator; is obtained by performing i consecutive operations on B, indicating an i-order lag on the time series; φ is the autoregressive coefficient, indicating the influence of the past p-period observations on the current value; θ is the moving average coefficient, indicating the influence of the past q-period prediction errors on the current value; d is the order of differencing; is the white noise error. Subsequently, determine the corresponding cache location score according to the relevant information of the cache data stored in the target cache location, and calculate the cache value of the cache data based on the cache location score, the predicted access frequency, and the importance of the cache data. The specific formula for calculating the cache value can be: ; Among them, Predicted_AF is the predicted access frequency, that is, the future data access frequency predicted by the ARIMA model; V represents the cache value; Cache_Location_Score is the cache location score, that is, the score calculated according to the storage cost and access speed of data in different cache locations. After obtaining the cache value, the cache location and retention strategy of the data can be determined according to the value of the cache value. It can be understood that the higher the value of the cache value, the higher the access frequency, importance, and cache location score of the data. Therefore, for the data with the highest cache value, it can be preferentially stored in the memory cache location for fast access; the data with a medium cache value has a relatively low access frequency and importance, but still requires a relatively fast access speed, and is suitable for storage in the flash cache location; the data with the lowest cache value has a low access frequency and low importance, and can be stored in the cloud cache location to save the storage resources of the edge device. At the same time, this embodiment can use the user behavior prediction data output by the preset user behavior model and the current state of the model collaborative operation framework to determine the data to be used in the cached data, so as to pre-load the data to be used into the memory cache location of the edge device. Among them, the user behavior prediction data is obtained by outputting using the preset user behavior model. This embodiment can collect user behavior data, including user interaction records with the system such as clickstream data, query history, operation time, etc. Through the user behavior data, the preset user behavior model can be constructed to analyze the usage habits and preferences of users. For example, in an industrial scenario, this embodiment can record the adjustment frequency and mode of device parameters by device operators at different time periods. Subsequently, the collected user behavior data is used to train the original machine learning model to predict the possible next operations of users. For example, by analyzing the adjustment mode of device parameters by the operator in the past week, this embodiment can predict which parameters the operator may need to adjust on the next working day and load these parameters into the cache in advance. It can be understood that this embodiment can continuously monitor the access situation of data and the state to dynamically adjust the cache management strategy to ensure that the cache is always in an optimal state. For example, for data predicted to be accessed soon, it can be pre-loaded into the memory cache; for data that has not been accessed for a long time and has low importance, it can be migrated from the memory cache to the flash cache or cloud cache.

[0038] It should be noted that in this embodiment, through the model collaborative operation framework for collaborative work, the edge server in the model collaborative operation framework can assist in processing some emergency tasks to ensure the operation of the optimized large model. Therefore, the processing flow in the process of performing the corresponding to-be-executed tasks based on the to-be-used data can also be as follows: Continuously monitor the task queue of the edge server; when it is detected that there is an emergency task in the task queue, pause or adjust other tasks in the task queue except the emergency task; determine the corresponding resource allocation ratio according to the target task priority of the emergency task, and determine the target computing resource amount of the emergency task and the current available computing resource amount of the edge server; calculate the resource allocation amount of the emergency task based on the resource allocation ratio, the target computing resource amount, and the current available computing resource amount, so as to allocate the corresponding computing resources from other tasks according to the obtained resource allocation amount for executing the emergency task by using the computing resources; if after allocating the computing resources, it is detected that the current edge server cannot execute the emergency task through the computing resources, unload the target computing link of the emergency task to the cloud to use the cloud to assist the edge server in executing the emergency task. Among them, the target computing link is a key computing link. That is, the edge server manages tasks in the form of a task queue, and the task processing order follows the principle of high to low priority. Continuously monitor the task queue of the edge server. When it is detected that there is an emergency task such as equipment failure warning processing in the task queue, pause or adjust other tasks in the task queue except the emergency task, insert the emergency task at the head of the task queue, and at the same time determine the corresponding resource allocation ratio according to the target task priority of the emergency task, and determine the target computing resource amount of the emergency task and the current available computing resource amount of the edge server, so as to calculate the resource allocation amount of the emergency task based on the resource allocation ratio, the target computing resource amount, and the current available computing resource amount. The specific calculation formula of the resource allocation amount can be: ; Among them, Emergency_Task_Requirement represents the target computing resource amount, that is, the computing resource amount required for the emergency task; Current_Resource is the current available computing resource amount, that is, the computing resource amount that the edge server can currently use to process tasks; Resource_Allocation_Ratio represents the resource allocation ratio, that is, the ratio of allocating resources from other tasks. After obtaining the resource allocation amount, in this embodiment, corresponding computing resources can be temporarily allocated from the other tasks according to the resource allocation amount to ensure that the emergency task has sufficient computing resources, and then the emergency task is executed using the computing resources. In addition, if after allocating the computing resources, it is detected that the current edge server cannot execute the emergency task through the computing resources, that is, the resources are still insufficient, communication with the cloud can be quickly established to offload the target computing link of the emergency task to the cloud, so as to complete the calculation using the powerful computing power of the cloud and quickly feedback the obtained calculation results to the edge device, thereby assisting the edge server to execute the emergency task. In this way, this embodiment uses the results output by the preset user behavior model to load the possible required data and model parameters from the cloud or low-speed storage into the high-speed cache of the edge device in advance, which can significantly reduce data access latency, improve task processing efficiency and response speed; classify and cache data according to the access frequency and importance of the data, optimizing the user experience; continuously monitor the data access situation and the status of the model collaborative operation framework to dynamically adjust the cache management strategy to ensure that the cache is always in the optimal state; assist in processing some emergency tasks through the edge server in the model collaborative operation framework to ensure the operation of the optimized large model and improve the operation efficiency.

[0039] As can be seen from the above, in the embodiment of the present application, the large model to be run is first preprocessed, and the preprocessed large model to be run is optimized according to the hardware characteristics of the edge device. At the same time, the task priority of the task to be executed by the optimized model is determined according to the resource situation of the current edge device, so as to allocate the task to be executed to the model collaborative operation framework according to the task priority. During the allocation process, the task to be executed is deployed to the load balancing node of the model collaborative operation framework according to the analysis result of the task to be executed. At the same time, the preset data cache management algorithm is used to manage the relevant parameter data, so as to preload the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative operation framework.In this way, through the above process of the embodiments of the present application, on the one hand, when a large model needs to run on an edge device, the large model to be run is preprocessed through techniques such as quantization, pruning, knowledge distillation, structural reparameterization, and sparsity utilization, which can reduce the number of parameters, computational complexity, and storage requirements of the large model while maintaining high accuracy, making it more suitable for running on resource-constrained edge devices; on the other hand, the preprocessed large model to be run is customized and optimized based on the hardware characteristics of the edge device to fully utilize the computing power of the hardware, achieve efficient hardware acceleration, reduce latency, increase computational throughput, and further improve the inference speed of the model; on the one hand, based on the resource situation of the current edge device and the relevant information of the task to be executed, the task priority is determined by comprehensively considering various factors such as task type, task priority processing requirements, task urgency, and task resource dependency, which can optimize resource allocation, meet business requirements, ensure the timely execution of critical tasks, and improve the operation efficiency of the large model; on the one hand, after determining the task priority, the tasks to be executed are dynamically allocated according to the task priority, which can optimize resource utilization and reduce latency; on the one hand, a model collaborative operation framework built by edge devices, edge servers, and the cloud is established, which can achieve efficient collaboration of data among edge devices, edge servers, and the cloud, ensure seamless communication between nodes, and thus execute the tasks to be executed more efficiently and improve the processing efficiency; on the one hand, tasks are reasonably allocated to edge devices, edge servers, and the cloud according to the resource status and task characteristics to avoid single-point overload and resource waste; on the one hand, the results output by the preset user behavior model are used to pre-load the data and model parameters that may be needed from the cloud or low-speed storage into the high-speed cache of the edge device in advance, which can significantly reduce data access latency, improve task processing efficiency and response speed; on the one hand, the data is hierarchically cached according to the access frequency and importance of the data, optimizing the user experience; on the one hand, the access situation of the data and the status of the model collaborative operation framework are continuously monitored to dynamically adjust the cache management strategy to ensure that the cache is always in the optimal state; on the other hand, the edge server in the model collaborative operation framework is used to assist in processing some urgent tasks to ensure the operation of the optimized large model, improve the operation efficiency, and thus optimize the operation process of the large model on the edge device to solve the application problems caused by the deficiencies of the current technology.

[0040] Based on the foregoing embodiments, it can be seen that through the method of the present application, in the process of executing the corresponding task to be executed using the data to be used that needs to be used next during the process of optimizing the large model based on the prediction, the data can be transmitted based on the preset data transmission mechanism when the data needs to be transmitted in the model collaborative operation framework to improve the data transmission efficiency. For this reason, this embodiment elaborates in detail on how to transmit data based on the preset data transmission mechanism. Refer to Figure 2 As shown, the embodiment of the present invention discloses a method for running a large model based on an edge device, including: Step S21: When the data to be used needs to be transmitted in the model collaborative operation framework, preprocessing operations of compressing and encapsulating the data to be used are performed to obtain processed data.

[0041] In this embodiment, when the data to be used needs to be transmitted in the model collaborative operation framework, the data to be used can be first compressed to reduce the amount of data transmitted, and preprocessing operations of encapsulation are performed to facilitate transmission and parsing, thereby obtaining processed data. Among them, the data to be used is the data that needs to be used next in the process of running and optimizing the large model predicted by the preset data cache management algorithm; the model collaborative operation framework is a framework constructed by edge devices, edge servers, and the cloud. Specifically, when performing the compression operation on the data to be used, this embodiment can adopt an improved LZ77 algorithm (a dictionary-based lossless compression algorithm), that is, by adding steps of a hash table and preprocessing to improve the compression efficiency, and counting the frequently occurring strings in the data to be used to establish a dictionary, so that when compressing, the dictionary can be used to quickly find matching strings and replace them with dictionary indexes, realizing efficient compression of the data to be used and reducing the amount of data transmitted. At the same time, when performing the preprocessing operation of encapsulating the data to be used, the data to be used can be encapsulated into data packets in a specific format, where the data packets contain information such as data content, source address, target address, and priority, facilitating transmission and parsing. In this way, in this embodiment, when the data to be used needs to be transmitted in the model collaborative operation framework, preprocessing operations of compressing and encapsulating it are performed to more efficiently utilize the bandwidth by reducing the amount of data transmitted, reducing the transmission time, and improving the efficiency of data transmission and data parsing through encapsulation.

[0042] Step S22: Determine the data priority and data type of the processed data, and calculate the target channel bandwidth for transmitting the processed data based on the task priority of the task to be executed and the total available bandwidth.

[0043] In this embodiment, the data priority and data type of the processed data are determined, and the target channel bandwidth for transmitting the processed data is calculated according to the task priority of the task to be executed corresponding to the processed data and the total available bandwidth. Among them, the specific calculation formula for the target channel bandwidth can be: ; Among them, is the bandwidth allocated to the i-th channel, is the task priority of the task to be executed, is the sum of all task priorities, is the total available bandwidth. It can be understood that in this embodiment, multiple data transmission channels are established for data transmission. By calculating the target channel bandwidth, a suitable channel can be selected for transmitting the data to be processed according to the target channel bandwidth. In this way, in this embodiment, the target channel bandwidth for transmitting the processed data is calculated based on the task priority of the task to be executed and the total available bandwidth, realizing flexible determination of the target channel bandwidth according to the characteristics of the task and facilitating adaptation to the complex and changeable edge computing environment.

[0044] Step S23: Determine the target transmission channel for the processed data according to the data priority, the data type, and the target channel bandwidth, so as to transmit the processed data using the target transmission channel.

[0045] In this embodiment, the target transmission channel for the processed data is determined according to the data priority, the data type, and the target channel bandwidth, so as to transmit the processed data using the target transmission channel. For example, for data with high real-time requirements, a high-speed and low-latency data transmission channel can be selected as its target transmission channel to transmit the data using the high-speed and low-latency data transmission channel; for batch data, a data transmission channel with a larger bandwidth but slightly higher latency can be selected as its target transmission channel to transmit the data using the data transmission channel with a larger bandwidth but slightly higher latency, thereby realizing dynamic allocation of the transmission channel according to the task priority and data type. In this way, in this embodiment, the target transmission channel for the processed data is determined according to the data priority, the data type, and the target channel bandwidth, which can make more efficient use of the bandwidth and improve the data transmission efficiency.

[0046] Step S24: During the process of transmitting the processed data, calculate the check value of the processed data using a preset data check algorithm to determine whether an abnormality occurs in the transmission of the processed data according to the check value.

[0047] In this embodiment, during the process of transmitting the processed data using the target transmission channel, the check value of the processed data can be calculated using a preset data check algorithm, so as to determine whether an abnormality occurs in the transmission of the processed data according to the check value, facilitating the discovery of transmission errors and timely retransmission to ensure the integrity of the data. Among them, the preset data check algorithm can be the CRC-32 check algorithm.

[0048] Specifically, in this embodiment, the check value of the processed data can be calculated during the transmission process, and verified at the receiving end according to the check value. If the verification fails, it means that the transmission of the current processed data is abnormal, and the retransmission strategy can be determined according to the importance and urgency of the processed data. Among them, for the importance level, if the importance level of the processed data with an abnormality is extremely high, it means that its loss may cause major losses such as system crashes. Therefore, the retransmission operation can be immediately triggered. At the same time, in order to ensure the success of the retransmission of this type of data, a backup channel can be established for parallel transmission; if the importance level of the processed data with an abnormality is relatively low, it means that the data has little impact on the system and can be processed when idle or directly discarded. For the urgency level, if the processed data with an abnormality has extremely high requirements for real-time performance, it means that it needs to arrive within a strict time limit. Therefore, the retransmission timeout time can be dynamically adjusted and multiple retransmissions can be quickly attempted; if the processed data with an abnormality has relatively low requirements for real-time performance, a certain delay can be allowed or retried when idle. It should be noted that this embodiment can also adopt a data synchronization mechanism during the transmission process of the processed data to ensure data consistency. In this way, this embodiment can effectively reduce the data transmission error rate and improve the reliability of data transmission by using the data synchronization and verification mechanisms.

[0049] As can be seen from the above, when the data to be used needs to be transmitted in the model collaborative operation framework in the embodiment of the present application, a preprocessing operation of compressing and encapsulating the data to be used is first performed to obtain the processed data. Subsequently, the data priority and data type of the processed data are determined. At the same time, the target channel bandwidth for transmitting the processed data is calculated based on the task priority of the task to be executed and the total available bandwidth, so as to determine the target transmission channel of the processed data according to the data priority, the data type, and the target channel bandwidth, and then use the target transmission channel to transmit the processed data. In addition, during the process of transmitting the processed data, a preset data verification algorithm is used to calculate the verification value of the processed data, so as to determine whether an abnormality occurs in the transmission of the processed data according to the verification value. In this way, through the above process of the embodiment of the present application, on the one hand, when the data to be used needs to be transmitted in the model collaborative operation framework, a preprocessing operation of compressing and encapsulating it is performed to reduce the amount of data transmitted through compression, more efficiently utilize the bandwidth, and reduce the transmission time. Through encapsulation, the efficiency of data transmission and data parsing is improved; on the one hand, the target channel bandwidth for transmitting the processed data is calculated based on the task priority of the task to be executed and the total available bandwidth, realizing the flexible determination of the target channel bandwidth according to the characteristics of the task, facilitating adaptation to the complex and changeable edge computing environment; on the one hand, the target transmission channel of the processed data is determined according to the data priority, the data type, and the target channel bandwidth, which can more efficiently utilize the bandwidth and improve the efficiency of data transmission; on the other hand, the use of the data synchronization and verification mechanism can effectively reduce the data transmission error rate, improve the reliability of data transmission, and further improve the efficiency of data transmission.

[0050] Based on the previous embodiments, the present application discloses a method for running a large model based on an edge device, which can optimize the running process of the large model on the edge device to solve the application problems caused by the deficiencies of the current technology. Next, a method for running a large model based on an edge device in the intelligent factory application scenario will be described in detail.

[0051] At present, the edge devices (such as industrial controllers) of this embodiment need to run a complex large model to monitor the equipment status and product quality on the production line in real time. In this embodiment, the weight of the large model, that is, the equipment status monitoring model, is first compressed from a 32-bit floating point number to an 8-bit integer through quantization technology, reducing the model volume by about 75%. At the same time, redundant parameters are further removed through pruning technology, reducing the number of model parameters by 30%. Through knowledge distillation technology, the knowledge of the large model is transferred to a lightweight small model to ensure that the reasoning accuracy of the small model reaches more than 95% of the large model. The model is further compressed by using structural reparameterization and sparsity. Finally, the model is customized and optimized for the built-in FPGA of the industrial controller, further improving the reasoning speed. After these optimizations, the industrial controller can efficiently run the equipment status monitoring model in a low-resource environment, reducing the reasoning time by 20%, significantly improving production efficiency. It is then detected that the CPU and memory resources of the industrial controller are tight, but the network bandwidth is sufficient. Based on this information, this embodiment can offload complex reasoning tasks in the model (such as feature extraction and classification in deep learning) to the cloud, while retaining simple preprocessing tasks (such as data collection and preliminary screening) for local execution, while ensuring seamless switching of tasks between the industrial controller and the cloud, thereby optimizing resource utilization and reducing latency without reducing user experience.

[0052] In the daily operation process, the industrial controller of this embodiment analyzes the operation mode of the production line through the preloading mechanism, and finds that when the production line starts running at 8 am every day, it is necessary to call specific model parameters for equipment status monitoring. Based on this prediction, the model collaborative operation framework automatically loads the relevant model parameters from the cloud or flash cache to the memory cache at around 7:50 am every day. When the production line starts at 8 am, the model parameters have been loaded and there is no need to wait for the data to be loaded from the cloud, thereby significantly reducing the waiting time for model reasoning and improving production efficiency. In addition, multiple industrial controllers on the production line work together through the model collaborative operation framework. When the resources of an industrial controller are tight, this embodiment can offload some tasks to a nearby edge server, or further offload to the cloud. At the same time, through the data synchronization mechanism, the data consistency between the edge device and the cloud is ensured. For example, when a device on the production line fails, the edge device can quickly upload the fault data to the cloud for in-depth analysis, and the edge server can assist in handling some emergency tasks to ensure the normal operation of the production line. Through this model collaborative operation framework, smart factories can handle complex tasks more efficiently, improve production efficiency and equipment reliability.

[0053] Accordingly, see Figure 3 As shown, the embodiment of the present application also provides a large model operation device based on an edge device, including: The model optimization module 11 is used to preprocess the large model to be run and optimize the preprocessed large model to be run based on the hardware characteristics of the edge device, so as to obtain an optimized model; The task priority determination module 12 is used to determine the task priority of the to-be-executed task corresponding to the optimized model based on the resource situation of the current edge device; The task analysis module 13 is used to allocate the to-be-executed task to a pre-built model collaborative operation framework according to the task priority, and during the allocation process, analyze the to-be-executed task, so as to deploy the to-be-executed task to the load balancing node of the model collaborative operation framework according to the obtained analysis result; the model collaborative operation framework is a framework constructed by the edge device, the edge server, and the cloud; The data management module 14 is used to manage the relevant parameter data required for running the optimized model by using a preset data cache management algorithm, so as to preload the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding to-be-executed task based on the data to be used through the load balancing node of the model collaborative operation framework; the data to be used is the data predicted by the preset data cache management algorithm and needed to be used in the next step during the process of running the optimized model.

[0054] As can be seen from the above, in the embodiments of the present application, the large model to be run is first preprocessed, and the preprocessed large model to be run is optimized according to the hardware characteristics of the edge device. At the same time, the task priority of the task to be executed by the optimized model is determined according to the resource situation of the current edge device, so as to allocate the task to be executed to the model collaborative operation framework according to the task priority. During the allocation process, the task to be executed is deployed to the load balancing node of the model collaborative operation framework according to the analysis result of the task to be executed. At the same time, the preset data cache management algorithm is used to manage the relevant parameter data, so as to preload the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative operation framework. In this way, through the above process of the embodiments of the present application, optimizing the large model to be run based on the hardware characteristics of the edge device can give full play to the hardware capabilities of the edge device to improve the operation efficiency of the large model; determining the task priority of the task to be executed according to the resource situation of the edge device and dynamically allocating the task to be executed according to the task priority can fully and reasonably realize resource allocation and improve system performance; using the preset data cache management algorithm to manage the relevant parameter data and preloading the data to be used, loading the relevant data into the cache of the edge device in advance, reducing the waiting time for model inference, and improving the operation efficiency; constructing a model collaborative operation framework composed of edge devices, edge servers and the cloud can effectively solve the latency problem, and further optimize the operation process of the large model on the edge device to solve the application problems caused by the deficiencies of the current technology.

[0055] In some specific embodiments, the model optimization module 11 may specifically include: The parameter removal unit is used to compress the weights of the large model to be run by using a preset quantization technique, and remove redundant parameters in the large model to be run by using a preset pruning technique to obtain a pruned large model; The parameter adjustment unit is used to transfer the knowledge of the pruned large model to the small model by using a preset knowledge distillation technique, and redesign and adjust the model structure of the transferred model to obtain an adjusted model; The feature mining unit is used to mine the sparse features of the adjusted model by using a gradient-based sparse training method to obtain a compressed model with a sparse structure; Correspondingly, the model optimization module 11 may specifically include: The model optimization unit is used to optimize the compressed model with a sparse structure based on the hardware characteristics of the edge device.

[0056] In some specific embodiments, the model optimization module 11 may specifically include: A module mapping unit, configured to map a target computing module in the to-be-run large model after preprocessing to the field programmable gate array by utilizing the characteristics of the field programmable gate array in the edge device; A frequency calculation unit, configured to calculate the computing frequency of a computing engine corresponding to the target computing module based on the clock speed of the field programmable gate array and the pipeline delay rate, so as to design the computing engine according to the computing frequency; A sub-batch size calculation unit, configured to, if it is detected that a graphics processor is configured in the edge device, utilize the single instruction multiple data architecture of the graphics processor to calculate the sub-batch size based on the total batch size of the batch data and the number of cores of the graphics processor, so as to split the batch data according to the sub-batch size and respectively allocate the split data to the cores of the graphics processor for parallel computing, so as to obtain an optimized model.

[0057] In some specific embodiments, the task priority determination module 12 may specifically include: A processing requirement weight determination unit, configured to determine a corresponding task type weight according to the task type of the to-be-executed task corresponding to the optimized model, and determine a corresponding task priority processing requirement weight based on whether the to-be-executed task meets the user-defined task priority processing condition; A dependency degree weight determination unit, configured to determine a corresponding urgency weight through the urgency of the to-be-executed task, and determine a corresponding resource dependency weight according to the degree of dependence of the to-be-executed task on resources; A task priority calculation unit, configured to calculate the task priority of the to-be-executed task based on the task type weight, the task priority processing requirement weight, the urgency weight, and the resource dependency weight; Correspondingly, the task analysis module 13 may specifically include: A first task allocation unit, configured to, when the task priority is not lower than a first preset priority, allocate the to-be-executed task to the edge device in a pre-established model collaborative operation framework; A second task allocation unit, configured to, when the task priority is not lower than a second preset priority and lower than the first preset priority, allocate the to-be-executed task to the edge server in the model collaborative operation framework; A third task allocation unit, configured to, when the task priority is lower than the second preset priority, allocate the to-be-executed task to the cloud in the model collaborative operation framework; Wherein, the first preset priority is higher than the second preset priority.

[0058] In some specific embodiments, the large model running device based on edge devices may further include: An information registration unit, configured to, after the edge device and the edge server in the model collaborative running framework are started, respectively register their respective resource information with the cloud in the model collaborative running framework through the edge device and the edge server, so that the cloud manages the allocation of the tasks to be executed according to the resource information and the task priorities; the edge device, the edge server, and the cloud in the model collaborative running framework communicate with each other through a preset communication protocol.

[0059] In some specific embodiments, the task analysis module 13 may specifically include: An index determination unit, configured to analyze the task to be executed to determine the computational complexity, data volume, and real-time requirements of the task to be executed, and determine the corresponding relevant performance indicators of several load balancing nodes in the model collaborative running framework; A fitness calculation unit, configured to calculate the task allocation fitness of the task to be executed on the corresponding load balancing node respectively based on the computational complexity, the data volume, the real-time requirements, and the relevant performance indicators, so as to allocate the task to be executed to the target load balancing node according to the task allocation fitness; the target load balancing node is the load balancing node corresponding to the target task allocation fitness among the several load balancing nodes, and the target task allocation fitness is the task allocation fitness with the smallest value among the task allocation fitnesses of the task to be executed corresponding to each load balancing node.

[0060] In some specific embodiments, the data management module 14 may specifically include: An importance determination unit, configured to calculate the data access frequency of the relevant parameter data required to run the optimized model by using the exponentially weighted moving average method, and determine the importance of the relevant parameter data; A data caching unit, configured to cache the relevant parameter data to the target cache location according to the importance and the data access frequency; the target cache location includes a memory cache location, a flash cache location, and a cloud cache location; A data analysis unit, configured to perform feature extraction and analysis on the cached data in the target cache location to predict the next predicted access frequency of the cached data through a preset autoregressive moving average model; A score determination unit, configured to determine the corresponding cache location score according to the relevant information of the cached data stored in the target cache location; A value calculation unit, configured to calculate the cache value of the cache data based on the cache location score, the predicted access frequency, and the importance level of the cache data, so as to adjust the cache location of the cache data according to the cache value. Meanwhile, the unit determines the data to be used in the cache data by using the user behavior prediction data output by a preset user behavior model and the current state of the model collaborative operation framework, and preloads the data to be used into the memory cache location of the edge device.

[0061] In some specific embodiments, the large model operation device based on an edge device may further include: A data preprocessing unit, configured to perform preprocessing operations of compressing and encapsulating the data to be used when the data to be used needs to be transmitted in the model collaborative operation framework, so as to obtain processed data; A bandwidth calculation unit, configured to determine the data priority and data type of the processed data, and calculate the target channel bandwidth for transmitting the processed data based on the task priority of the to-be-executed task and the total available bandwidth; A channel determination unit, configured to determine the target transmission channel of the processed data according to the data priority, the data type, and the target channel bandwidth, so as to transmit the processed data by using the target transmission channel; Moreover, during the process of transmitting the processed data, a preset data verification algorithm is used to calculate the verification value of the processed data, so as to determine whether an abnormality occurs in the transmission of the processed data according to the verification value.

[0062] In some specific embodiments, the large model operation device based on an edge device may further include: A queue monitoring unit, configured to continuously monitor the task queue of the edge server; A task adjustment unit, configured to pause or adjust other tasks in the task queue except the urgent task when it is detected that there is an urgent task in the task queue; A resource amount determination unit, configured to determine the corresponding resource allocation ratio according to the target task priority of the urgent task, and determine the target computing resource amount of the urgent task and the current available computing resource amount of the edge server; A deployment amount calculation unit, configured to calculate the resource deployment amount of the urgent task based on the resource allocation ratio, the target computing resource amount, and the current available computing resource amount, so as to allocate corresponding computing resources from the other tasks according to the obtained resource deployment amount, and use the computing resources to execute the urgent task; A link unloading unit is configured to unload the target computing link of the emergency task to the cloud if it is detected that the current edge server cannot execute the emergency task through the computing resources after allocating the computing resources, so as to use the cloud to assist the edge server in executing the emergency task.

[0063] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the method for running a large model based on an edge device disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0064] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0065] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0066] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of implementing the method for running a large model based on an edge device executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0067] Further, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned large model operation method based on edge devices. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0068] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0069] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0070] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0071] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0072] The above has introduced the technical solution provided by the present application in detail. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for running a large model based on an edge device, characterized in that, Including: Preprocessing the large model to be run and optimizing the preprocessed large model to be run based on the hardware characteristics of the edge device to obtain an optimized model; Determining the task priority of the task to be executed corresponding to the optimized model based on the current resource situation of the edge device; Allocating the task to be executed to a pre-built model collaborative running framework according to the task priority, and during the allocation process, analyzing the task to be executed to deploy the task to be executed to the load balancing node of the model collaborative running framework according to the obtained analysis result; the model collaborative running framework is a framework built by the edge device, the edge server, and the cloud; Managing the relevant parameter data required for running the optimized model by using a preset data cache management algorithm to preload the data to be used in the relevant parameter data into the edge device, so that the optimized model can execute the corresponding task to be executed based on the data to be used through the load balancing node of the model collaborative running framework; The data to be used is the data predicted by the preset data cache management algorithm and needed to be used next during the running of the optimized model.

2. The method for running a large model based on an edge device according to claim 1, wherein, The preprocessing of the large model to be run includes: Compressing the weights of the large model to be run by using a preset quantization technique and removing redundant parameters in the large model to be run through a preset pruning technique to obtain a pruned large model; Transferring the knowledge of the pruned large model to a small model through a preset knowledge distillation technique and redesigning the model structure and adjusting the parameters of the transferred model to obtain an adjusted model; Mining the sparse characteristics of the adjusted model by using a gradient-based sparse training method to obtain a compressed model with a sparse structure; Correspondingly, the optimization of the preprocessed large model to be run based on the hardware characteristics of the edge device includes: Optimizing the compressed model with a sparse structure based on the hardware characteristics of the edge device.

3. The method for running a large model based on an edge device according to claim 1, wherein The optimization of the preprocessed large model to be run based on the hardware characteristics of the edge device to obtain an optimized model includes: Using the characteristics of the field programmable gate array in the edge device to map the target computing module in the preprocessed large model to be run to the field programmable gate array; Calculating the computing frequency of the computing engine corresponding to the target computing module based on the clock speed and pipeline delay rate of the field programmable gate array, and designing the computing engine according to the computing frequency; If it is detected that a graphics processor is configured in the edge device, using the single instruction multiple data architecture of the graphics processor to calculate the sub-batch size based on the total batch size of the batch data and the number of cores of the graphics processor, so as to split the batch data according to the sub-batch size and allocate the split data to the cores of the graphics processor for parallel computing to obtain an optimized model.

4. The method for running a large model based on an edge device according to claim 1, wherein The determination of the task priority of the task to be executed corresponding to the optimized model based on the current resource situation of the edge device includes: Determine the corresponding task type weight according to the task type of the to-be-executed task corresponding to the optimized model, and determine the corresponding task priority processing requirement weight based on whether the to-be-executed task meets the user-defined task priority processing condition; Determine the corresponding urgency weight according to the urgency of the to-be-executed task, and determine the corresponding resource dependence weight according to the dependence degree of the to-be-executed task on resources; Calculate the task priority of the to-be-executed task based on the task type weight, the task priority processing requirement weight, the urgency weight, and the resource dependence weight; Correspondingly, the allocation of the to-be-executed task to the pre-built model collaborative operation framework according to the task priority includes: When the task priority is not lower than the first preset priority, allocate the to-be-executed task to the edge device in the pre-built model collaborative operation framework; When the task priority is not lower than the second preset priority and lower than the first preset priority, allocate the to-be-executed task to the edge server in the model collaborative operation framework; When the task priority is lower than the second preset priority, allocate the to-be-executed task to the cloud in the model collaborative operation framework; Wherein, the first preset priority is higher than the second preset priority.

5. The method for running a large model based on an edge device according to claim 1, wherein Before allocating the to-be-executed task to the pre-built model collaborative operation framework according to the task priority, it further includes: After the edge device and the edge server in the model collaborative operation framework are started, register their respective resource information to the cloud in the model collaborative operation framework through the edge device and the edge server respectively, so that the cloud manages the allocation of the to-be-executed task according to the resource information and the task priority; The edge device, the edge server, and the cloud in the model collaborative operation framework communicate with each other through a preset communication protocol.

6. The method for running a large model based on an edge device according to claim 1, wherein The analysis of the to-be-executed task to deploy the to-be-executed task to the load balancing node of the model collaborative operation framework according to the obtained analysis result includes: Analyze the to-be-executed task to determine the computational complexity, data volume, and real-time requirement of the to-be-executed task, and determine the corresponding relevant performance indicators of several load balancing nodes of the model collaborative operation framework; Calculate the task allocation fitness of the to-be-executed task on the corresponding load balancing node based on the computational complexity, the data volume, the real-time requirement, and the relevant performance indicators, and allocate the to-be-executed task to the target load balancing node according to the task allocation fitness; The target load balancing node is the load balancing node corresponding to the target task allocation fitness among the several load balancing nodes, and the target task allocation fitness is the task allocation fitness with the smallest value among the task allocation fitnesses of the to-be-executed task corresponding to each load balancing node.

7. The method for running a large model based on an edge device according to claim 1, wherein Managing the relevant parameter data required for running the optimized model by using a preset data cache management algorithm to preload the data to be used in the relevant parameter data into the edge device, including: Calculating the data access frequency of the relevant parameter data required for running the optimized model by using the exponential weighted moving average method, and determining the importance level of the relevant parameter data; Caching the relevant parameter data into a target cache location according to the importance level and the data access frequency; the target cache location includes a memory cache location, a flash cache location, and a cloud cache location; Performing feature extraction and analysis on the cached data in the target cache location to predict the next predicted access frequency of the cached data through a preset autoregressive moving average model; Determining a corresponding cache location score according to the relevant information of the cached data stored in the target cache location; Calculating the cache value of the cached data based on the cache location score, the predicted access frequency, and the importance level of the cached data, so as to adjust the cache location of the cached data according to the cache value, and at the same time determining the data to be used in the cached data by using the user behavior prediction data output by a preset user behavior model and the current state of the model collaborative operation framework, so as to preload the data to be used into the memory cache location of the edge device.

8. The method for running a large model based on an edge device according to claim 1, wherein During the process of executing the corresponding to-be-executed task based on the data to be used, it further includes: When the data to be used needs to be transmitted in the model collaborative operation framework, performing a preprocessing operation of compressing and encapsulating the data to be used to obtain processed data; Determining the data priority and data type of the processed data, and calculating the target channel bandwidth for transmitting the processed data based on the task priority of the to-be-executed task and the total available bandwidth; Determining the target transmission channel of the processed data according to the data priority, the data type, and the target channel bandwidth, so as to transmit the processed data by using the target transmission channel; And, during the process of transmitting the processed data, calculating a check value of the processed data by using a preset data verification algorithm to determine whether an abnormality occurs in the transmission of the processed data according to the check value.

9. The method for running a large model based on an edge device according to claim 1, wherein, During the process of executing the corresponding to-be-executed task based on the data to be used, it further includes: Continuously monitoring the task queue of the edge server; When it is detected that there is an urgent task in the task queue, pausing or adjusting other tasks in the task queue except the urgent task; Determining a corresponding resource allocation ratio according to the target task priority of the urgent task, and determining the target computing resource amount of the urgent task and the current available computing resource amount of the edge server; Calculating the resource allocation amount of the urgent task based on the resource allocation ratio, the target computing resource amount, and the current available computing resource amount, so as to allocate corresponding computing resources from the other tasks according to the obtained resource allocation amount, so as to execute the urgent task by using the computing resources. If it is detected that the current edge server cannot execute the emergency task through the computing resources after allocating the computing resources, the target computing link of the emergency task is offloaded to the cloud to utilize the cloud to assist the edge server in executing the emergency task.

10. An operation device for large models based on edge devices, characterized in that, Including: A model optimization module for preprocessing the to-be-run large model and optimizing the preprocessed to-be-run large model based on the hardware characteristics of the edge device to obtain an optimized model; A task priority determination module for determining the task priority of the to-be-executed task corresponding to the optimized model based on the resource situation of the current edge device; A task analysis module for allocating the to-be-executed task to a pre-built model collaborative operation framework according to the task priority, and analyzing the to-be-executed task during the allocation process to deploy the to-be-executed task to the load balancing node of the model collaborative operation framework according to the obtained analysis result; the model collaborative operation framework is a framework built by the edge device, the edge server, and the cloud; A data management module for managing the relevant parameter data required for running the optimized model by using a preset data cache management algorithm to preload the data to be used in the relevant parameter data into the edge device so that the optimized model can execute the corresponding to-be-executed task based on the data to be used through the load balancing node of the model collaborative operation framework; The data to be used is the data predicted by the preset data cache management algorithm and needed to be used in the next step during the running of the optimized model.

11. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for executing the computer program to implement the method for running a large model based on an edge device according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the method for running a large model based on an edge device according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Task unloading and resource optimization method based on edge cache

    CN111552564A

  • Retrieval enhancement generation deployment method based on edge calculation

    CN118689491A

  • Cross-level data prefetching method and device, equipment and medium

    CN119003387A

  • Dynamic component loading system based on React framework

    CN119536836A

  • Cloud storage acceleration method and system based on edge computing

    CN119621686A