Load resource collaborative optimization method and device of edge intelligent drive, equipment and medium
By monitoring the resource status of edge devices in real time, dynamically adjusting task and model pruning strategies, and combining workload fragmentation technology, the drone target detection process is optimized, solving the problem of insufficient resource utilization in edge collaboration scenarios and achieving efficient and reliable target detection.
Patent Information
- Application Number
- CN202510953295.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-11-18
Smart Images

Figure CN120976789A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of target detection, and in particular to an edge intelligent driving load resource collaborative optimization method, device, equipment and medium. BACKGROUND
[0002] In recent years, unmanned aerial vehicles (UAVs) play an increasingly important role in post-disaster rescue, environmental monitoring and other scenarios. With the help of deep learning technology, UAVs can perform real-time perception tasks such as target detection and image segmentation, providing key information support for post-disaster decision-making. However, the post-earthquake site environment is complex and variable, and UAVs face severe challenges when performing perception tasks. On the one hand, the available power, storage space and computing resources of the UAVs will dynamically change with the increase of task intensity and duration; on the other hand, traditional lightweight strategies based on fixed parameter design are difficult to adapt to this dynamic changing hardware environment, resulting in decreased inference efficiency or excessive energy consumption.
[0003] In the field of deep learning, target detection technology has made significant progress, and models such as YOLO and Faster R-CNN have shown strong performance in various application scenarios. However, these models usually have high computational complexity, and direct deployment on resource-constrained UAVs can lead to increased latency, increased energy consumption, and other problems. To address these challenges, existing technologies mainly seek solutions from two directions: model compression technology and collaborative inference technology. Model compression reduces the number of model parameters and computations through pruning, quantization, and knowledge distillation, but existing methods mostly use static compression strategies that cannot adapt to dynamic changes in device resources. Collaborative inference uses multi-device distributed computing to share the computational load, but there is still room for optimization in task allocation and resource scheduling among heterogeneous devices.
[0004] Therefore, the existing solutions have the problems of insufficient utilization of heterogeneous resources and insufficient dynamic adaptation of inference models to device heterogeneity in edge collaborative scenarios. SUMMARY
[0005] The present application provides an edge intelligent driving load resource collaborative optimization method, device, equipment and medium, which solves the defects of insufficient utilization of heterogeneous resources and insufficient dynamic adaptation of inference models to device heterogeneity in edge collaborative scenarios in the prior art.
[0006] The present application provides an edge intelligent driving load resource collaborative optimization method, comprising the following steps: Real-time monitoring of each edge device to obtain current resource state data of each edge device, the current resource state data including current computing power, current network bandwidth, current remaining power and current memory capacity; According to the current resource state data of each edge device and the global task demand, a target detection model deployed on each edge device is subjected to structured pruning processing to obtain a pruned target detection model; The high-definition image captured by the unmanned aerial vehicle is uniformly fragmented, and an overlapping area of a preset size is retained to obtain a plurality of image fragments; According to the current resource state data of each edge device, the plurality of image fragments are assigned to a plurality of edge devices, so that each edge device applies a locally deployed target detection model to perform target inference on the assigned image fragments, and outputs a partition target inference result; The partition target inference result output by each edge device is obtained, and the partition target inference result output by each edge device is integrated to eliminate redundant detection boxes corresponding to the overlapping area, thereby obtaining a global target inference result.
[0007] According to the edge intelligent driving load resource collaborative optimization method provided by the application, the current resource state data of each edge device and the global task demand are used to determine the pruning rate corresponding to the target detection model deployed on each edge device, and the pruning rate corresponding to the target detection model deployed on each edge device is determined according to the current resource state data of each edge device and the global task demand. According to the current resource state data of each edge device and the global task demand, the pruning rate corresponding to the target detection model deployed on each edge device is determined. For each target detection model deployed on an edge device, the target detection model is subjected to structured pruning processing according to the pruning rate corresponding to the target detection model, to obtain a pruned target detection model.
[0008] According to the edge intelligent driving load resource collaborative optimization method provided by the application, the current resource state data of each edge device and the global task demand are used to determine the pruning rate corresponding to the target detection model deployed on each edge device, and the pruning rate corresponding to the target detection model deployed on each edge device is determined according to the current resource state data of each edge device and the global task demand. The current resource state data of each edge device and the global task demand are input into a resource demand optimization model to obtain the pruning rate corresponding to the target detection model deployed on each edge device output by the resource demand optimization model.
[0009] According to the edge intelligent driving load resource collaborative optimization method provided by the application, the method further comprises: A pruning rate and accuracy dataset and a pruning rate and energy consumption dataset are constructed. The pruning rate and accuracy dataset is input into an accuracy predictor to obtain the relationship between the pruning rate and the accuracy output by the accuracy predictor. The pruning rate and energy consumption dataset is input into an energy consumption predictor to obtain the relationship between the pruning rate and the energy consumption output by the energy consumption predictor. determining the relationship between the pruning rate and the model memory occupation; using the relationship between the pruning rate and the accuracy, the relationship between the pruning rate and the energy consumption, and the relationship between the pruning rate and the model memory occupation as constraint conditions, defining solving the pruning rate optimization problem by using a linear programming solver, and constructing a resource demand optimization model.
[0010] According to the edge intelligent driving load resource collaborative optimization method provided by the application, the multiple image slices are distributed to multiple edge devices according to the current resource state data of each edge device, and the method comprises the steps that: According to the current computing capacity and the current network bandwidth of each edge device, the computing time delay and the transmission time delay of each edge device are calculated; According to the computing time delay and the transmission time delay of each edge device, the maximum total time delay is calculated, the total time delay is the sum of the transmission time delays of all edge devices and the sum of the maximum values of the computing time delays of all edge devices; The slice distribution problem is modeled as an integer linear programming problem, the target is to minimize the maximum total time delay, and the constraint conditions are that the total size of the non-overlapping areas of all image slices is equal to the size of the high-definition image, and the number of slices of each edge device is a non-negative integer; The integer linear programming problem is solved to obtain a current slice distribution strategy; According to the current slice distribution strategy, the multiple image slices are distributed to multiple edge devices.
[0011] According to the edge intelligent driving load resource collaborative optimization method provided by the application, the partitioned target reasoning results output by each edge device are integrated, the redundant detection boxes corresponding to the overlapping areas are eliminated, and a global target reasoning result is obtained, and the method comprises the steps that: According to the coordinate information and the redundant area design at the time of slicing, the partitioned target reasoning results output by each edge device are mapped back to the original image coordinate system corresponding to the high-definition image; The redundant detection boxes corresponding to the overlapping areas are filtered using a non-maximum suppression algorithm, and the detection box with the highest confidence is retained to obtain a global target reasoning result.
[0012] According to the edge intelligent driving load resource collaborative optimization method provided by the application, the high-definition image photographed by the unmanned aerial vehicle is uniformly sliced, and a preset size of overlapping area is retained to obtain multiple image slices, and the method comprises the steps that: Based on the image slice size that can be processed by the edge device with the minimum current memory capacity, the minimum size of the image slice is determined; According to the size of the high-definition image and the minimum size of the image slice, the number of image slices is calculated; The step size of the image segment is calculated based on the minimum size of the image segment and the size of the overlapping area. The high-definition image is segmented according to the number of image segments and the segmentation step size to obtain multiple image segments.
[0013] The present invention also provides an edge-intelligent driven load resource collaborative optimization device, comprising the following modules: The monitoring module is used to monitor each edge device in real time and obtain the current resource status data of each edge device. The current resource status data includes the current computing power, current network bandwidth, current remaining power, and current memory capacity. The pruning module is used to perform structured pruning on the target detection model deployed on each edge device based on the current resource status data and global task requirements of each edge device, so as to obtain the pruned target detection model. The slicing module is used to uniformly slice the high-definition images captured by the drone, and retain the overlapping area of the preset size to obtain multiple image slices; The allocation module is used to allocate the multiple image slices to multiple edge devices according to the current resource status data of each edge device, so that each edge device applies the locally deployed target detection model to perform target inference on the allocated image slices and outputs the partition target inference results; An integration module is used to acquire the partition target inference results output by each edge device, integrate the partition target inference results output by each edge device, eliminate redundant detection boxes corresponding to the overlapping areas, and obtain global target inference results.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the load resource collaborative optimization method of any of the above-described edge intelligent driving methods.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the load resource collaborative optimization method of edge intelligence driving as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the load resource collaborative optimization method of any of the above-described edge intelligent driving methods.
[0017] This invention provides an edge-intelligent driven load resource collaborative optimization method, apparatus, device, and medium. By monitoring each edge device in real time and obtaining its current resource status data, it can dynamically adjust task allocation and model optimization strategies, ensuring that each edge device operates efficiently within its capabilities, thus improving the stability and reliability of target detection. Furthermore, based on the current resource status data and global task requirements, structured pruning of the target detection model reduces the number of model parameters and computational load, ensuring optimal performance and resource utilization efficiency across different devices and task scenarios. The pruned model maintains high accuracy while significantly reducing computational load, thereby improving inference speed and meeting real-time requirements. Further, by uniformly segmenting high-definition images captured by UAVs and retaining overlapping areas of a preset size, multiple image slices are obtained. This allows tasks to be distributed to multiple edge devices for parallel processing, significantly reducing the computational burden on individual devices compared to processing the entire image on a single device, thereby significantly improving computational efficiency and accelerating target detection. On the other hand, retaining overlapping areas of a preset size to support subsequent result fusion processing avoids the loss of target information due to fragmentation, thereby improving the accuracy of target detection. Furthermore, fragmentation processing can better handle these complex scenarios, improving the robustness of target detection. Moreover, dynamically allocating image fragments based on the current resource status data of the devices can achieve load balancing, preventing some devices from being overloaded while others are idle, thus improving the overall resource utilization of the system. Furthermore, by integrating the partitioned target inference results of each device and eliminating redundant detection boxes in overlapping areas, duplicate detection can be avoided, resulting in globally consistent target inference results and ensuring the accuracy and reliability of the final result. Therefore, the solution of this invention, by optimizing the image processing flow and adaptively pruning the target detection model, rationally utilizes the resources of multiple heterogeneous devices, not only improving the real-time performance of target detection but also enhancing the overall energy efficiency of the system and the utilization rate of edge devices. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the edge-intelligent driven load resource collaborative optimization method provided by the present invention.
[0020] Figure 2 This is a schematic diagram illustrating the relationship between accuracy and pruning rate provided by the present invention.
[0021] Figure 3 This is a schematic diagram illustrating the relationship between energy consumption and pruning rate provided by the present invention.
[0022] Figure 4 This is a schematic diagram illustrating the impact of sparsification training on the model pruning accuracy provided by the present invention.
[0023] Figure 5 This is a schematic diagram illustrating the generalization performance of the accuracy predictor provided by the present invention.
[0024] Figure 6 This is a schematic diagram illustrating the generalization performance of the energy consumption predictor provided by the present invention.
[0025] Figure 7 This is a schematic diagram of the edge-intelligent driven load resource collaborative optimization device provided by the present invention.
[0026] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0028] In recent years, drones have played an increasingly important role in disaster relief and environmental monitoring. Leveraging deep learning technology, drones can perform real-time perception tasks such as target detection and image segmentation, providing crucial information support for post-disaster decision-making. However, the complex and ever-changing environment at earthquake sites presents significant challenges for drones performing perception tasks. On the one hand, the available power, storage space, and computing resources of drones dynamically change with the intensity and duration of the task; on the other hand, traditional lightweight strategies based on fixed parameters struggle to adapt to this dynamic hardware environment, leading to decreased inference efficiency or excessive energy consumption.
[0029] In the field of deep learning, object detection technology has made significant progress, with models such as YOLO and Faster R-CNN demonstrating powerful performance in various application scenarios. However, these models typically have high computational complexity, and direct deployment on resource-constrained drones can lead to increased latency and energy consumption. To address these challenges, solutions are sought primarily from two directions: model compression technology and collaborative inference technology. Model compression reduces the number of model parameters and computational load through methods such as pruning, quantization, and knowledge distillation, but most existing methods employ static compression strategies, which cannot adapt to dynamic changes in device resources. Collaborative inference utilizes distributed computing across multiple devices to share the computational load, but there is still room for optimization in task allocation and resource scheduling among heterogeneous devices.
[0030] Therefore, existing solutions suffer from insufficient utilization of heterogeneous resources and inadequate dynamic adaptability of inference models to device heterogeneity in edge collaboration scenarios.
[0031] To address the aforementioned technical challenges, the following technical concept is adopted: combining load partitioning with resource adaptation techniques to achieve efficient collaborative target inference. First, workload sharding reduces the computational burden on individual devices, and multi-UAV collaborative computing optimizes resource utilization. In the image processing stage, high-definition images are uniformly sharded while retaining appropriate overlapping areas. Then, computational tasks are dynamically allocated based on the resource status of each device, allowing high-performance devices to handle more load, thus achieving a balanced distribution of computational resources. Simultaneously, a data-driven resource monitoring mechanism is constructed to track the current resource status data of edge devices in real time. Based on the current resource status data of each edge device and global task requirements, the model pruning strategy is dynamically adjusted to dynamically solve for the optimal pruning scheme. This allows the lightweight model to efficiently adapt to dynamically changing hardware environments, minimizing energy consumption while ensuring inference accuracy.
[0032] The technical solution of this application and how it solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The following is a combination of... Figures 1-7 This invention describes a load resource collaborative optimization method for edge-driven intelligent systems.
[0033] Figure 1 This is a flowchart illustrating the edge-intelligent driven load resource collaborative optimization method provided by the present invention, as shown below. Figure 1 As shown, the method includes the following: Step 101: Monitor each edge device in real time and obtain the current resource status data of each edge device. The current resource status data includes the current computing power, current network bandwidth, current remaining power, and current memory capacity. Step 102: Based on the current resource status data of each edge device and the global task requirements, perform structured pruning on the target detection model deployed on each edge device to obtain the pruned target detection model; Step 103: Divide the high-definition image captured by the drone into uniform slices, and retain the overlapping area of the preset size to obtain multiple image slices; Step 104: Based on the current resource status data of each edge device, allocate multiple image slices to multiple edge devices, so that each edge device can apply the locally deployed target detection model to perform target inference on the allocated image slices and output the partition target inference results. Step 105: Obtain the partition target inference results output by each edge device, integrate the partition target inference results output by each edge device, eliminate redundant detection boxes corresponding to overlapping areas, and obtain the global target inference results.
[0034] In practical applications, the execution entity of this edge-intelligent load resource collaborative optimization method can be an edge-intelligent load resource collaborative optimization device. There are various ways to implement this device. For example, it can be implemented through computer programs, such as application software; or, for example, chips. It can also be implemented as a medium storing relevant computer programs, such as a USB flash drive or cloud storage; or, it can be implemented through a physical device that integrates or installs relevant computer programs, such as a server or smart device.
[0035] The following explanation uses an edge-intelligent-driven load resource collaborative optimization device as the main body for implementing the edge-intelligent-driven load resource collaborative optimization method.
[0036] Specifically, within the collaborative reasoning architecture, this invention optimizes in two parts. The first part focuses on workload partitioning, determining the minimum size and number of image fragments. Based on network bandwidth and the computing power of each device, it calculates the number of image fragments allocated to each device while minimizing latency. First, the device acquiring the workload is designated as the master device, and the other devices as slave devices. The objective is to minimize the sum of computation latency and transmission latency, where the computation latency is the maximum among all devices, and the transmission latency is the latency for the master device to send all fragments.
[0037] in, For equipment i The computational delay, Main equipment to equipment i Transmission delay for sending fragments.
[0038] The second part focuses on resource adaptive pruning and deployment. First, it monitors device energy consumption and performs inference using target detection models with different pruning rates. This yields datasets on inference accuracy and pruning rate, as well as datasets on device energy consumption and pruning rate. Accuracy predictors and energy consumption predictors are then built using these datasets. During inference on a single device, a resource requirement optimization model is constructed based on the predicted workload, accuracy predictor, and energy consumption predictor. This model calculates the optimal pruning rate, pruning the target detection model to reduce computational load and accelerate inference, thereby reducing the computational energy consumption of each edge device.
[0039] Where N represents the number of edge devices, and K is the number of shards for the current device. This represents the pruning rate corresponding to the object detection model.
[0040] Specifically, step 101 includes: real-time monitoring of each edge device to obtain the current resource status data of each edge device, including current computing power, current network bandwidth, current remaining power and current memory capacity.
[0041] In this embodiment, the current resource status data refers to the resource availability and performance indicators of the edge device at the current moment. It should be noted that the present invention does not impose specific limitations on the current resource status data. For example, the current resource status data includes current computing power, current network bandwidth, current remaining power, and current memory capacity. Optionally, the current resource status data may also include other indicators, such as current CPU utilization, current disk I / O utilization, current network latency, current energy consumption, and current task queue length.
[0042] The current computing power of an edge device refers to the computing resources it can provide at any given moment. In practice, the current computing power directly affects the speed at which an edge device processes image segments; the stronger the current computing power, the faster the processing speed.
[0043] The current network bandwidth of an edge device refers to the network transmission rate that the edge device can provide at the current moment. In practice, the current network bandwidth directly affects the speed at which data is transmitted between devices; the higher the current network bandwidth, the shorter the data transmission time.
[0044] The current remaining battery power of an edge device refers to the remaining battery power of the edge device at the current moment. In practice, the current remaining battery power directly affects the device's battery life. The lower the current remaining battery power, the more energy-efficient the device may need to operate in, and it may even need to stop certain tasks to save power.
[0045] The current memory capacity of the edge device refers to the amount of memory available to the edge device at the current moment. In practice, the current memory capacity directly affects the amount of data the device can process. The smaller the memory capacity, the smaller the shard size or the more lightweight model the device may need.
[0046] Understandably, by monitoring each edge device in real time and obtaining its current resource status data, task allocation and model optimization strategies can be dynamically adjusted to ensure that each edge device operates efficiently within its capabilities, thereby improving the stability and reliability of target detection.
[0047] Further, step 102 includes: performing structured pruning on the target detection model deployed on each edge device based on the current resource status data of each edge device and the global task requirements, to obtain the pruned target detection model.
[0048] Structured pruning is a technique used to reduce the number of model parameters and computational cost. Unlike unstructured pruning, structured pruning removes not only individual weights, but entire neurons, channels, or filters, thus maintaining the structural integrity of the model. This makes the pruned model more suitable for efficient operation on hardware.
[0049] Specifically, the main purpose of structured pruning is to reduce the number of model parameters and computational cost while maintaining model accuracy, thereby improving inference speed and reducing energy consumption. Types of structured pruning include, but are not limited to: channel pruning, neuron pruning, and layer pruning.
[0050] In one example, the importance of each channel, neuron, or layer in the object detection network is evaluated based on the current resource status data of each edge device and the global task requirements. Based on the importance evaluation results, channels, neurons, or layers that need to be pruned are selected. Further, the selected channels, neurons, or layers are removed, and the model structure is updated. Optionally, the model's weights and structure are further updated to ensure that the pruned model can still perform forward and backward propagation. Optionally, the pruned model is further fine-tuned to recover the accuracy loss caused by pruning.
[0051] Optionally, in one possible implementation, step 102 above includes: Based on the current resource status data of each edge device and the global task requirements, determine the pruning rate corresponding to the target detection model deployed on each edge device; For each target detection model deployed on an edge device, a structured pruning process is performed on the target detection model according to the pruning rate corresponding to the target detection model to obtain the pruned target detection model.
[0052] The pruning rate is a crucial concept in structured pruning, representing the proportion of model parameters (such as weights, channels, and neurons) removed during the pruning process. The pruning rate directly impacts the size and performance of the pruned model. In practice, a higher pruning rate results in more parameters being removed, a smaller model size, and lower computational cost.
[0053] For example, global task requirements include accuracy requirements, inference speed requirements, and energy budget. In practice, higher accuracy requirements result in lower pruning rates. Higher inference speed requirements result in higher pruning rates. Higher energy budgets result in lower pruning rates.
[0054] Based on the above explanation, current resource status data refers to the resource availability and performance indicators of edge devices at the current moment. Current resource status data includes current computing power, current network bandwidth, current remaining battery power, and current memory capacity.
[0055] Specifically, in one example, the above-mentioned determination of the pruning rate corresponding to the target detection model deployed on each edge device based on the current resource status data of each edge device and the global task requirements includes: The current resource status data and global task requirements of each edge device are input into the resource requirement optimization model to obtain the pruning rate corresponding to the target detection model deployed on each edge device, as output by the resource requirement optimization model.
[0056] The resource requirement optimization model dynamically adjusts the pruning rate of the target detection model based on the current resource status of each edge device and the global task requirements. The goal of this model is to optimize resource utilization while meeting task requirements (such as accuracy, latency, and energy consumption), ensuring that each edge device operates efficiently within its resource constraints.
[0057] For example, the optimization objective of a resource requirement optimization model is to minimize inference time, energy consumption, or model size while maintaining model accuracy within a predetermined range. The constraints of a resource requirement optimization model include: the pruned model accuracy meets the task's accuracy requirements; the inference time is within the task's latency budget; the model's energy consumption is within the task's energy consumption budget; and the model's memory usage is within the device's available memory range.
[0058] Understandably, the choice of pruning rate needs to strike a balance between model size and performance. A higher pruning rate can significantly reduce the computational load of the model, but may sacrifice some accuracy. In this implementation, considering the current resource status data of each edge device and the global task requirements, the pruning rate corresponding to the target detection model deployed on each edge device is determined. For each target detection model deployed on the edge device, structured pruning is performed on the target detection model according to the corresponding pruning rate, resulting in a pruned target detection model. This pruned model can significantly improve the inference speed of the model and reduce energy consumption while maintaining high accuracy.
[0059] In practical applications, it is necessary to pre-build a resource demand optimization model. Optionally, in one possible implementation, the above method further includes: Construct datasets for pruning rate and accuracy, and pruning rate and energy consumption. Input the pruning rate and accuracy dataset into the accuracy predictor to obtain the relationship between the pruning rate and accuracy output by the accuracy predictor. Input the pruning rate and energy consumption dataset into the energy consumption predictor to obtain the relationship between the pruning rate and energy consumption output by the energy consumption predictor. Determine the relationship between pruning rate and model memory usage; The relationships between pruning rate and accuracy, pruning rate and energy consumption, and pruning rate and model memory usage are used as constraints. A resource requirement optimization model is constructed by using a linear programming solver to solve the pruning rate optimization problem.
[0060] Specifically, in edge computing scenarios, the deployment of deep learning models must simultaneously meet the dual constraints of accuracy assurance and energy efficiency. For example, the battery capacity limitations of edge devices mean that excessive inference power consumption may lead to performance degradation or even abnormal shutdowns, and the coupling relationship between accuracy and energy consumption during model compression has not been fully quantified. To address this issue, this invention proposes to simultaneously construct pruning rate and accuracy datasets, as well as pruning rate and energy consumption datasets, and systematically reveal the joint impact of model compression on accuracy and energy efficiency through experimental methods.
[0061] In one example, two related experimental groups were used to construct pruning rate and accuracy datasets, and pruning rate and energy consumption datasets. Specifically, the first experimental group performed structured pruning on the target detection model, generating 99 derived models with pruning rates ranging from 1% to 99% at 1% intervals, forming a model set covering the entire pruning range. Model accuracy data at each pruning rate was collected to form a pruning rate and accuracy dataset. Subsequently, accuracy verification was performed on a standard detection dataset to establish the mapping relationship between pruning rate and model accuracy. The second experimental group deployed the above pruned model to an edge intelligent terminal, collected real-time power data during the inference process through professional power consumption monitoring equipment, and obtained accurate energy consumption values through time integration calculation. Energy consumption data at each pruning rate was collected to form a pruning rate and energy consumption dataset. The experimental design considered the heterogeneity of devices, and by reproducing the experiment on various typical edge devices, a multi-dimensional energy consumption dataset with device difference characterization capabilities was constructed.
[0062] Furthermore, the pruning rate and accuracy datasets are input into the accuracy predictor to obtain the relationship between the pruning rate and accuracy output by the accuracy predictor.
[0063] Specifically, to clarify the relationship between model accuracy and model pruning rate, this invention constructs an accuracy predictor to explore the model accuracy under different pruning rates. To adapt to heterogeneous edge devices, an accuracy predictor needs to be established. Taking the YOLOv8 model as the object detection model as an example, real data constructed using pruning rate and accuracy datasets are used. The mAP50 of the object detection model is used to represent the accuracy under different pruning rates, and Gompertz curves are used for data fitting. The fitting function used is as follows: in, a , b , c It is a coefficient. This indicates the pruning rate.
[0064] Figure 2 This is a schematic diagram illustrating the relationship between accuracy and pruning rate provided by the present invention. Figure 2 (a) is a schematic diagram showing the relationship between the accuracy and pruning rate of the yolov8m model. Figure 2 (b) is a schematic diagram showing the relationship between the accuracy and pruning rate of the YOLOv5S model. According to... Figure 2 It can be seen that the larger the model, the more room there is for pruning. The yolov8m model has more parameters than the yolov5s model, so the yolov8m model is better able to fit the shape of the s-curve.
[0065] Furthermore, the pruning rate and energy consumption dataset are input into the energy consumption predictor to obtain the relationship between the pruning rate and energy consumption output by the energy consumption predictor.
[0066] Specifically, to clarify the relationship between equipment energy consumption and model pruning rate, this invention constructs an energy consumption predictor to explore the energy consumption of the equipment during operation under different pruning rates. Energy is the product of equipment power and time. This invention employs a novel method to calculate the time required for the equipment to perform computational tasks, namely: Among them, the calculated strength Defined as the number of clock cycles the GPU processes for every 1KB of data input. For equipment i The calculation frequency, For equipment i computational efficiency This represents the standby power consumption offset determined through residual analysis. Computational efficiency is defined as the number of floating-point operations performed per watt of power. The rated power is then obtained based on the number of floating-point operations performed after model pruning.
[0067] Figure 3 This is a schematic diagram illustrating the relationship between energy consumption and pruning rate provided by the present invention. For example... Figure 3 As shown, the curve fitted using the yolov8m model demonstrates the relationship between the model's pruning rate and energy consumption. It can be seen that energy consumption and... The relationship is quadratic, and the two curves follow the same trend, which can accurately predict energy consumption under different pruning rates.
[0068] Furthermore, the relationship between the pruning rate and the model's memory usage was determined.
[0069] For example, the model's memory usage can be divided into four parts: input layer memory, feature map memory, parameter memory, and output layer memory. (Example: The model's memory usage...) The expression indicating the memory usage of the model is as follows: in, This refers to the memory occupied by the input layer, i.e., the memory used by the input image. This refers to the memory used for model parameters, specifically the memory occupied by model weights. This refers to the memory occupied by the feature maps of intermediate layers, such as convolutional layers. This refers to the memory usage of the output layer, specifically the memory occupied by the detection boxes output by the model. Let i be the available memory capacity of device i.
[0070] Specifically, a memory prediction model based on batch normalization (BN) layer sensitivity is established. Because of the structured BN layer channel pruning strategy, the pruning rate and the number of channels retained in the BN layer have a strictly linear relationship, which makes the model memory consumption exhibit a significantly linear characteristic. in, The pruning rate, This represents the slope of the linear relationship between memory consumption and pruning rate. This is the intercept of the linear relationship between memory consumption and pruning rate, representing the baseline memory usage of the unpruned model.
[0071] Furthermore, the relationships between pruning rate and accuracy, pruning rate and energy consumption, and pruning rate and model memory usage are used as constraints. A resource demand optimization model is constructed by using a linear programming solver to solve the pruning rate optimization problem.
[0072] Specifically, linear programming (LP) is used to solve for the optimal solution of a resource demand optimization model with linear constraints. In this case, LP is used to find the optimal pruning rate that minimizes inference time while satisfying resource constraints. The fitted relation is then substituted and transformed into the following form: in, For pruning rate The computational quantity function under, For equipment i computational efficiency For equipment i The calculated intensity, For pruning rate The size of the model below For equipment i The calculation frequency, This is the standby power consumption offset. This is the preset energy consumption budget. a, b, c The fitting coefficients for the Gompertz curve are given. This is the preset accuracy budget. This represents the slope of the linear relationship between memory consumption and pruning rate. This is the intercept of the linear relationship between memory consumption and pruning rate, representing the baseline memory usage of the unpruned model. This is the preset memory budget.
[0073] In this invention, to solve the aforementioned LP problem, a standard LP solver based on the simplex method is used. The LP solver will find the optimal value of the pruning ratio. By using this pruning ratio, a resource requirement optimization model that meets resource constraints and the required accuracy can be generated.
[0074] Further, step 103 includes: uniformly segmenting the high-definition image captured by the drone, and retaining the overlapping area of a preset size to obtain multiple image segments.
[0075] In practical applications, edge intelligence-driven load resource collaborative optimization devices acquire high-definition images captured by drones and divide these images into multiple image slices of the same size, so that tasks can be assigned to multiple edge devices for parallel processing.
[0076] Specifically, due to the special nature of convolution operations, each convolution needs to acquire features within the receptive field. After image segmentation, the edge parts may lack a small number of features, so redundant segmentation is required. Therefore, the overlapping area of the preset size is retained.
[0077] Understandably, the edge-intelligent-driven load resource collaborative optimization device uniformly segments high-resolution images captured by drones, retaining overlapping regions of a preset size to obtain multiple image segments. This serves several purposes: firstly, it allows the task to be distributed to multiple edge devices for parallel processing, significantly reducing the computational burden on individual devices compared to a single device processing the entire image, thereby significantly improving computational efficiency and accelerating target detection. Secondly, retaining overlapping regions of a preset size supports subsequent result fusion processing, avoiding the loss of target information due to segmentation, thus improving the accuracy of target detection. Furthermore, segmentation processing allows for better handling of complex scenarios, improving the robustness of target detection.
[0078] Optionally, in one possible implementation, step 103 above includes: The minimum size of an image slice is determined based on the image slice size that the edge device with the smallest current memory capacity can process. Calculate the number of image pieces based on the size of the high-definition image and the minimum size of the image piece; The step size of the image segment is calculated based on the minimum size of the image segment and the size of the overlapping area. Based on the number of image slices and the step size of the slices, the high-definition image is sliced to obtain multiple image slices.
[0079] In practical applications, edge devices have limited memory capacity, especially in resource-constrained environments. If the image tile size is too large, it may lead to insufficient memory on the edge device, preventing it from processing the allocated image tiles. Determining the minimum image tile size based on the image tile size that the edge device with the smallest current memory capacity can process ensures that all edge devices can process image tiles within their memory limits, avoiding performance degradation of the entire system due to resource limitations of a single edge device. Furthermore, since the resource status of edge devices is dynamic, by monitoring these resource statuses in real time and adjusting the image tile size according to the image tile size that the edge device with the smallest current memory capacity can process, the system can dynamically adapt to environmental changes, ensuring stability and reliability.
[0080] Specifically, the edge intelligence-driven load resource collaborative optimization device monitors the current memory capacity of each edge device in real time, identifies the edge device with the smallest memory capacity, determines the size of image segments that this edge device can process, and uses this size as the minimum size of the image segments. This ensures that the memory occupied by each image segment does not exceed the available memory of the edge device with the smallest memory capacity.
[0081] For example, for a CNN-based object detection model, the memory occupied by the object detection task can be calculated as follows. The model's memory usage is generally divided into four parts: input layer memory, feature map memory, parameter memory, and output layer memory. Input layer memory is the size of the input image, which is the minimum acceptable slice size. In practice, due to different network structures, there will be multiple different intermediate layer feature maps. The feature map of each layer can be obtained by multiplying the feature map's length, width, number of channels, batch size, and single-precision floating-point number. Depending on the model structure, the sum of the feature maps of all layers yields the feature map memory. Parameter memory depends on the number of parameters and their storage precision. Taking 32-bit float single-precision floating-point numbers as an example, parameter memory is the product of the number of parameters and 4 bytes. Output layer memory is the memory occupied by the detection boxes output by the model. After obtaining the above information, the remaining available memory of the edge device can be determined, i.e., the current memory capacity of the edge device. Further, based on the image slice size that the edge device with the smallest current memory capacity can process, the minimum size of the image slice is determined.
[0082] It is worth noting that, according to the methods of parallel computing on hardware, the hardware's computational performance is best utilized when the width and height of the input image are multiples of 32. In one example, the minimum size of the image slices is set to a multiple of 32.
[0083] For example, the minimum size of an image patch is High-resolution image size Image slicing retains overlapping areas of a preset size, while the size of non-overlapping areas is... The sizes of the overlapping regions are respectively , as well as Specifically, based on the size of the non-overlapping region. and the size of high-definition images Calculate the number of image slices obtained.
[0084] Furthermore, based on the minimum size of the image patch and the size of the overlapping region, the width and height of the non-overlapping region are determined. The width of the non-overlapping region is used as the patch step size in the horizontal direction, and the height of the non-overlapping region is used as the patch step size in the vertical direction.
[0085] Furthermore, based on the number of image slices and the step size of the slices, the high-definition image is sliced to obtain multiple image slices.
[0086] Furthermore, step 104 above includes: allocating multiple image slices to multiple edge devices based on the current resource status data of each edge device, so that each edge device applies a locally deployed target detection model to perform target inference on the allocated image slices and outputs the partition target inference result.
[0087] Edge devices refer to computing devices deployed at the network edge that can directly process and analyze data from data sources without sending the data to the cloud or a central server. Examples of edge devices include, but are not limited to: drones, mobile devices (smartphones, tablets, etc.), and edge servers.
[0088] Specifically, in one example, step 104 above, which allocates multiple image slices to multiple edge devices based on the current resource status data of each edge device, includes: Calculate the computation latency and transmission latency of each edge device based on its current computing power and network bandwidth. The maximum total delay is calculated based on the computation and transmission delay of each edge device. The total delay is the sum of the transmission delays of all edge devices and the maximum value of the computation delays of all edge devices. The image fragment allocation problem is modeled as an integer linear programming problem. The objective is to minimize the maximum total latency. The constraints are that the sum of the sizes of all image fragments after removing overlapping regions is equal to the size of the high-resolution image, and the number of fragments for each edge device is a non-negative integer. The maximum total latency is the maximum value of the total latency of all edge devices. Solve the integer linear programming problem to obtain the current partitioning strategy; According to the current fragmentation allocation strategy, multiple image fragments are allocated to multiple edge devices.
[0089] Specifically, the image partitioning problem is modeled as an integer linear programming problem. For example, the Python `pulp` library is used to solve it and obtain the optimal partitioning solution. The constraints for image partitioning are as follows: the sum of all partitions after removing redundant parts equals the size of the original image. in, It represents the number of image slices allocated to each edge device, and is also a set of solutions that an integer linear programming problem needs to find. This refers to the redundancy during image fragmentation, where A is the workload size. Latency is constrained by pre-setting a maximum acceptable total latency D. The total latency is mainly divided into two parts: computation latency and transmission latency. Computation latency is the total time for all devices to calculate the detection task, and transmission latency is the transmission latency of image fragments under different bandwidths. The maximum total latency satisfies the following condition: in, For equipment i The computational delay, For equipment i Let B be the transmission delay, and B be the transmission bandwidth between devices. Furthermore, since bandwidth fluctuates and is difficult to predict in a short time, it is assumed that the transmission bandwidth is stable during this transmission period. Therefore, the following constraints apply: Specifically, for input images of different sizes, the computation time can be estimated based on the device's computing power. For example, the thop library can be used to calculate the computational cost for input images of different sizes. Convolution operations scale proportionally with the input image size; the higher the resolution of the input image, the more computational cost increases exponentially. After obtaining the computational cost, the fastest time for a device to compute image slices can be calculated based on its computing power (i.e., maximum floating-point operations per second). Where F represents the computational quantity function, Representative equipment i The available computational power for inference tasks. Given the constraints mentioned above, we treat the problem as an integer linear programming problem and calculate a set of optimal solutions. This allows devices with strong computing power and good bandwidth to undertake more tasks, ensuring balanced utilization and efficient operation of system resources.
[0090] In this embodiment, each edge device is equipped with an object detection model. The object detection model is a trained model. For example, the object detection model can be built and trained based on YOLO (You Only Look Once) series models, SSD (Single Shot MultiBox Detector), Faster R-CNN, etc.
[0091] Understandably, after each edge device obtains its assigned image slice, it applies a locally deployed object detection model to perform object inference on the assigned image slice and outputs the partition object inference result.
[0092] Furthermore, it is necessary to integrate the partition target inference results returned by each device to obtain the global target inference result.
[0093] Step 105 includes: obtaining the partition target inference results output by each edge device, integrating the partition target inference results output by each edge device, eliminating redundant detection boxes corresponding to overlapping areas, and obtaining global target inference results.
[0094] Understandably, by integrating the partitioned target inference results from various devices and eliminating redundant detection boxes in overlapping areas, duplicate detection can be avoided, resulting in globally consistent target inference results and ensuring the accuracy and reliability of the final result.
[0095] Optionally, in one example, step 105 above includes: Obtain the partition target inference results output by each edge device; Based on the coordinate information and redundant area design during the slicing process, the partition target inference results output by each edge device are mapped back to the original image coordinate system corresponding to the high-definition image; The non-maximum suppression algorithm is used to filter redundant detection boxes corresponding to overlapping regions, retaining the detection boxes with the highest confidence, and obtaining the global target inference result.
[0096] Specifically, firstly, based on the coordinate information and redundant region design during image segmentation, the target inference results output by each edge device are mapped back to the original image coordinate system corresponding to the high-resolution image. Since overlapping areas exist at the edges of image segments, the same target may be repeatedly detected in multiple segments. Therefore, a non-maximum suppression (NMS) algorithm is needed to filter redundant detection boxes. Specifically, the detection boxes are first sorted according to their confidence level, retaining the detection box with the highest confidence level while removing other detection boxes whose intersection-union ratio (IU) exceeds a set threshold. This process is iterated until all detection boxes have been processed. Through NMS processing, duplicate boxes can be eliminated while retaining high-confidence detection results, ultimately yielding accurate and non-redundant target inference results from the complete image. This effectively solves the problem of result redundancy caused by segmentation calculations, ensuring the accuracy and reliability of the final detection results.
[0097] In addition, to verify the edge intelligent driving load resource collaborative optimization method provided in this application, experiments and analyses were conducted to verify the effectiveness of the edge intelligent driving load resource collaborative optimization method. The following is an explanation with specific examples.
[0098] Specifically, the VisDrone 2019 dataset was used for the experiments. This dataset is primarily designed for target detection and multi-target tracking tasks in drone aerial photography scenarios. The dataset was collected from diverse environments including urban, suburban, and rural areas, covering 10 typical target categories (such as pedestrians, cars, and bicycles), exhibiting high scene complexity and target scale variability. The dataset images have a resolution of 1920×1080, with the training set containing over 5000 labeled images and the test set containing 1500 images, making it suitable for research and validation of computer vision tasks such as target detection and behavior analysis.
[0099] Regarding the experimental equipment, a high-performance computing server equipped with an NVIDIA GeForce RTX 3090 Ti GPU (24GB VRAM) was used for model training and large-scale inference testing; an edge computing device, NVIDIA Jetson AGXXavier, was used to simulate the edge deployment environment and verify the lightweight nature and computational efficiency of the model.
[0100] Regarding model training and optimization strategies, the experiments were conducted based on the YOLOv8 object detection framework, using three different sized models: YOLOv8s, YOLOv8m, and YOLOv8l, for comparative studies. The standard 50-epoch training strategy was employed, and during the pruning and optimization phase, the impact of different training epochs (30, 50, and 80 epochs) on model sparsity was further explored to balance model accuracy and computational efficiency.
[0101] Specifically, four key metrics are used to comprehensively evaluate the target detection system. For detection accuracy, the mean average accuracy (mAP) is used as the core metric, evaluating model performance by calculating the average accuracy of each category when the intersection-union ratio (IU) threshold is 0.5. This metric comprehensively reflects the model's ability to recognize different categories of targets. For inference speed evaluation, frames per second (FPS) is used as the quantification standard to measure the system's ability to process images per unit of time. This metric directly determines the system's real-time performance in practical applications, and is particularly important for scenarios requiring rapid response, such as drones. The end-to-end latency metric comprehensively considers two key factors: communication latency and inference latency. Communication latency reflects data transmission efficiency, while inference latency reflects model computational efficiency; both together determine the overall response speed of the system. This metric is significant for evaluating the system's practicality. For energy consumption, joules are used as the unit of measurement, evaluating system energy efficiency by measuring power consumption during inference. This metric has important guiding value for the selection and optimization of edge computing devices, especially on resource-constrained drone platforms, where energy efficiency often determines system feasibility.
[0102] Furthermore, comparative experiments were conducted. Specifically, this application achieved better overall performance on the VisDrone dataset, with an mAP50 of 39.8%, significantly better than the Sahi method's 38.5%, but slightly lower than the aSahi method's 41.6%. Notably, this application has a significant advantage in inference speed, achieving an FPS of 19.56, which is 5.3 times and 3.9 times faster than the Sahi and aSahi methods, respectively, demonstrating better real-time processing capabilities. Although the aSahi method performs better in large target detection (mAP50_l), this method maintains a good balance in the detection of medium-sized targets (mAP50_m) and small targets (mAP50_s).
[0103] Furthermore, energy consumption predictors and accuracy predictors were constructed. Specifically, pruning and sparsification training was performed: the sparsification effect of the YOLOv8m model was tested at different training epochs (Epoch = 30, 50, 80). Experimental results show that the model's performance gradually improves with the increase of sparsification training epochs, thereby expanding the space for pruning. Figure 4 This is a schematic diagram illustrating the impact of sparsification training on the model pruning accuracy provided by the present invention. Figure 4 (a) is a schematic diagram illustrating the impact of sparsity training on model pruning accuracy over 30 epochs. Figure 4 (b) is a schematic diagram illustrating the impact of sparsity training on model pruning accuracy at 50 epochs. Figure 4 (c) is a schematic diagram illustrating the impact of sparsity training on model pruning accuracy at 80 epochs. Figure 4As shown, at 30 epochs, the model has almost no room for pruning; at 50 epochs, the pruning effect is still unsatisfactory; however, at 80 epochs, the accuracy remains basically unchanged while maintaining a pruning rate of 38%, and the model's memory usage decreases from 98.81MB to 52.00MB, significantly reducing storage overhead. Compared to the original model, memory usage is reduced by approximately 47.4%, and accuracy does not show a significant decrease.
[0104] Specifically, regarding the model generalization of the accuracy predictor: experiments were conducted using different YOLOv8 (YOLOv8s, YOLOv8m, YOLOv8l) models, all of which were able to fit a good S-shaped curve. Figure 5 This is a schematic diagram illustrating the generalization performance of the accuracy predictor provided by the present invention. Figure 5 (a) is a schematic diagram illustrating the generalization performance of the YOLOv8s model accuracy predictor. Figure 5 (b) is a schematic diagram illustrating the generalization performance of the YOLOv8m model accuracy predictor. Figure 5 (c) is a schematic diagram illustrating the generalization performance of the YOLOv8l model accuracy predictor. Experiments show that as the number of model parameters increases, the S-curve shifts to the left, indicating that the more complex the model, the larger the pruning space. While maintaining essentially the same accuracy, the memory usage of the YOLOv8s model decreased from 42.66MB to 24.54MB (a reduction of approximately 42.5%); the YOLOv8m model decreased from 98.81MB to 52.00MB (a reduction of 47.4%); and the YOLOv8l model decreased from 166.61MB to 66.37MB (a reduction of approximately 60.2%).
[0105] Furthermore, experiments were conducted on different hardware platforms to verify the generalization performance of the energy consumption predictor. Figure 6 This is a schematic diagram illustrating the generalization performance of the energy consumption predictor provided by the present invention. Figure 6 (a) is a schematic diagram illustrating the generalization performance of the power consumption predictor on the 3090 Ti graphics card. Figure 6 (b) A schematic diagram illustrating the generalization performance of the energy consumption predictor on the Jetson AGX Orin development board. First, models with different pruning rates (0-99%) were tested on a 3090 Ti graphics card. For each pruning rate, 100 images were inferred, and the average inference energy consumption per image was calculated. Subsequently, the predicted energy consumption curve was obtained using the energy consumption prediction formula, as shown below. Figure 6 As shown in (a). Furthermore, the same method was verified on a Jetson AGX Orin (64GB) development board. In this experiment, the development board was set to 30W power mode, and the experimental results are as follows. Figure 6As shown in (b). It should be noted that the predicted energy consumption curve is not a perfectly smooth quadratic curve. This is because the computational cost of the model does not increase strictly under different pruning rates, but rather shows an overall increasing trend. This phenomenon indicates that the energy consumption prediction model can capture energy consumption changes under different pruning rates relatively accurately, but under certain pruning rates, the growth rate of computational cost may not be linear.
[0106] Further ablation experiments were conducted. Inference time and energy consumption under different fragment sizes: Experimental results show that when using the YOLOv8m model, the 640×640 single-fragment scheme performs best in detection accuracy (mAP50=0.417), but has higher inference latency (62.6ms) and energy consumption (1.50J). As the fragment size decreases and the number increases, the accuracy decreases slightly (mAP50=0.358 for the 320×320 scheme), but the computational efficiency improves significantly. The inference energy consumption of the 6-fragment scheme is reduced by 32%, while the latency only increases by 30%. This indicates that in scenarios where accuracy requirements are not stringent, the strategy of using small fragments and multiple concurrent operations can effectively improve computational efficiency.
[0107] Inference time and power consumption under different latencies: Experiments show that with the 320×320 fragmentation scheme, the increase in bandwidth significantly improves performance: from 1Mbps to 100Mbps, inference latency decreases by 53% (389ms→183ms), and power consumption decreases by 53% (9.54J→4.46J). However, even at 100Mbps, the performance is still more than twice as bad as that of a single device, indicating that network communication is the main bottleneck.
[0108] Computational latency and energy consumption under different devices: In a 640×640 single-sharding scheme, both the AGX development board and the 3090Ti server achieve a detection accuracy of 0.417 mAP50. However, the AGX development board demonstrates superior edge computing performance with an inference latency of only 62.6ms and a low energy consumption of 1.50J. Although the 3090Ti server still has advantages in latency and energy consumption, the AGX development board, while maintaining the same detection accuracy, fully meets the real-time requirements of edge devices. This fully demonstrates the applicability of this method in edge computing scenarios: it can guarantee detection accuracy while achieving efficient inference on resource-constrained edge devices, providing a feasible solution for real-time target detection on mobile platforms such as UAVs.
[0109] The edge intelligence-driven load resource collaborative optimization method provided in this embodiment obtains the current resource status data of each edge device in real time, and can dynamically adjust task allocation and model optimization strategies to ensure that each edge device operates efficiently within its capabilities, thereby improving the stability and reliability of target detection. Furthermore, based on the current resource status data and global task requirements, structured pruning of the target detection model can reduce the number of model parameters and computational load, ensuring that the model achieves optimal performance and resource utilization efficiency under different devices and task scenarios. The pruned model maintains high accuracy while significantly reducing computational load, thereby improving inference speed and meeting real-time requirements. Further, by uniformly segmenting high-definition images captured by UAVs and retaining overlapping regions of a preset size, multiple image slices are obtained. This allows tasks to be distributed to multiple edge devices for parallel processing, significantly reducing the computational burden on a single device compared to processing the entire image on a single device, thus significantly improving computational efficiency and accelerating target detection. On the other hand, retaining overlapping regions of a preset size supports subsequent result fusion processing, avoiding the loss of target information due to segmentation, thereby improving the accuracy of target detection. On the other hand, image fragmentation allows for better handling of complex scenarios and improves the robustness of object detection. Furthermore, dynamically allocating image fragments based on the current resource status data of the devices enables load balancing, preventing some devices from being overloaded while others remain idle, thus improving the overall resource utilization of the system. Moreover, by integrating the partitioned object inference results from each device and eliminating redundant detection boxes in overlapping areas, duplicate detection can be avoided, resulting in globally consistent object inference results and ensuring the accuracy and reliability of the final outcome. Therefore, the solution in this embodiment, by optimizing the image processing flow and adaptively pruning the object detection model, rationally utilizes the resources of multiple heterogeneous devices, not only improving the real-time performance of object detection but also enhancing the overall energy efficiency of the system and the utilization rate of edge devices.
[0110] The edge intelligent drive load resource collaborative optimization device provided by the present invention will be described below. The edge intelligent drive load resource collaborative optimization device described below and the edge intelligent drive load resource collaborative optimization method described above can be referred to in correspondence.
[0111] Figure 7 This is a schematic diagram of the edge-intelligent driven load resource collaborative optimization device provided by the present invention, as shown below. Figure 7 As shown, the device includes: Monitoring module 71 is used to monitor each edge device in real time and obtain the current resource status data of each edge device. The current resource status data includes the current computing power, current network bandwidth, current remaining power and current memory capacity. The pruning module 72 is used to perform structured pruning on the target detection model deployed on each edge device based on the current resource status data and global task requirements of each edge device, so as to obtain the pruned target detection model. The slicing module 73 is used to uniformly slice the high-definition images captured by the drone and retain the overlapping area of a preset size to obtain multiple image slices. The allocation module 74 is used to allocate multiple image slices to multiple edge devices according to the current resource status data of each edge device, so that each edge device can apply the locally deployed target detection model to perform target inference on the allocated image slices and output the partition target inference results. The integration module 75 is used to obtain the partition target inference results output by each edge device, integrate the partition target inference results output by each edge device, eliminate redundant detection boxes corresponding to overlapping areas, and obtain the global target inference results.
[0112] Optionally, in one possible implementation, the pruning module 72 includes: The determination unit is used to determine the pruning rate corresponding to the target detection model deployed on each edge device based on the current resource status data of each edge device and the global task requirements; The pruning unit is used to perform structured pruning on the target detection model deployed on each edge device, based on the pruning rate corresponding to the target detection model, to obtain the pruned target detection model.
[0113] Optionally, in one possible implementation, the determining unit described above is specifically used for: The current resource status data and global task requirements of each edge device are input into the resource requirement optimization model to obtain the pruning rate corresponding to the target detection model deployed on each edge device, as output by the resource requirement optimization model.
[0114] Optionally, in one possible implementation, the above-described edge-intelligent driven load resource collaborative optimization device further includes: a construction module; the construction module is used for: Construct datasets for pruning rate and accuracy, and pruning rate and energy consumption. Input the pruning rate and accuracy dataset into the accuracy predictor to obtain the relationship between the pruning rate and accuracy output by the accuracy predictor. Input the pruning rate and energy consumption dataset into the energy consumption predictor to obtain the relationship between the pruning rate and energy consumption output by the energy consumption predictor. Determine the relationship between pruning rate and model memory usage; The relationships between pruning rate and accuracy, pruning rate and energy consumption, and pruning rate and model memory usage are used as constraints. A resource requirement optimization model is constructed by using a linear programming solver to solve the pruning rate optimization problem.
[0115] Optionally, in one possible implementation, the above-mentioned segmentation module 73 is specifically used for: The minimum size of an image slice is determined based on the image slice size that the edge device with the smallest current memory capacity can process. Calculate the number of image pieces based on the size of the high-definition image and the minimum size of the image piece; The step size of the image segment is calculated based on the minimum size of the image segment and the size of the overlapping area. Based on the number of image slices and the step size of the slices, the high-definition image is sliced to obtain multiple image slices.
[0116] Optionally, in one possible implementation, when the allocation module 74 allocates multiple image slices to multiple edge devices based on the current resource status data of each edge device, it is specifically used for: Calculate the computation latency and transmission latency of each edge device based on its current computing power and network bandwidth. The maximum total delay is calculated based on the computation and transmission delay of each edge device. The total delay is the sum of the transmission delays of all edge devices and the maximum value of the computation delays of all edge devices. The image fragment allocation problem is modeled as an integer linear programming problem. The objective is to minimize the maximum total latency. The constraints are that the sum of the sizes of all image fragments after removing overlapping regions is equal to the size of the high-resolution image, and the number of fragments for each edge device is a non-negative integer. The maximum total latency is the maximum value of the total latency of all edge devices. Solve the integer linear programming problem to obtain the current partitioning strategy; According to the current fragmentation allocation strategy, multiple image fragments are allocated to multiple edge devices.
[0117] Optionally, in one possible implementation, the above-mentioned integration module 75 is specifically used for: Obtain the partition target inference results output by each edge device; Based on the coordinate information and redundant area design during the slicing process, the partition target inference results output by each edge device are mapped back to the original image coordinate system corresponding to the high-definition image; The non-maximum suppression algorithm is used to filter redundant detection boxes corresponding to overlapping regions, retaining the detection boxes with the highest confidence, and obtaining the global target inference result.
[0118] The edge-driven load resource collaborative optimization device provided in this embodiment has a monitoring module that monitors each edge device in real time to obtain the current resource status data of each edge device. This allows for dynamic adjustment of task allocation and model optimization strategies, ensuring that each edge device operates efficiently within its capabilities and improving the stability and reliability of target detection. Furthermore, the pruning module performs structured pruning on the target detection model based on the current resource status data and global task requirements. This reduces the number of model parameters and computational load, ensuring optimal performance and resource utilization efficiency across different devices and task scenarios. The pruned model maintains high accuracy while significantly reducing computational load, thereby improving inference speed and meeting real-time requirements. Further, the slicing module uniformly slices the high-definition images captured by the UAV, retaining overlapping areas of a preset size to obtain multiple image slices. This allows tasks to be distributed to multiple edge devices for parallel processing, significantly reducing the computational burden on individual devices compared to processing the entire image on a single device, thus significantly improving computational efficiency and accelerating target detection. On the other hand, retaining overlapping areas of a preset size to support subsequent result fusion processing avoids the loss of target information due to fragmentation, thereby improving the accuracy of target detection. Furthermore, fragmentation processing allows for better handling of complex scenarios, improving the robustness of target detection. Moreover, the allocation module dynamically allocates image fragments based on the current resource status data of the devices, achieving load balancing and preventing some devices from being overloaded while others are idle, thus improving the overall resource utilization of the system. Furthermore, the integration module integrates the partitioned target inference results from each device and eliminates redundant detection boxes in overlapping areas, avoiding duplicate detection and obtaining globally consistent target inference results, ensuring the accuracy and reliability of the final result. Therefore, the solution in this embodiment, by optimizing the image processing flow and adaptively pruning the target detection model, rationally utilizes the resources of multiple heterogeneous devices, not only improving the real-time performance of target detection but also enhancing the overall energy efficiency of the system and the utilization rate of edge devices.
[0119] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 8As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communications bus 840. The processor 810 can call logical instructions in the memory 830 to execute an edge intelligence-driven load resource collaborative optimization method. This method includes: real-time monitoring of each edge device to obtain current resource status data for each edge device, including current computing power, current network bandwidth, current remaining power, and current memory capacity; performing structured pruning on the target detection model deployed on each edge device based on the current resource status data and global task requirements to obtain a pruned target detection model; uniformly segmenting high-definition images captured by the drone, retaining overlapping areas of a preset size, to obtain multiple image segments; allocating the multiple image segments to multiple edge devices based on the current resource status data of each edge device, so that each edge device applies its locally deployed target detection model to perform target inference on the allocated image segments and outputs partitioned target inference results; acquiring the partitioned target inference results output by each edge device, integrating the partitioned target inference results output by each edge device, eliminating redundant detection boxes corresponding to overlapping areas, and obtaining a global target inference result.
[0120] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0121] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the edge intelligence-driven load resource collaborative optimization method provided by the above methods. The method includes: real-time monitoring of each edge device to obtain the current resource status data of each edge device, including current computing power, current network bandwidth, current remaining power, and current memory capacity; performing structured pruning on the target detection model deployed on each edge device according to the current resource status data of each edge device and global task requirements to obtain a pruned target detection model; uniformly segmenting the high-definition images captured by the UAV and retaining overlapping areas of a preset size to obtain multiple image segments; allocating the multiple image segments to multiple edge devices according to the current resource status data of each edge device, so that each edge device applies the locally deployed target detection model to perform target inference on the allocated image segments and outputs partitioned target inference results; obtaining the partitioned target inference results output by each edge device and integrating the partitioned target inference results output by each edge device to eliminate redundant detection boxes corresponding to overlapping areas to obtain a global target inference result.
[0122] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the edge-driven load resource collaborative optimization method provided by the above methods. The method includes: real-time monitoring of each edge device to obtain current resource status data of each edge device, including current computing power, current network bandwidth, current remaining power, and current memory capacity; performing structured pruning on the target detection model deployed on each edge device according to the current resource status data of each edge device and global task requirements to obtain a pruned target detection model; uniformly segmenting high-definition images captured by a drone and retaining overlapping areas of a preset size to obtain multiple image segments; allocating the multiple image segments to multiple edge devices according to the current resource status data of each edge device, so that each edge device applies its locally deployed target detection model to perform target inference on the allocated image segments and outputs partitioned target inference results; obtaining the partitioned target inference results output by each edge device and integrating the partitioned target inference results output by each edge device to eliminate redundant detection boxes corresponding to overlapping areas to obtain a global target inference result.
[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A load resource collaborative optimization method driven by edge intelligence, characterized in that, include: Real-time monitoring of each edge device to obtain the current resource status data of each edge device, including current computing power, current network bandwidth, current remaining power, and current memory capacity; Based on the current resource status data and global task requirements of each edge device, the target detection model deployed on each edge device is subjected to structured pruning to obtain the pruned target detection model. The high-definition images captured by the drone are uniformly segmented, and overlapping areas of a preset size are retained to obtain multiple image segments; Based on the current resource status data of each edge device, the multiple image slices are allocated to multiple edge devices, so that each edge device applies a locally deployed target detection model to perform target inference on the allocated image slices and outputs the partition target inference results; The partition target inference results output by each edge device are obtained, and the partition target inference results output by each edge device are integrated to eliminate redundant detection boxes corresponding to the overlapping areas, thereby obtaining the global target inference result.
2. The edge-driven load resource collaborative optimization method according to claim 1, characterized in that, The step of performing structured pruning on the target detection model deployed on each edge device based on the current resource status data and global task requirements of each edge device to obtain the pruned target detection model includes: Based on the current resource status data of each edge device and the global task requirements, determine the pruning rate corresponding to the target detection model deployed on each edge device; For each target detection model deployed on an edge device, a structured pruning process is performed on the target detection model according to the pruning rate corresponding to the target detection model to obtain the pruned target detection model.
3. The edge-driven load resource collaborative optimization method according to claim 2, characterized in that, The step of determining the pruning rate corresponding to the target detection model deployed on each edge device based on the current resource status data and global task requirements of each edge device specifically includes: The current resource status data of each edge device and the global task requirements are input into the resource requirement optimization model to obtain the pruning rate corresponding to the target detection model deployed on each edge device, as output by the resource requirement optimization model.
4. The edge-driven load resource collaborative optimization method according to claim 3, characterized in that, The method further includes: Construct datasets for pruning rate and accuracy, and pruning rate and energy consumption. The pruning rate and accuracy dataset are input into the accuracy predictor to obtain the relationship between the pruning rate and accuracy output by the accuracy predictor. The pruning rate and energy consumption dataset are input into the energy consumption predictor to obtain the relationship between the pruning rate and energy consumption output by the energy consumption predictor. Determine the relationship between pruning rate and model memory usage; The relationships between the pruning rate and accuracy, the pruning rate and energy consumption, and the pruning rate and model memory usage are used as constraints. A resource requirement optimization model is constructed by solving the pruning rate optimization problem using a linear programming solver.
5. The edge-driven load resource collaborative optimization method according to claim 1, characterized in that, The step of allocating the multiple image slices to multiple edge devices based on the current resource status data of each edge device includes: Calculate the computation latency and transmission latency of each edge device based on its current computing power and network bandwidth. The maximum total latency is calculated based on the computation latency and transmission latency of each edge device. The total latency is the sum of the transmission latency of all edge devices and the maximum value of the computation latency of all edge devices. The image fragment allocation problem is modeled as an integer linear programming problem, with the goal of minimizing the maximum total latency. The constraints are that the sum of the sizes of all image fragments after removing overlapping regions is equal to the size of the high-resolution image, and the number of fragments for each edge device is a non-negative integer. Solve the integer linear programming problem to obtain the current partitioning strategy; According to the current fragmentation allocation strategy, the multiple image fragments are allocated to multiple edge devices.
6. The edge-driven load resource collaborative optimization method according to claim 1, characterized in that, The step of integrating the partitioned target inference results output by each edge device, eliminating redundant detection boxes corresponding to the overlapping regions, and obtaining global target inference results includes: Based on the coordinate information and redundant region design during the slicing process, the partition target inference results output by each edge device are mapped back to the original image coordinate system corresponding to the high-definition image; The non-maximum suppression algorithm is used to filter the redundant detection boxes corresponding to the overlapping regions, and the detection boxes with the highest confidence are retained to obtain the global target inference result.
7. The edge-driven load resource collaborative optimization method according to any one of claims 1-6, characterized in that, The process of uniformly segmenting the high-definition images captured by the drone, while retaining overlapping areas of a preset size, to obtain multiple image segments includes: The minimum size of an image slice is determined based on the image slice size that the edge device with the smallest current memory capacity can process. The number of image segments is calculated based on the size of the high-definition image and the minimum size of the image segment. The step size of the image segment is calculated based on the minimum size of the image segment and the size of the overlapping area. The high-definition image is segmented according to the number of image segments and the segmentation step size to obtain multiple image segments.
8. An edge-intelligent driven load resource collaborative optimization device, characterized in that, include: The monitoring module is used to monitor each edge device in real time and obtain the current resource status data of each edge device. The current resource status data includes the current computing power, current network bandwidth, current remaining power, and current memory capacity. The pruning module is used to perform structured pruning on the target detection model deployed on each edge device based on the current resource status data and global task requirements of each edge device, so as to obtain the pruned target detection model. The slicing module is used to uniformly slice the high-definition images captured by the drone, and retain the overlapping area of the preset size to obtain multiple image slices; The allocation module is used to allocate the multiple image slices to multiple edge devices according to the current resource status data of each edge device, so that each edge device applies the locally deployed target detection model to perform target inference on the allocated image slices and outputs the partition target inference results; An integration module is used to acquire the partition target inference results output by each edge device, integrate the partition target inference results output by each edge device, eliminate redundant detection boxes corresponding to the overlapping areas, and obtain global target inference results.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the load resource collaborative optimization method of edge intelligent driving as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the load resource collaborative optimization method of edge intelligence drive as described in any one of claims 1 to 7.