Low-delay edge end large model reasoning acceleration method and system
By constructing a computing bottleneck identification system and a hierarchical acceleration optimization strategy, the problem of high inference latency for large models in UAV edge computing environments was solved, achieving efficient, low-latency inference execution and flight safety.
Patent Information
- Application Number
- CN202511349615.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-28
AI Technical Summary
High latency in large model inference, uneven distribution of computing resources, and insufficient coordination between flight control and inference tasks in UAV edge computing environments make it impossible to meet real-time flight control requirements.
By constructing a multi-level computing bottleneck identification system and a hierarchical acceleration optimization strategy, including hierarchical sparsity extraction, operator fusion analysis, parallelism optimization, and resource scheduling matrix, the allocation of computing resources is dynamically adjusted to optimize inference tasks.
It achieves efficient and low-latency inference execution on the UAV platform, ensuring flight safety and mission completion quality, reducing overall inference latency and improving computational efficiency.
Smart Images

Figure CN120849128A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence inference technology, and in particular to a low-latency edge-end large model inference acceleration method and system. Background Technology
[0002] In the edge computing environment of drones, large-model inference faces multiple challenges, including stringent flight safety requirements, tight real-time constraints, and high priority of flight control tasks. Existing drone inference systems generally adopt a strategy of directly deploying models in the cloud, failing to fully consider the special characteristics of drone flight control tasks and the resource constraints of onboard hardware. Traditional methods typically rely on simple model quantization and pruning techniques for model compression, lacking precise localization of computational bottlenecks during drone inference and coordinated optimization between flight control and AI inference tasks, resulting in inference latency that fails to meet the real-time flight control requirements of drones.
[0003] The main problems with current technologies include: first, the lack of methods for identifying computational bottlenecks specific to the characteristics of UAV onboard equipment, making it impossible to accurately pinpoint key aspects affecting flight control and inference tasks; second, the lack of consideration for flight control task latency constraints in model optimization strategies, failing to comprehensively plan the allocation of computing resources between flight control and AI inference; and third, the lack of adaptive resource scheduling to UAV flight status, making it difficult to dynamically adjust according to flight mission requirements and hardware conditions. These technical limitations severely restrict the practical deployment and widespread application of large-scale models on UAV platforms. Summary of the Invention
[0004] This invention discloses a low-latency edge-end large model inference acceleration method and system, aiming to solve the technical problems of high latency in large model inference at the edge of UAVs, uneven distribution of computing resources, and insufficient coordination between flight control tasks and inference tasks in the prior art. By constructing a multi-level computing bottleneck identification system and a hierarchical acceleration optimization strategy, it achieves efficient inference and low-latency execution of large models on the UAV platform, ensuring flight safety and mission completion quality.
[0005] The first aspect of this invention proposes a low-latency edge-end large model inference acceleration method, comprising the following steps: Acquire UAV computing power resource signals and flight control mission parameters, perform hierarchical sparsity extraction on the computing power resource signals to identify computing bottleneck features, derive time delay constraint coefficients from the flight control mission parameters, and dynamically associate the computing bottleneck features with the time delay constraint coefficients to establish an inference task mapping relationship. Based on the inference task mapping relationship, the model hotspot tracking is used to locate high-frequency computing cores. Operator fusion analysis is performed on the high-frequency computing cores to form acceleration paths. Delay distribution data is collected along the acceleration paths, and the inference pipeline network is mapped based on the delay data. Parallelism analysis is performed on the inference pipeline network to identify branch merging points, the throughput parameters of the branch merging points are extracted, the optimal splitting position is determined by the matching degree between the throughput parameters and the computing bottleneck features, and the optimal splitting position is connected to the high-frequency computing core to determine the optimization path; The cache parameter chain is reconstructed based on the optimization path, the cache parameter chain is aligned with the latency constraint coefficient to determine the low latency processing window, and an adaptive quantization matrix is formed based on the low latency processing window. The adaptive quantization matrix is coupled with the delay distribution data to identify the precision loss region. Redundant computation is extracted from the precision loss region. A resource scheduling matrix is established based on the redundant computation and the inference pipeline network. A hierarchical acceleration strategy is formed based on the resource scheduling matrix, and a latency curve is formed by tracking the execution jitter of the hierarchical acceleration strategy. Inference acceleration execution instructions are generated based on the latency curve.
[0006] A second aspect of this invention proposes a low-latency edge-end large-model inference acceleration system, comprising: The resource analysis module is used to acquire UAV computing power resource signals and flight control mission parameters, perform hierarchical sparsity extraction on the computing power resource signals to identify computing bottleneck features, derive time delay constraint coefficients from the flight control mission parameters, and dynamically associate the computing bottleneck features with the time delay constraint coefficients to establish an inference task mapping relationship. The hotspot localization module is used to perform model hotspot tracking and localize high-frequency computing cores based on the inference task mapping relationship, perform operator fusion analysis on the high-frequency computing cores to form acceleration paths, collect delay distribution data along the acceleration paths, and map the inference pipeline network based on the delay distribution data. The parallel analysis module is used to perform parallelism analysis on the inference pipeline network to identify branch merging points, extract the throughput parameters of the branch merging points, determine the optimal splitting position by the matching degree between the throughput parameters and the computing bottleneck features, and connect the optimal splitting position to the high-frequency computing core to determine the optimization path. The parameter configuration module is used to reconstruct the cache parameter chain according to the optimization path, align the cache parameter chain with the latency constraint coefficient to determine the low latency processing window, and form an adaptive quantization matrix based on the low latency processing window. The resource optimization module is used to couple the adaptive quantization matrix with the delay distribution data to identify the accuracy loss region, extract redundant computation in the accuracy loss region, and establish a resource scheduling matrix based on the redundant computation and the inference pipeline network. The execution control module is used to form a hierarchical acceleration strategy based on the resource scheduling matrix, track the execution jitter of the hierarchical acceleration strategy to form a latency curve, and generate inference acceleration execution instructions based on the latency curve.
[0007] The beneficial effects of this invention are reflected in the following points: First, the hierarchical sparsity extraction technology can accurately locate the computational bottlenecks in the UAV inference process. Through operator memory usage graph construction and gradient sensitivity analysis, it identifies the key operators that have the greatest impact on inference latency. Combined with the inference task mapping relationship established by the flight control task latency constraint coefficient, the UAV's computing resources can be prioritized for tasks critical to flight safety, effectively reducing the overall inference latency. Second, the operator fusion analysis and parallelism optimization technology, by identifying high-frequency computing cores and forming acceleration paths, combined with optimal splitting position determination and branch merging point identification, transforms the originally serially executed computing tasks into a parallel execution mode, while reducing data transmission overhead and memory access frequency between operators, significantly improving the UAV's inference throughput and computational efficiency. Third, the hierarchical acceleration strategy, combined with adaptive quantization matrix and resource scheduling matrix, through accuracy loss zone identification and redundant computation extraction, minimizes unnecessary computational operations while ensuring inference accuracy. Combined with the inference acceleration execution commands generated by jitter tracking, it achieves dynamic optimization control of the UAV inference process, enabling the UAV to complete real-time inference tasks of complex large models with limited onboard computing resources, ensuring flight safety and mission execution efficiency.
[0008] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating a low-latency edge-end large model inference acceleration method of the present invention.
[0010] Figure 2 This is a block diagram of a low-latency edge-end large-model inference acceleration system according to the present invention. Detailed Implementation
[0011] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0012] It should be noted that all directional indicators (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.
[0013] It should also be noted that when a component is described as "fixed to" or "set on" another component, it can be directly on the other component or there may be an intervening component present. When a component is described as "connected to" another component, it can be directly connected to the other component or there may be an intervening component present.
[0014] In addition, the descriptions of "first", "second", etc. in this application are for descriptive purposes only and should not be understood as indicating or implying their relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0015] like Figure 1 As shown, this embodiment of the invention provides a low-latency edge-end large model inference acceleration method, including the following steps S110-S160: Step S110: Obtain UAV computing power resource signals and flight control mission parameters; perform hierarchical sparsity extraction on the computing power resource signals to identify computing bottleneck features; derive time delay constraint coefficients from the flight control mission parameters; and dynamically associate the computing bottleneck features with the time delay constraint coefficients to establish an inference mission mapping relationship.
[0016] Specifically, the system acquires UAV computing resource signals and flight control mission parameters. Through the hardware monitoring interface of the UAV's onboard computing platform, it collects multi-dimensional computing resource signals in real time, including CPU utilization, GPU memory usage, memory bandwidth utilization, and power consumption distribution. These signals reflect the real-time load status of the UAV's computing system. When the UAV simultaneously performs multiple AI tasks such as target detection, path planning, and obstacle avoidance decision-making, the CPU may reach over 90% utilization due to complex algorithms, and the GPU memory may approach its capacity limit due to large-scale image processing. These resource signals can promptly reflect the system's computational pressure. The computing resource signal acquisition system employs a high-frequency sampling mechanism to ensure the capture of transient changes in resource usage. It synchronously monitors the execution status of the UAV's onboard inference model, including the resource usage of AI inference tasks such as target detection, path planning, and obstacle avoidance decision-making. It obtains process-level resource allocation information through system call interfaces, including memory allocation, computation unit usage time, and I / O access frequency for each inference task. A correlation mechanism is established between system-level resource monitoring and model-level execution analysis to track the actual resource consumption of each operator in the inference model. Simultaneously collect flight control mission parameters, including flight path planning data, attitude control commands, sensor data processing requirements, and a real-time decision-making task list. Flight control mission parameters are acquired via the flight control communication bus and transmitted using standard communication protocols. A mission priority classification system is established to manage flight control missions hierarchically according to safety levels.
[0017] In some embodiments, the step of performing hierarchical sparsity extraction to identify computational bottleneck features on the computing resource signal includes: constructing an operator memory occupancy graph based on the computing resource signal; performing gradient sensitivity analysis on the operator memory occupancy graph to obtain pruning candidate points; injecting sparsity perturbations into the pruning candidate points to identify and activate sparse paths; and extracting structured pruning parameters along the activated sparse paths and encoding them as computational bottleneck features.
[0018] A memory usage graph for operators is constructed based on computing resource signals. Combining system-level resource monitoring data and model execution analysis results, a detailed analysis and visualization of the memory usage of each operator in the UAV inference model is provided. Computational bottleneck characteristics describe the key resource constraints limiting overall performance in the system. When the UAV target detection model processes 4K resolution images, convolutional layers may become memory bottlenecks due to their large number of parameters, and fully connected layers may become CPU bottlenecks due to their high computational complexity. These bottleneck characteristics determine the execution efficiency of the entire inference task. Memory tracking is performed on the model execution process of inference tasks such as target detection and path planning through deep performance analysis, recording the memory usage of each operator at different execution stages. Operators are classified according to their hierarchical structure in the neural network, including convolutional layers, fully connected layers, activation layers, pooling layers, and batch normalization layers. For typical UAV inference tasks, the resource consumption patterns of computationally intensive and memory-intensive operators are analyzed in detail. The average memory usage, peak memory usage, and duration of each operator type are statistically analyzed to establish a statistical model of operator memory usage. The memory usage data is organized into a graph structure according to the execution order and hierarchical relationship of the operators. The nodes of the graph represent operators, and the edges represent data flow and memory dependencies. A multi-dimensional representation of the memory usage graph is established, including a memory usage curve in the time dimension, an operator distribution graph in the spatial dimension, and a memory consumption structure in the hierarchical dimension.
[0019] Gradient sensitivity analysis is performed on the operator memory usage graph to obtain pruning candidate points. Based on the constructed operator memory usage graph, the sensitivity of each operator to the overall performance of the UAV inference task is analyzed, and candidate positions that can be safely pruned are identified. Sensitivity analysis is used to evaluate the impact of each operator parameter on inference accuracy, and pruning priority is determined through parameter importance assessment. The sensitivity of the computational model output to each operator parameter is analyzed through forward and backward propagation, and a sensitivity evaluation system is constructed. The sensitivity evaluation comprehensively considers the impact of parameters on task accuracy and system resource consumption. The sensitivity evaluation results are jointly analyzed with memory usage to calculate the pruning benefit ratio of each operator, which is equal to the ratio of resource savings to performance loss. A sensitivity threshold suitable for the real-time inference requirements of the UAV is set, and operators with sensitivity below the threshold and high memory usage are marked as pruning candidate points. Cluster analysis is performed on the candidate points to identify operator groups with similar characteristics, and each group can be subject to a unified pruning strategy.
[0020] Injecting sparsity perturbations into pruning candidate points identifies and activates sparse paths. The identified pruning candidate points are then sparsified, and the impact of operator pruning on UAV inference tasks is tested by simulating different resource-constrained scenarios. Multiple sparsity perturbation modes are designed, including parameter zeroing, structured pruning, channel pruning, and weight quantization. For typical UAV flight scenarios, the impact of different sparsity strategies on target detection accuracy, path planning accuracy, and obstacle avoidance response speed is tested. By gradually increasing the perturbation intensity, the critical point at which model performance begins to decline is identified, determining the sparsity tolerance of each operator. A progressive strategy is adopted for sparsity perturbation testing to ensure the safety and reliability of UAV inference tasks. During perturbation, the data flow propagation path is monitored, and computational paths less affected by the perturbation are identified. Paths with less impact are marked as activated sparse paths; these paths retain good information transmission capabilities after sparsification. Optimal sparsity combinations are determined through path analysis to maximize computational resource savings while ensuring UAV flight safety.
[0021] Structured pruning parameters are extracted and encoded as computational bottleneck features along the activated sparse path. Based on the identified activated sparse path, structured pruning parameters are extracted and encoded as feature vectors describing computational bottlenecks. The activated sparse path refers to the key computational channels in the network that have a significant impact on the results. For example, in target tracking algorithms, convolutional layers specifically responsible for identifying the contours of moving targets are more important than layers responsible for background texture analysis; the former constitutes the core part of the activated sparse path. The distribution pattern of retained parameters is analyzed along the sparse path to identify the network structural features most important for key UAV inference tasks. Connectivity indices of the sparse path are calculated, including path length, number of branches, convergence point location, and importance of key nodes. Path connectivity focuses on computational paths critical to flight safety and mission execution; for example, computational channels responsible for obstacle detection must remain fully connected. The proportion of parameter retention, computational reduction, and memory saving for each layer after pruning is extracted to form a multi-dimensional description of the pruning effect. An encoding system for pruning parameters is established, organizing information such as network layers, pruning type, sparsity, and performance impact into structured feature vectors. Computational bottleneck features reflect the key factors limiting the overall performance of the network. By encoding the pruning results into these feature vectors, subsequent network acceleration decisions can be guided.
[0022] Derivation of latency constraint coefficients from flight control mission parameters. Based on the collected flight control mission parameters, the sensitivity and tolerance of various tasks to execution latency are analyzed, quantifying latency constraint requirements. The latency constraint coefficient quantifies the sensitivity of flight control tasks to response time. When the UAV performs emergency obstacle avoidance tasks, the latency constraint coefficient for obstacle detection is very high, requiring identification and decision-making to be completed within milliseconds. The latency constraint coefficient for environmental temperature monitoring is lower, tolerating processing delays in the second range. A tiered standard for latency constraints is established based on UAV flight safety requirements and the real-time characteristics of flight control tasks. Critical flight control tasks have the most stringent latency requirements, followed by navigation and positioning tasks, while environmental perception and auxiliary function tasks have relatively lenient latency requirements. A latency sensitivity index is calculated for each specific inference task, comprehensively considering flight status, environmental complexity, and task urgency. The latency constraint coefficient is calculated using the formula C = Tr / Td, where C is the latency constraint coefficient, Tr is the required response time of the task, and Td is the task deadline. A smaller constraint coefficient indicates a more latency-sensitive task requiring a higher scheduling priority. By analyzing the dependencies of flight control tasks, critical task chains affecting flight safety are identified, and the total delay of each chain and the delay contribution ratio of each task are calculated. A delay budget allocation mechanism is established, allocating a reasonable delay budget to each task based on its importance and computational complexity. The delay constraint coefficients of all tasks are arranged according to their task identifiers to form a delay constraint coefficient vector.
[0023] A dynamic mapping relationship between computational bottleneck features and latency constraints is established. This mapping relationship is defined as the optimal matching strategy between AI inference tasks and computational resources. When a UAV needs to simultaneously process visual target detection and LiDAR point cloud analysis, the system allocates image processing to GPU resources and point cloud analysis to CPU multi-core resources based on the computational requirements and latency constraints of the two tasks, achieving efficient resource utilization. The matching degree between each inference task and each computational bottleneck feature is calculated, based on the task's computational requirement pattern and the resource supply characteristics of the bottleneck feature. The matching strength is calculated through similarity analysis between the task requirement vector and the bottleneck feature vector, prioritizing tasks with high similarity to their corresponding computational resources. The matching results are weighted and adjusted based on the safety level and latency constraint coefficient of the flight control task to ensure that critical flight tasks receive priority resource guarantees. A mapping relationship matrix M = W × (1 - αC) is established, where M is the mapping relationship matrix, W is the basic matching weight between tasks and resources, C is the latency constraint coefficient, and α is the adjustment parameter. The mapping relationship is adjusted through optimization algorithms to maximize computational resource utilization efficiency while meeting flight safety requirements and latency constraints.
[0024] Step S120: Based on the inference task mapping relationship, perform model hotspot tracking to locate high-frequency computing cores, perform operator fusion analysis on high-frequency computing cores to form acceleration paths, collect delay distribution data along the acceleration paths, and map the inference pipeline network based on the delay distribution data.
[0025] In some embodiments, the step of performing model hotspot tracking to locate high-frequency computing kernels based on the inference task mapping relationship includes: extracting an operator dependency graph from the inference task mapping relationship; applying delay accumulation analysis to the operator dependency graph to obtain a critical path; performing operator reordering on the critical path to identify a fusion candidate set; and determining a high-frequency computing kernel based on the fusion candidate set.
[0026] Extract operator dependency graphs from the inference task mapping relationships. Analyze the task allocation information within the inference task mapping relationships to identify the operator sequences and execution order involved in each inference task. High-frequency computational kernels refer to core operator combinations that are frequently invoked and consume significant computational resources in inference tasks. For example, the edge detection convolution operator in infrared images needs to process high-resolution thermal imaging data, and the A* algorithm matrix operation in path planning needs to calculate safe flight paths in complex terrain in real time. These are typical high-frequency computational kernels. Using data flow analysis techniques, trace the transmission paths of tensors between different operators in UAV inference tasks, and the input-output dependencies between operators. Construct a graph structure to represent operator dependencies, where nodes represent computational operators, and edges represent data dependencies and control dependencies. Assign weights to each dependency edge, reflecting the magnitude of data transmission and the impact of transmission latency. Calculate the dependency strength by comprehensively considering data transmission size, frequency, and latency. Check the rationality of dependencies to ensure there are no circular dependencies or logical errors.
[0027] Latency accumulation analysis is applied to the operator dependency graph to identify critical paths. Path analysis is performed on the constructed operator dependency graph to identify the critical execution paths that have the greatest impact on the overall execution latency. Path analysis methods are used to calculate all possible paths from the start node to the end node of the graph, and the cumulative execution latency of each path is calculated. For each operator node, its earliest start time and latest finish time are calculated to identify critical operators with zero time margin. The formula for calculating path latency accumulation is L_path=Σ(Ti+Di), where L_path is the total path latency, Ti is the execution time of the i-th operator, and Di is the data transmission latency of the i-th operator. Through critical path analysis, the longest execution path that determines the overall execution time is identified, and any increase in latency on this path will directly affect the total execution time. The sensitivity of each path to changes in UAV inference task performance is analyzed. Segmented analysis is performed on the critical paths to identify bottleneck operators and performance hotspots within the paths.
[0028] Operator reordering is performed on the critical path to identify fusion candidate sets. For the identified critical path, operator reordering techniques are used to optimize the execution order and identify operator combinations with fusion potential. Data dependencies between bottleneck operators and adjacent operators within performance hotspot regions are analyzed to identify bottleneck operator pairs whose execution order can be adjusted. Dependency relaxation techniques are used to increase the flexibility of operator scheduling while ensuring correctness. A search algorithm is employed to explore different operator ordering schemes, evaluating the impact of each scheme on performance hotspot regions and the optimization effect on bottleneck operators. The benefits of reordering are evaluated by comparing execution times before and after reordering. Operator combinations with similar execution patterns and compatible data types are identified, with particular attention paid to combinations containing bottleneck operators; these combinations are marked as fusion candidate sets. Memory access pattern analysis prioritizes operator combinations that reduce memory round trips. The goal is to ensure that candidate operators maintain computational accuracy and functional correctness after fusion.
[0029] For example, determining the high-frequency computing core based on the fusion candidate set includes: evaluating the fusion benefits based on the fusion candidate set to determine the optimization priority, wherein the fusion benefits include memory access reduction, computational density improvement, and parallelization potential; formulating an operator fusion strategy according to the optimization priority; and generating the high-frequency computing core through the operator fusion strategy.
[0030] The optimization priority is determined by evaluating the fusion benefits based on the fusion candidate set. A multi-dimensional benefit analysis is performed on each fusion candidate set to quantify the performance improvements and resource savings brought by the fusion operation. Quantitative calculations are performed on the reduction in memory accesses, statistically analyzing changes in memory read / write operations, cache hit rate, and bandwidth utilization before and after fusion. After fusing the convolutional layer and batch normalization layer in the infrared target detection network, normalization calculations can be directly performed within the GPU registers, avoiding data movement between GPU memory and system memory and reducing memory access overhead. Detailed memory behavior during operator execution is analyzed using memory access tracing tools to identify redundant memory operations that can be eliminated through fusion. The degree of computational density improvement is evaluated by calculating the percentage increase in the number of effective computational operations completed per unit time after fusion. The impact of operator fusion on computational resource utilization is analyzed, including ALU utilization, vector unit utilization, and dedicated computational unit utilization. The parallelization potential is quantified, evaluating the parallel execution capability of the fused operators on multi-core processors, GPUs, and dedicated accelerators. A benefit weight model is established to determine the importance weights of different benefit indicators based on the performance requirements of the target application scenario.
[0031] Operator fusion strategies are formulated based on optimization priorities. Detailed implementation strategies and execution plans for operator fusion are developed based on benefit assessment results. When UAVs perform maritime rescue missions, they need to simultaneously handle visible light vessel detection and infrared personnel search. The system prioritizes fusion optimization of vessel identification operator combinations with high detection accuracy requirements, while relegating latency-sensitive personnel search operator combinations to a secondary position. The fusion candidate set is sorted from highest to lowest comprehensive benefit, prioritizing fusion opportunities with the highest benefit. A reasonable fusion execution order is determined considering the mutual influence and dependencies between fusion operations. Complex multi-operator fusion is decomposed into multiple simple two-operator fusion steps, with each fusion involving only two adjacent compatible operators, reducing the complexity and risk of fusion implementation. After each fusion step, the correctness of the results and the performance improvement effect are verified by comparing the computational accuracy and execution speed before and after fusion. When fusion operations cause performance degradation or functional abnormalities, the system can quickly recover to the pre-fusion state, ensuring system stability and reliability. Considering the computing power and memory limitations of the UAV hardware platform, the fusion strategy is adjusted to adapt to the resource constraints of edge computing devices.
[0032] High-frequency computational kernels are generated through an operator fusion strategy. Following the execution order of fusing object detection convolutional layers first, followed by path planning computational layers, related operators are progressively fused and merged to form more powerful composite operators. For UAV vision tasks, convolution-batch normalization-activation functions are fused into a single operator, reducing memory accesses during image processing. For flight control algorithms, continuous matrix operations are fused into an integrated computational kernel, improving attitude control response speed. Operator interface compatibility is maintained during the fusion process, ensuring seamless integration of the fused computational kernel into the existing flight control system without affecting the normal operation of other sensor data processing operators. Execution code optimized for the UAV ARM processor is generated for the fused computational kernel, fully utilizing the parallel computing capabilities of the onboard hardware. Parameters of the fused operators are uniformly managed, simplifying the model loading process during UAV system startup. Flight environment adaptability tests are conducted on the generated high-frequency computational kernels to verify their processing accuracy and execution efficiency under different lighting conditions, wind disturbances, and target complexity scenarios.
[0033] Operator fusion analysis is performed on high-frequency computing kernels to form acceleration paths. Based on the identified high-frequency computing kernels, the data dependencies and fusion potential between adjacent operators are analyzed. The acceleration path is an efficient computational execution sequence optimized through operator fusion. Originally scattered operators such as LiDAR point cloud processing, GPS signal correction, and wind compensation calculation are fused and optimized to form a continuous acceleration path, improving response speed under adverse weather conditions. Dependency analysis is performed on operators in the high-frequency computing kernels to construct data flow graphs and control flow graphs between operators. Operator combinations that can be optimized through fusion are identified, including convolution-batch normalization fusion, activation-pooling fusion, and continuous linear layer fusion. Memory access pattern analysis is used to calculate the reduction in memory access and the improvement in cache hit rate after operator fusion. The impact of operator fusion on computational parallelism is evaluated, and the improvement in computational density and thread utilization after fusion is analyzed.
[0034] Delay distribution data was collected along the acceleration path. Based on the formed acceleration path, the execution delay of each computational stage was measured and statistically analyzed in detail. The delay distribution data describes the time consumption pattern of the computational task at different execution stages. When using UAVs for forest fire monitoring, the fire detection algorithm processes faster in clear weather, but under conditions of dense smoke obscuring the view, the same detection task takes significantly longer due to the need for additional image enhancement and noise filtering. This distribution characteristic of delay variation reflects the specific impact of different environmental conditions on algorithm performance. Performance monitoring probes were deployed at each key node of the acceleration path to record the operator's startup delay, execution delay, and data transmission delay in real time. High-precision timestamp technology was used to ensure the accuracy and consistency of delay measurements. CPU computation delay, GPU core delay, memory access delay, and inter-device communication delay were statistically analyzed separately. The distribution characteristics of the delay data were analyzed using statistical analysis methods to identify key parameters such as the mean, variance, and range of variation of the delay distribution.
[0035] Inference Pipeline Network Based on Latency Distribution Data Mapping. This paper designs an efficient pipeline network architecture suitable for the real-time inference needs of UAVs using collected latency distribution data. The inference pipeline network is a parallel execution architecture optimized based on latency characteristics. By analyzing the latency distribution characteristics of different tasks, the pipeline is designed to allow tasks with different processing speeds to execute in parallel, significantly reducing the total execution time. Based on the latency characteristics of each operator, the computational task is divided into multiple parallel pipeline stages. Through latency balancing analysis, the execution time of each pipeline stage is ensured to be as uniform as possible, avoiding obvious bottleneck stages. The mapping formula for the inference pipeline network is N_stage = ceil(L_max / L_target), where N_stage is the number of pipeline stages, L_max is the longest operator latency, and L_target is the target stage latency. Resource allocation for each pipeline stage is adjusted based on real-time latency feedback to maximize the opportunity for parallel execution while ensuring data dependencies. The optimized operator execution sequence, parallel relationships, and timing constraints are encoded into the topological description of the inference pipeline network, completing the mapping process from latency distribution data to the pipeline network.
[0036] Step S130: Perform parallelism analysis on the inference pipeline network to identify branch merging points, extract the throughput parameters of the branch merging points, determine the optimal splitting position by matching the throughput parameters with the characteristics of the computational bottleneck, and connect the optimal splitting position with the high-frequency computing core to determine the optimization path.
[0037] In some embodiments, performing parallelism analysis to identify branch merging points on the inference pipeline includes: performing tensor parallel scans on the inference pipeline to extract split boundaries; identifying data synchronization overhead points from the split boundaries; performing pipeline depth evaluation based on the synchronization overhead points to generate a parallel efficiency index; and marking the locations where the parallel efficiency index exceeds a threshold as branch merging points.
[0038] Tensor parallel scanning is performed on the inference pipeline network to extract partition boundaries. By analyzing the hierarchical structure and data flow characteristics of the inference pipeline network, the dimensional characteristics of tensors in each layer are analyzed to identify tensor operations that can be partitioned dimensionally. Branch merging points are key locations where multiple parallel branches converge or separate in the inference network. In UAV multi-sensor fusion tasks, three data streams from visible light cameras, infrared cameras, and LiDAR need to be merged at the feature fusion layer. This fusion layer is a typical branch merging point, which determines the efficiency of multi-path parallel processing. Tensor shape analysis is used to determine the feasibility and granularity of partitioning each tensor in different dimensions. The data volume and computational cost distribution of each partitioned tensor are calculated to evaluate the load balancing effect. The parallelism is determined based on the branch structure of the inference pipeline network; the parallelism is the number of network branches that can be processed independently in parallel. Data dependencies between tensor operations are identified, and the constraints that partition boundaries must satisfy are determined. Dependency analysis is used to identify combinations of tensor operations that can be processed independently in parallel. The impact of different partitioning strategies on memory usage patterns is analyzed, and the partitioning scheme with the highest memory efficiency is selected.
[0039] Identify data synchronization overhead points from the partition boundaries. Based on the extracted partition boundaries, analyze the key locations where data synchronization is required during parallel execution. Data synchronization overhead points refer to the locations where different processing branches in parallel computing need to wait and coordinate. For example, when multiple UAVs cooperate in a search, each UAV processes images of different regions, but when performing target fusion, it must wait for all UAVs to complete their respective detection tasks. This fusion waiting point is a typical data synchronization overhead point. Identify data dependencies that cross partition boundaries; these dependencies require synchronization mechanisms to ensure correctness. Analyze the data exchange patterns between different parallel branches, including point-to-point communication, broadcast communication, and reduction communication. Calculate the time overhead of various synchronization operations, including waiting time, transmission time, and coordination overhead. Identify the data exchange points with the highest synchronization frequency; these locations have the greatest impact on overall parallel efficiency. Analyze the relationship between synchronization overhead and parallelism, and find the pattern of synchronization overhead changing with increasing parallelism. Reduce unnecessary synchronization operations and data transmission through communication mode improvements. Identify operations that can reduce synchronization overhead through asynchronous processing and design asynchronous execution strategies.
[0040] Parallel efficiency index is generated based on pipeline depth evaluation using synchronization overhead points. The identified synchronization overhead points are used to evaluate the parallel execution efficiency under different pipeline depth configurations. The parallel efficiency index quantifies the overall effect of pipeline parallel processing. When UAVs perform search and rescue missions, they need to run three algorithms simultaneously: thermal imaging target detection, terrain analysis, and flight path planning. If a single-stage pipeline is used, the three algorithms must be executed sequentially, resulting in a long total time consumption. If a three-stage pipeline is used to run the three algorithms in parallel, although the parallelism is increased, the synchronization overhead is also increased. In this case, the parallel efficiency index can comprehensively evaluate the trade-off between the increased parallelism and the increased synchronization overhead, helping to determine the optimal pipeline configuration. The actual synchronization time consumption of the synchronization overhead points is calculated, converting the time overhead into a synchronization overhead value. The pipeline utilization rate under different depth configurations is calculated to evaluate the load balancing degree of each stage. The parallel efficiency index is calculated using the formula I = P × U / (S + 1), where I is the parallel efficiency index, P is the parallelism, U is the utilization rate, and S is the synchronization overhead.
[0041] Locations with parallel efficiency indices exceeding a threshold are designated as branch merging points. Using the calculated parallel efficiency indices, key locations with high parallel potential in the UAV inference network are identified. An efficiency threshold is set based on the UAV's real-time requirements, prioritizing network layers with minimal impact on flight safety as branch merging points. Candidate locations are sorted by efficiency index, and the location with the highest index that does not affect critical flight control functions is selected as the primary branch merging point. The distribution of the selected branch merging points in the target detection and path planning networks is analyzed to ensure that the obstacle recognition branch and the flight path calculation branch can be executed in parallel reasonably. The computational dependencies between branch merging points are considered to avoid selecting location combinations that would cause flight control data processing delays, preventing multiple merging points from competing for limited onboard computing resources. Candidate locations with similar functions are grouped, and a representative location is selected from each group to ensure a balanced distribution of computational load across UAV tasks. The data processing capabilities of the branch merging points are evaluated to ensure that the selected locations can handle the concurrent demands of real-time image streams, sensor data, and control commands during flight.
[0042] Extract throughput parameters from branch merging points. Based on the identified branch merging points, quantitatively analyze the data processing capacity and transmission efficiency of each merging location. Throughput parameters reflect the data processing capacity of the branch merging point. In a real-time UAV target tracking system, when multiple targets appear simultaneously, the target feature merging layer, acting as a branch merging point, needs to process feature data from multiple detection branches. Its throughput parameter determines the upper limit of the number of targets the system can track simultaneously. Measure the data convergence rate of each branch merging point and statistically analyze the data flow that can be processed per unit time. Use throughput monitoring tools to collect the input and output data rates of the merging points in real time. Analyze the buffering behavior of data at the merging points, calculating buffer occupancy and queuing latency. Measure the synchronization overhead of the merging points, including waiting time between branches and the additional processing time required for data synchronization. Analyze the impact of different data types and sizes on the throughput of the merging points, establishing the relationship between throughput and data features. Calculate the resource utilization efficiency of the merging points, evaluating the use of resources such as CPU, memory, and bandwidth during the data merging process.
[0043] The optimal segmentation position is determined by the matching degree between throughput parameters and computational bottleneck features. The matching degree between throughput parameters and computational bottleneck features is evaluated to quantify the degree of fit between the two. The optimal segmentation position is the best node selection for parallel segmentation in the network. In UAV obstacle avoidance systems, when the LiDAR point cloud processing network needs parallel acceleration, the system searches for the optimal segmentation position between the point cloud segmentation algorithm and the path planning algorithm, placing the computationally intensive point cloud segmentation in parallel processing on the GPU and the logically complex path planning in serial processing on the CPU, achieving optimal allocation of computing resources. The matching strength between each candidate segmentation position and the bottleneck feature is evaluated using a vector similarity calculation method. The impact of the segmentation position on the overall network performance is analyzed, including indicators such as parallel efficiency, resource utilization, and execution latency. The matching degree calculation formula is M=(T·B) / (|T|×|B|), where M is the matching degree, T is the throughput parameter vector, and B is the bottleneck feature vector. The weight allocation of the matching degree evaluation is adjusted considering the implementation complexity of the segmentation operation and the degree of hardware support. An optimal balance is sought between performance improvement and implementation cost. Constraints on the segmentation location are established to ensure sufficient computational independence of the resulting network segments. The impact of different segmentation schemes on memory usage and data transmission is analyzed, and the segmentation strategy with the lowest resource overhead is selected.
[0044] The optimal path is determined by connecting the optimal segmentation point to the high-frequency computing core. The connection relationship between the optimal segmentation point and the high-frequency computing core is analyzed, and an efficient data transmission interface is designed. The optimized path is an efficient data transmission channel connecting the segmentation point and the high-frequency computing core. In a UAV image recognition system, when the image preprocessing module is segmented to the edge AI chip and the feature extraction module remains on the main processor, an optimized path needs to be established between them to transmit the preprocessed image data, ensuring that data transmission latency does not offset the performance improvement brought by parallel computing. The shortest data path connecting the segmentation point and the computing core is found through path planning algorithms. The impact of different connection schemes on the network topology is considered, and a connection strategy that maintains the optimal overall network performance is selected. Standardized design of connection interfaces ensures data compatibility and transmission efficiency between different network segments. Potential bottlenecks on the connection path are analyzed, and performance bottlenecks are eliminated through buffer optimization and pipeline technology. A fault tolerance mechanism is designed so that when a connection path fails, it can quickly switch to a backup path. Load balancing design of connection paths avoids the impact of single path overload on overall performance.
[0045] Step S140: Reconstruct the cache parameter chain according to the optimized path, align the cache parameter chain with the latency constraint coefficient to determine the low latency processing window, and form an adaptive quantization matrix based on the low latency processing window.
[0046] Specifically, the cache parameter chain is reconstructed based on the optimized path. Based on the execution characteristics of the optimized path, the parameter access patterns of each computing node are analyzed to identify the frequency of parameter usage and access timing characteristics. The cache parameter chain is a parameter storage sequence reorganized according to the computation execution order. When a UAV performs real-time obstacle avoidance tasks, traditional parameter storage may randomly distribute the convolution kernel parameters of the obstacle detection algorithm, the weight matrix of the path planning algorithm, and the PID parameters of the control algorithm in memory, leading to frequent cache invalidations. The reconstructed cache parameter chain stores these closely related parameters consecutively according to the execution order, reducing memory access latency. Through parameter dependency analysis, a reference graph and sharing relationship network between parameters are constructed. Frequently accessed parameters are grouped into cache clusters to reduce memory access fragmentation and randomness. The cache hit rate is calculated using the formula H = Na / Nt, where H is the cache hit rate, Na is the number of cache hits, and Nt is the total number of accesses. Based on the execution order of the optimized path, the storage layout of parameters in memory is rearranged to optimize the spatial locality of data. The capacity limitations and access latency characteristics of different cache levels are analyzed, and a multi-level cache parameter allocation strategy is designed. By analyzing parameter lifecycles, we can identify temporary parameters that can be released early and persistent parameters that need to be maintained for a long time.
[0047] In some embodiments, aligning the cache parameter chain with the latency constraint coefficient to determine the low-latency processing window includes: performing memory pooling analysis on the cache parameter chain to identify reuse patterns; performing dynamic batch processing on the reuse patterns and the latency constraint coefficient to obtain batch configurations; performing prefetch optimization analysis based on the batch configurations to determine buffer boundaries; and performing window partitioning along the buffer boundaries to generate low-latency processing windows.
[0048] Memory pooling analysis is performed on the cached parameter chain to identify reuse patterns. Based on the reconstructed cached parameter chain, the memory usage patterns of parameters are analyzed in depth and organized into pools. The access frequency and access patterns of each parameter in different time periods are statistically analyzed to identify parameter combinations with similar access characteristics. In UAV swarm collaborative operations, the path planning algorithms of multiple UAVs repeatedly access the same map data and constraint parameters. Memory pooling analysis reveals the reuse patterns of these shared parameters, allowing parameters such as map grid data, obstacle information, and flight restriction areas to be organized into the same memory pool, reducing redundant loading. Time series analysis reveals periodic patterns and repetitive patterns in parameter access. Parameters with close reuse distances are organized into the same memory pool to improve the spatial locality of the cache. The capacity requirements and access characteristics of different memory pools are analyzed to optimize the size and number of memory pools. Parameter lifecycle overlap analysis identifies parameter pairs that can share memory space. Memory pool allocation and reclamation strategies reduce memory fragmentation and allocation overhead.
[0049] The system dynamically batches data to obtain batch configurations by combining reuse patterns with latency constraints. An efficient batch processing execution strategy is designed based on the identified reuse patterns and latency constraints. The time window and data grouping method for batch processing are determined according to the temporal characteristics of the parameter reuse patterns. In UAV video analysis tasks, when a parameter reuse pattern is detected showing multiple consecutive frames using the same background feature parameters, the system organizes these frames into a batch for processing. Simultaneously, the latency constraint ensures that the batch size does not exceed real-time requirements, avoiding flight safety issues caused by excessive batch processing latency. The limitations of latency constraints on batch size are analyzed to find a balance between latency requirements and batch processing efficiency. The batch size is chosen to be the smaller value between the maximum batch capacity and the time constraint, ensuring both improved processing efficiency and compliance with latency constraints. The batch processing configuration parameters are adjusted based on future computational needs, employing a predictive batch planning strategy. Intra-batch task scheduling strategies optimize the execution order and resource allocation of tasks within the batch. The impact of different batch configurations on memory usage and computational efficiency is analyzed, and the configuration scheme with optimal resource consumption is selected.
[0050] Based on batch configuration, prefetch optimization analysis is performed to determine buffer boundaries. Using the obtained batch configuration information, intelligent data prefetching strategies and buffer management schemes are designed. Data access patterns during batch execution are analyzed to predict data requirements for future batches. In an autonomous UAV inspection mission, when the current batch processes image data of building A, prefetch analysis predicts the data of building B that the next batch needs to process based on the inspection path, preloading relevant detection model parameters and historical inspection records into the buffer to ensure no data waiting delay occurs during batch switching. Access sequence analysis identifies the optimal timing and quantity for data prefetching. Prefetching effectiveness is evaluated to quantify the contribution of prefetching operations to performance improvement. The overhead of prefetching operations is analyzed, including additional memory usage and bandwidth consumption. Prefetch distance optimization determines how far in advance to start prefetching operations for optimal results. A multi-level prefetching strategy employs different prefetch intensities based on data importance and access frequency. The optimal buffer capacity is calculated to find the best balance between memory consumption and caching effectiveness. Accurate determination of buffer boundaries provides precise boundary information for the division of low-latency processing windows.
[0051] Low-latency processing windows are generated by partitioning along the buffer boundaries. Based on the defined buffer boundaries, the entire processing flow is divided into multiple independent low-latency processing units. A low-latency processing window is an independent computational unit that can be completed within latency constraints. During an UAV emergency landing, the system needs to complete ground detection, landing point selection, and trajectory generation within a very short time. Traditional methods may time out due to excessive computational tasks, while low-latency processing windows decompose the entire landing decision-making process into multiple small computational windows, each of which can be completed within milliseconds, ensuring the real-time nature of the landing decision. The start and end boundaries of the processing windows are determined based on the location of the buffer boundaries. Data dependencies between adjacent processing windows are analyzed, and an efficient data transfer mechanism between windows is designed. The computational load and memory requirements of each processing window are calculated to ensure that tasks within the window can be completed within latency constraints. The window latency calculation formula is L = Tc + Tm + Ts, where L is the total window latency, Tc is the computation time, Tm is the memory access time, and Ts is the synchronization time. Processing latency and resource utilization efficiency are balanced by adjusting the window size. Window scheduling strategy determines the parallel execution scheme and priority of multiple processing windows.
[0052] An adaptive quantization matrix is formed based on a low-latency processing window. Utilizing a defined low-latency processing window, the computational latency within the window is further optimized through adaptive adjustment of quantization precision. The adaptive quantization matrix is a configuration table that dynamically adjusts numerical precision according to the latency budget of different processing windows. In UAV visual navigation systems, when lighting conditions are good and image features are clear, lower precision quantization can still achieve accurate recognition; in this case, the quantization matrix reduces numerical precision to improve computational speed. When lighting is dim and image details are blurred, the system requires higher precision computation to ensure recognition accuracy; the quantization matrix automatically increases numerical precision to ensure navigation safety. The computational precision requirements of different operators within the processing window are analyzed to establish a trade-off between precision and latency. Parameters and activation values that are sensitive to and insensitive to quantization operations are identified. A multi-level quantization strategy is designed to allocate different quantization precisions to data of varying importance. The quantization matrix has a row-column structure, with rows corresponding to different data types and columns corresponding to different processing windows. The precision allocation in the quantization matrix is adaptively adjusted according to the latency budget of the processing window.
[0053] Step S150: Couple the adaptive quantization matrix with the delay distribution data to identify the precision loss region, extract redundant computational load in the precision loss region, and establish a resource scheduling matrix based on the redundant computational load and the inference pipeline network.
[0054] In some embodiments, the step of coupling the adaptive quantization matrix with the delay distribution data to identify the accuracy loss region includes: extracting the quantization sensitive layer distribution from the adaptive quantization matrix; performing mixed precision mapping on the quantization sensitive layer distribution and the delay distribution data to obtain weight shared points; applying incremental inference to the weight shared points to generate an accuracy decay response; and determining the accuracy loss region based on the accuracy decay response.
[0055] The distribution of quantization-sensitive layers is extracted from the adaptive quantization matrix. Based on the precision configuration information of the adaptive quantization matrix, the sensitivity distribution of each layer in the network to quantization operations is analyzed. Through quantization precision gradient analysis, the rate of change of precision loss for each network layer during the quantization process is calculated. The performance differences of each layer under different quantization levels are statistically analyzed to identify the most sensitive and least sensitive network layers. The sensitivity calculation formula is Se = ΔA / ΔQ, where Se is the sensitivity, ΔA is the precision change, and ΔQ is the quantization level change. Through analytical methods, network layers with similar quantization sensitivity characteristics are grouped into the same sensitivity group. The distribution pattern of sensitive layers in the network topology is analyzed to identify network regions where sensitive layers are concentrated. The importance of sensitive layers is ranked by weight to determine the key sensitive layers that have the greatest impact on overall performance. Through sensitivity propagation analysis, the impact range of single-layer sensitivity on adjacent layers and the entire network is studied.
[0056] We perform mixed-precision mapping on the quantization-sensitive layer distribution and latency distribution data to obtain weight sharing points. Combining the extracted quantization-sensitive layer distribution and latency distribution data, we identify locations where weight sharing can be achieved. Weight sharing points are key locations where multiple network layers can use the same parameter configuration. In UAV multi-target recognition, when detecting vehicles, pedestrians, and buildings, although the target categories are different, the convolutional kernel parameters for edge detection have strong similarities. By setting weight sharing points in the first few layers of feature extraction, different detection branches can share the same edge detection weights, reducing memory usage while maintaining detection accuracy. We analyze the correspondence between sensitive layers and latency hotspots, examining their overlap in the network space. Through numerical calculation methods, we find the quantization configuration that minimizes accuracy loss while satisfying latency constraints. We design a mixed-precision mapping algorithm to assign different quantization precisions to layers with different sensitivities. We identify combinations of sensitive layers with similar weight distribution characteristics, clustering and sharing the weights of these layers. We analyze the comprehensive impact of weight sharing on computational accuracy and execution latency to ensure that the sharing strategy does not lead to a significant performance degradation.
[0057] Incremental inference is applied to weight sharing points to generate an accuracy decay response. Based on the determined weight sharing points, the impact of weight sharing on model accuracy is analyzed through stepwise testing. In UAV terrain recognition tasks, when multiple recognition branches share convolutional weights, it is necessary to test the impact of the degree of sharing on recognition accuracy. First, two similar terrain branches (such as grassland and farmland) share weights, and the change in recognition accuracy is recorded. Then, the sharing range is gradually increased to three or four branches, and the decrease in accuracy after each increase is observed. The accuracy change curve generated by this stepwise testing is the accuracy decay response. Through small-batch data testing, the accuracy decay under different sharing configurations can be quickly understood. The accuracy decay response is determined based on the functional relationship between the degree of sharing, quantization intensity, and the number of network layers. The nonlinear characteristics of accuracy decay are analyzed to identify the critical point and sharing threshold of sharp accuracy decline. The sharing parameters most sensitive to the accuracy decay response are determined through gradient analysis. Fine-tuning and compensation techniques are used to reduce the accuracy loss caused by weight sharing.
[0058] Accuracy loss regions are determined based on accuracy decay response. Statistical analysis of the accuracy decay response identifies network regions where accuracy loss exceeds an acceptable threshold. Accuracy loss regions are areas in the network where computational accuracy significantly decreases due to quantization and processing operations. In UAV night vision navigation, when quantizing infrared image processing networks, the high-frequency filtering layer responsible for detail enhancement may fail to accurately extract image details due to reduced quantization accuracy, forming accuracy loss regions. Higher-precision computation or compensation algorithms are needed for these regions to maintain navigation accuracy. A boundary detection algorithm for loss regions accurately delineates the spatial range of accuracy loss regions. The distribution density and clustering of loss regions are analyzed to identify hotspots with the most concentrated loss. Accuracy loss regions are classified into three levels—mild, moderate, and severe—based on the severity of loss. The repair difficulty and processing potential of regions with different loss levels are analyzed to formulate differentiated processing strategies. The impact analysis of loss regions calculates the contribution of each region to the overall model performance.
[0059] Redundant computational costs are extracted from the accuracy loss region. Based on the spatial distribution characteristics of the accuracy loss region, the computational complexity of operators within the region is analyzed. Redundant computational costs refer to computational operations that can be omitted or simplified while ensuring functional correctness. In UAV path planning, when processing complex terrain data, traditional algorithms perform detailed safety calculations for each grid point. However, in reality, most obviously safe areas (such as flat open areas) and obviously dangerous areas (such as inside buildings) do not require complex calculations; the detailed calculations in these areas constitute redundant computational costs. The number of multiplication operations, addition operations, and memory accesses for each operator are counted. The distribution of zero values in the parameter matrix and activation tensor is identified through sparsity analysis, and the number of computational operations that can be skipped is calculated. The computational dependencies between operators are analyzed, and intermediate computational steps that can be reduced through operator fusion are identified. The formula for calculating the redundant computational ratio is R=(C_total-C_essential) / C_total, where R is the redundant computational ratio, C_total is the total computational cost, and C_essential is the required computational cost. Activation value analysis identifies computational pathways that contribute little to the final result and marks them as redundant computations. A computational importance system determines the necessity of computation based on the operator's impact on output accuracy.
[0060] In some embodiments, establishing a resource scheduling matrix based on the redundant computational load and the inference pipeline network includes: performing operator-level decomposition on the redundant computational load to obtain a set of prunable operators; performing speculative decoding on the set of prunable operators and the inference pipeline network to identify jump opportunities; performing computation graph processing based on the jump opportunities to generate scheduling paths; and weaving heterogeneous mapping relationships along the scheduling paths to establish a resource scheduling matrix.
[0061] Redundant computational resources are decomposed at the operator level to obtain a set of pruning operators. Based on the extracted redundant computational information, each operator in the network is examined, and the difference between the actual computational operation and the theoretical minimum computational operation for each operator is calculated. The set of pruning operators refers to the combination of operators in the network that can be safely removed or simplified without significantly affecting functionality. For example, convolutional layers that specifically identify the fine texture of leaves can extract rich texture details, but contribute little to macroscopic tasks such as forest cover statistics. Such refined operators can be simplified or removed. Convolutional operators are traversed, and the proportion of zero-value weights in their convolutional kernels is calculated. When the proportion of zero-value weights exceeds a set threshold, the operator is marked as a pruning object. The weight matrices of fully connected layers are examined, and low-rank components are identified through mathematical decomposition. Fully connected layers that can be approximated by low-dimensional matrices are added to the pruning candidate list. Batch normalization and activation function operators are scanned to identify operators with small output value variations, as these operators have limited impact on the network output. The data flow connections between operators are checked to confirm that pruning a certain operator will not cause a dimensionality mismatch in the input data of subsequent operators.
[0062] Speculative decoding is performed on the prunable operator set and the inference pipeline network to identify skip opportunities. Combining the structural features of the prunable operator set and the inference pipeline network, conditional decision nodes are inserted at each processing stage of the pipeline network. A skip opportunity refers to a moment during computation when some computational steps can be terminated or skipped early. For example, when inspecting a transmission line, if the current convolutional layers have clearly identified it as a normal straight line, there is no need to continue executing subsequent complex defect detection operators; the process can directly jump to the result output, saving computation time. At the conditional decision node, the current intermediate result is compared with a preset reference pattern; when the similarity reaches a threshold, the skip mechanism is triggered. For the feature map of the convolutional layer, its correlation with typical feature templates is calculated; when the coefficient indicates that the features have been sufficiently extracted, subsequent refined convolutional operations are skipped. During the computation of the fully connected layer, the magnitude of the output vector change is monitored; when the output change of several consecutive layers is very small, the final output is directly estimated using interpolation methods.
[0063] Scheduling paths are generated by performing computational graph processing based on jump opportunities. The original linear computational graph is reconstructed using identified jump opportunities, with branches inserted and nodes merged. Scheduling paths consist of multiple parallel execution channels designed according to varying computational complexity requirements. For example, when processing aerial images, simple images with uniform lighting and clearly visible targets can follow a fast processing path, while complex images with blurred targets require a complete, refined processing path. The coexistence of these two paths allows the algorithm to flexibly choose based on the actual situation. At each jump opportunity location, the original single execution path is split into multiple parallel paths: a standard path performing the complete computation and a fast path performing simplified computation. A switching node is set between the standard and fast paths, determining the path selection based on intermediate results from real-time computation. For network segments with multiple jump opportunities, the jump decision points are concatenated in the computational order to form a cascaded path selection structure. The data dependencies of each path are recalculated to ensure that path switching does not lead to data loss or computational errors.
[0064] A resource scheduling matrix is established by weaving heterogeneous mapping relationships along the scheduling path. Based on the generated scheduling path, each computing node in the path is scanned to extract its computing type, data scale, and latency requirements. The resource scheduling matrix defines the allocation strategy for computing tasks and hardware resources. In UAV swarm collaborative tasks, when multiple UAVs need to collaboratively process large-scale image stitching, the scheduling matrix will allocate image preprocessing tasks to the edge AI chips of each UAV for parallel processing, feature matching tasks to the high-performance GPU of the master UAV, and the final stitching task to the ground station server, achieving reasonable allocation of computing resources. CPU-intensive tasks are allocated to general-purpose processors, and parallel computing tasks are allocated to GPUs or dedicated accelerators. For memory-intensive tasks, priority is given to processing units with large-capacity high-speed caches. Data transfer requirements between tasks are detected, and tasks that require frequent data exchange are allocated to computing units with close physical proximity. A task-resource mapping table is created, with entries including task identifier, resource type, expected execution time, and resource consumption.
[0065] Step S160: Based on the resource scheduling matrix, a hierarchical acceleration strategy is formed, the execution jitter of the hierarchical acceleration strategy is tracked to form a latency curve, and inference acceleration execution instructions are generated based on the latency curve.
[0066] A hierarchical acceleration strategy is formed based on a resource scheduling matrix. Using the constructed resource scheduling matrix, multi-level acceleration processing strategies and implementation schemes are designed. The hierarchical acceleration strategy is a differentiated acceleration scheme designed for different computing levels. When a UAV needs to process multi-sensor data in real time, GPUs can process image data in parallel at the hardware level, unimportant computational steps can be simplified at the algorithm level, and the amount of data transmitted can be compressed at the data level. Real-time requirements are met through collaborative acceleration at multiple levels. Based on the resource allocation results of the scheduling matrix, the acceleration strategy is divided into three levels: hardware-level acceleration, algorithm-level acceleration, and data-level acceleration. At the hardware level, the parallel utilization efficiency of hardware resources is maximized through the scheduling of dedicated computing units. The speedup ratio is calculated as S = T_baseline / T_optimized, where S is the speedup ratio, T_baseline is the baseline execution time, and T_optimized is the accelerated execution time. At the algorithm level, unnecessary computational operations are reduced through operator fusion, quantization acceleration, and sparse computation techniques. At the data level, overall processing efficiency is improved through memory management improvements, caching strategy adjustments, and data stream processing.
[0067] In some embodiments, the step of tracking the execution jitter of the hierarchical acceleration strategy to form a latency curve includes: extracting inter-layer activation transfer overhead from the hierarchical acceleration strategy; performing a latency sensitivity test on the activation transfer overhead to obtain a bottleneck distribution; identifying latency mutation points based on the bottleneck distribution; and fitting the temporal features of the latency mutation points into a latency curve.
[0068] This paper extracts the inter-layer activation transfer overhead from a hierarchical acceleration strategy. Based on the execution configuration of the hierarchical acceleration strategy, it analyzes in depth the overhead composition of the data activation transfer process between each layer. Inter-layer activation transfer overhead refers to the time and resource consumption incurred when data is transferred between different network layers. When processing aerial images, the features extracted by the convolutional layer need to be transferred to the fully connected layer for classification. This transfer process includes multiple overheads such as data format conversion, memory copy, and synchronization wait. The paper statistically analyzes the amount of activation data transferred between different network layers, including the size of the activation tensor, the transfer frequency, and the transfer direction. Through memory access monitoring, it tracks the movement process and access patterns of activation data between different storage levels. The paper analyzes the time overhead composition of activation transfer, which includes the sum of data copy time, format conversion time, and synchronization wait time. It calculates the proportion of activation transfer overhead in the total execution time, quantifying the impact of the transfer process on overall performance. It distinguishes between necessary transfers and redundant transfers, and statistically analyzes the overhead distribution of each type of transfer.
[0069] Latency sensitivity testing was performed on the activation transmission overhead to obtain bottleneck distribution. Based on the extracted activation transmission overhead, sensitivity testing analysis was used to identify performance bottleneck locations in the transmission process. Transmission performance under different configurations was tested, and the rate of change of transmission latency with respect to various influencing factors was calculated. Bottleneck distribution reflects the spatial distribution of performance limitations in the transmission process. For example, when processing high-resolution surveillance video, data transmission between certain network layers may become a bottleneck due to insufficient memory bandwidth, while transmission between certain layers may become a bottleneck due to complex data format conversion. The key parameters and configuration options with the greatest impact on transmission latency were identified. By traversing each node in the transmission link, the processing capacity and response time of each node were detected to identify performance bottleneck locations. The causes and characteristics of bottlenecks were analyzed, including factors such as hardware limitations, software bottlenecks, and data characteristics. Through bottleneck correlation analysis, combinations of mutually influencing and restrictive bottlenecks were identified to understand the correlation between bottlenecks. The severity and scope of impact of bottlenecks were recorded, and a detailed distribution map of bottleneck locations was established. Accurate identification of bottleneck distribution provides important reference information for subsequent latency mutation analysis.
[0070] Identifying Delay Inflection Points Based on Bottleneck Distribution. Using the obtained bottleneck distribution information, mathematical analysis methods are employed to identify abrupt changes and jumps in the delay curve. Delay inflection points are time points where delay changes drastically during execution. For example, when processing images of densely built-up areas, the computational load increases sharply when the algorithm switches from processing simple flat areas to complex building clusters, forming obvious delay inflection points. The location of these inflection points is detected by calculating the derivative of the delay change, identifying abrupt inflection points. A threshold standard for inflection detection is set, and locations with delay change rates exceeding the threshold are marked as candidate inflection points. Statistical analysis methods are used to distinguish between genuine inflections and noise interference, improving the accuracy of inflection detection. The temporal distribution characteristics of inflection points are analyzed to identify periodic patterns and triggering conditions. Correlation analysis is used to identify inflection points closely related to the bottleneck distribution, filtering out irrelevant random inflections.
[0071] The temporal characteristics of delay mutation points are fitted into a delay curve. Temporal features such as distribution density, interval distance, and trend of change of mutation points on the time axis are extracted. The delay curve fitting process needs to capture the delay change pattern during execution, including different patterns such as gradual delay growth and abrupt delay changes when processing images of varying complexity. An appropriate fitting function type is selected for curve construction, considering the continuity and differentiability requirements of the curve. The delay fitting quality is calculated using the formula Q = 1 - MSE / VAR, where Q is the fitting quality, MSE is the mean squared error, and VAR is the variance of the delay data. The smoothness and continuity of the fitted curve are analyzed to ensure that the curve accurately reflects the true trend of delay change. The fitting parameters are adjusted based on new monitoring data to maintain the accuracy and timeliness of the curve. Outliers and noisy data are processed during the fitting process to improve the stability of the curve. The delay curve fitting completes the transformation from discrete monitoring data to a continuous performance model.
[0072] Inference-accelerated execution instructions are generated based on latency curves. These instructions are control commands that automatically adjust the computation strategy according to latency variation patterns. When the latency curve indicates an impending high-load range, the instructions proactively reduce image processing resolution, simplify algorithm complexity, and increase the number of parallel processing threads to address upcoming performance challenges. By analyzing the morphology of the latency curve, key nodes and inflection points in performance changes are identified. For example, in forest fire prevention patrols, when the latency curve predicts an approach to a densely wooded and complex area, the instructions automatically adjust the accuracy requirements and processing priorities of the detection algorithm. Geometric features of the curve are extracted to guide the instruction generation strategy. Differentiated instruction generation strategies are designed based on the curve features corresponding to different execution scenarios, such as conventional processing instructions for stable latency regions and adaptive adjustment instructions for fluctuating latency regions. A conditional execution mechanism is designed to select the appropriate instruction branch based on real-time performance status, and to switch to a backup instruction sequence when the actual latency deviates from the predicted latency. Ultimately, the generation of inference-accelerated execution instructions achieves a closed-loop process from performance monitoring to intelligent control.
[0073] To implement the low-latency edge-end large model inference acceleration method corresponding to the above method embodiments, in order to achieve the corresponding functions and technical effects. See also Figure 2 , Figure 2 This diagram illustrates a structural block diagram of a low-latency edge-end large-model inference acceleration system 200 provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The low-latency edge-end large-model inference acceleration system 200 provided in this embodiment includes: Resource analysis module 201 is used to acquire UAV computing power resource signals and flight control mission parameters, perform hierarchical sparsity extraction on the computing power resource signals to identify computing bottleneck features, derive time delay constraint coefficients from the flight control mission parameters, and dynamically associate the computing bottleneck features with the time delay constraint coefficients to establish an inference task mapping relationship. Hotspot location module 202 is used to perform model hotspot tracking to locate high-frequency computing cores according to the inference task mapping relationship, perform operator fusion analysis on the high-frequency computing cores to form acceleration paths, collect delay distribution data along the acceleration paths, and map the inference pipeline network based on the delay distribution data; The parallel analysis module 203 is used to perform parallelism analysis on the inference pipeline network to identify branch merging points, extract the throughput parameters of the branch merging points, determine the optimal splitting position by the matching degree between the throughput parameters and the computing bottleneck features, and connect the optimal splitting position with the high-frequency computing core to determine the optimization path. Parameter configuration module 204 is used to reconstruct the cache parameter chain according to the optimization path, align the cache parameter chain with the latency constraint coefficient to determine the low latency processing window, and form an adaptive quantization matrix based on the low latency processing window; Resource optimization module 205 is used to couple the adaptive quantization matrix with the delay distribution data to identify the precision loss region, extract redundant computation in the precision loss region, and establish a resource scheduling matrix based on the redundant computation and the inference pipeline network. The execution control module 206 is used to form a hierarchical acceleration strategy based on the resource scheduling matrix, track the execution jitter of the hierarchical acceleration strategy to form a latency curve, and generate inference acceleration execution instructions based on the latency curve.
[0074] The low-latency edge-end large model inference acceleration system 200 described above can implement a low-latency edge-end large model inference acceleration method according to the above method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining content of this application embodiment can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.
[0075] The above description is only a part or preferred embodiment of this application. Neither the text nor the drawings should limit the scope of protection of this application. All equivalent structural transformations made using the content of this application's specification and drawings under the overall concept of this application, or direct / indirect applications in other related technical fields, are included within the scope of protection of this application.
Claims
1. A low-latency edge-end large model inference acceleration method, characterized in that, include: Acquire UAV computing power resource signals and flight control mission parameters, perform hierarchical sparsity extraction on the computing power resource signals to identify computing bottleneck features, derive time delay constraint coefficients from the flight control mission parameters, and dynamically associate the computing bottleneck features with the time delay constraint coefficients to establish an inference task mapping relationship. Based on the inference task mapping relationship, the model hotspot tracking is used to locate high-frequency computing cores. Operator fusion analysis is performed on the high-frequency computing cores to form acceleration paths. Delay distribution data is collected along the acceleration paths, and the inference pipeline network is mapped based on the delay data. Parallelism analysis is performed on the inference pipeline network to identify branch merging points, the throughput parameters of the branch merging points are extracted, the optimal splitting position is determined by the matching degree between the throughput parameters and the computing bottleneck features, and the optimal splitting position is connected to the high-frequency computing core to determine the optimization path; The cache parameter chain is reconstructed based on the optimization path, the cache parameter chain is aligned with the latency constraint coefficient to determine the low latency processing window, and an adaptive quantization matrix is formed based on the low latency processing window. The adaptive quantization matrix is coupled with the delay distribution data to identify the precision loss region. Redundant computation is extracted from the precision loss region. A resource scheduling matrix is established based on the redundant computation and the inference pipeline network. A hierarchical acceleration strategy is formed based on the resource scheduling matrix, and a latency curve is formed by tracking the execution jitter of the hierarchical acceleration strategy. Inference acceleration execution instructions are generated based on the latency curve.
2. The method according to claim 1, characterized in that, The step of performing hierarchical sparsity extraction and identification of computational bottleneck features on the computing resource signal includes: Construct an operator memory usage graph based on the aforementioned computing power resource signals; Gradient sensitivity analysis is performed on the memory usage graph of the operator to obtain pruning candidate points; Inject sparsity perturbations into the pruning candidate points to identify and activate sparse paths; Structured pruning parameters are extracted and encoded as computational bottleneck features along the activated sparse path.
3. The method according to claim 1, characterized in that, The step of performing model hotspot tracking and locating high-frequency computing kernels based on the inference task mapping relationship includes: Extract the operator dependency graph from the inference task mapping relationship; Apply delay accumulation analysis to the operator dependency graph to obtain the critical path; Operator reordering is performed on the critical path to identify the fusion candidate set; High-frequency computing cores are determined based on the fusion candidate set.
4. The method according to claim 1, characterized in that, The step of performing parallelism analysis to identify branch merging points on the inference pipeline network includes: Tensor parallel scanning is performed on the inference pipeline network to extract the segmentation boundaries; Identify data synchronization overhead points from the segmentation boundaries; Based on the aforementioned synchronization overhead points, a pipeline depth evaluation is performed to generate a parallel efficiency index. The points where the parallel efficiency index exceeds the threshold are marked as branch merging points.
5. The method according to claim 1, characterized in that, The step of aligning the cache parameter chain with the latency constraint coefficient to determine the low-latency processing window includes: Perform memory pooling analysis on the cache parameter chain to identify reuse patterns; The reuse mode and the delay constraint coefficient are dynamically batch processed to obtain the batch configuration; Based on the batch configuration, perform prefetch optimization analysis to determine the buffer boundaries; A low-latency processing window is generated by dividing the window along the buffer boundary.
6. The method according to claim 1, characterized in that, The step of coupling the adaptive quantization matrix with the delay distribution data to identify the accuracy loss region includes: Extract the quantization-sensitive layer distribution from the adaptive quantization matrix; Perform mixed-precision mapping between the quantization-sensitive layer distribution and the delay distribution data to obtain weight sharing points; Incremental inference is applied to the weighted shared points to generate a precision decay response; The accuracy loss zone is determined based on the accuracy attenuation response.
7. The method according to claim 1, characterized in that, The step of establishing a resource scheduling matrix based on the redundant computational load and the inference pipeline network includes: Perform operator-level decomposition on the redundant computational load to obtain a set of prunable operators; Speculative decoding is performed on the set of prunable operators and the inference pipeline network to identify jump opportunities; Based on the aforementioned jump opportunities, computational graph processing is performed to generate scheduling paths; A resource scheduling matrix is established by weaving heterogeneous mapping relationships along the scheduling path.
8. The method according to claim 1, characterized in that, The tracking of the execution jitter of the hierarchical acceleration strategy to form a latency curve includes: Extract the inter-layer activation propagation overhead from the aforementioned hierarchical acceleration strategy; Perform a latency sensitivity test on the activation propagation overhead to obtain the bottleneck distribution; Identify delay mutation points based on the bottleneck distribution; The temporal characteristics of the delayed mutation point are fitted into a delay curve.
9. The method according to claim 3, characterized in that, The determination of high-frequency computing kernels based on the fusion candidate set includes: Based on the fusion candidate set, the fusion benefits are evaluated to determine the optimization priority. The fusion benefits include the reduction in memory access, the improvement in computational density, and the potential for parallelization. Develop an operator fusion strategy according to the aforementioned optimization priorities; High-frequency computing kernels are generated through the operator fusion strategy.
10. A low-latency edge-end large-model inference acceleration system, characterized in that, include: The resource analysis module is used to acquire UAV computing power resource signals and flight control mission parameters, perform hierarchical sparsity extraction on the computing power resource signals to identify computing bottleneck features, derive time delay constraint coefficients from the flight control mission parameters, and dynamically associate the computing bottleneck features with the time delay constraint coefficients to establish an inference task mapping relationship. The hotspot localization module is used to perform model hotspot tracking and localize high-frequency computing cores based on the inference task mapping relationship, perform operator fusion analysis on the high-frequency computing cores to form acceleration paths, collect delay distribution data along the acceleration paths, and map the inference pipeline network based on the delay distribution data. The parallel analysis module is used to perform parallelism analysis on the inference pipeline network to identify branch merging points, extract the throughput parameters of the branch merging points, determine the optimal splitting position by the matching degree between the throughput parameters and the computing bottleneck features, and connect the optimal splitting position to the high-frequency computing core to determine the optimization path. The parameter configuration module is used to reconstruct the cache parameter chain according to the optimization path, align the cache parameter chain with the latency constraint coefficient to determine the low latency processing window, and form an adaptive quantization matrix based on the low latency processing window. The resource optimization module is used to couple the adaptive quantization matrix with the delay distribution data to identify the accuracy loss region, extract redundant computation in the accuracy loss region, and establish a resource scheduling matrix based on the redundant computation and the inference pipeline network. The execution control module is used to form a hierarchical acceleration strategy based on the resource scheduling matrix, track the execution jitter of the hierarchical acceleration strategy to form a latency curve, and generate inference acceleration execution instructions based on the latency curve.
Citation Information
Cited By
Load balancing method and system of low-power AI processor
CN121070629A
Load balancing method and system for low-power ai processor
CN121070629B